Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through July 28, 2026 — 22:57:29 EST

0

Posted in cs.AI · 2026-01-08 · Qiang Yu, Xinran Cheng, Chuanyi Liu

Defense Against Indirect Prompt Injection via Tool Result Parsing

As LLM agents transition from digital assistants to physical controllers in autonomous systems and robotics, they face an escalating threat from indirect prompt injection. By embedding adversarial instructions into the results of tool calls, attackers can hijack the agent's decision-making process to execute unauthorized actions. This...

💬 0 commentsarXiv:2601.04795v1PDF
0

Posted in cs.AI · 2026-01-08 · Chengxin Shi, Qinnan Cai, Zeyuan Chen, Long Zeng, Yibo Zhao, Jing Yu, Jianxiang Yu, Xiang Li

APEX: Academic Poster Editing Agentic Expert

Designing academic posters is a labor-intensive process requiring the precise balance of high-density content and sophisticated layout. While existing paper-to-poster generation methods automate initial drafting, they are typically single-pass and non-interactive, often fail to align with complex, subjective user intent. To bridge...

💬 0 commentsarXiv:2601.04794v1PDF
0

Posted in cs.CV · 2026-01-08 · Denis Korzhenkov, Adil Karjauv, Animesh Karnewar, Mohsen Ghafoorian, Amirhossein Habibian

PyramidalWan: On Making Pretrained Video Model Pyramidal for Efficient Inference

Recently proposed pyramidal models decompose the conventional forward and backward diffusion processes into multiple stages operating at varying resolutions. These models handle inputs with higher noise levels at lower resolutions, while less noisy inputs are processed at higher resolutions. This hierarchical approach significantly...

💬 0 commentsarXiv:2601.04792v1PDF
0

Posted in cs.CV · 2026-01-08 · Lee Hyoseok, Sohwi Lim, Eunju Cha, Tae-Hyun Oh

Measurement-Consistent Langevin Corrector for Stabilizing Latent Diffusion Inverse Problem Solvers

While latent diffusion models (LDMs) have emerged as powerful priors for inverse problems, existing LDM-based solvers frequently suffer from instability. In this work, we first identify the instability as a discrepancy between the solver dynamics and stable reverse diffusion dynamics learned by the diffusion model, and show that...

💬 0 commentsarXiv:2601.04791v4PDF
0

Posted in cs.CL · 2026-01-08 · Junhyuk Choi, Jeongyoun Kwon, Heeju Kim, Haeun Cho, Hayeong Jung, Sehee Min, Bugeun Kim

Belief in Authority: Impact of Authority in Multi-Agent Evaluation Framework

Multi-agent systems utilizing large language models often assign authoritative roles to improve performance, yet the impact of authority bias on agent interactions remains underexplored. We present the first systematic analysis of role-based authority bias in free-form multi-agent evaluation using ChatEval. Applying French and Raven's...

💬 0 commentsarXiv:2601.04790v1PDF
0

Posted in cs.CL · 2026-01-08 · Xinyue Peng, Yanming Liu, Yihan Cang, Yuwei Zhang, Xinyi Wang, Songhang Deng, Jiannan Cao

NC2C: Automated Convexification of Generic Non-Convex Optimization Problems

Non-convex optimization problems are pervasive across mathematical programming, engineering design, and scientific computing, often posing intractable challenges for traditional solvers due to their complex objective functions and constrained landscapes. To address the inefficiency of manual convexification and the over-reliance on...

💬 0 commentsarXiv:2601.04789v1PDF
0

Posted in cs.LG · 2026-01-08 · Lang Feng, Fuchao Yang, Feng Chen, Xin Cheng, Haiyang Xu, Zhenglin Wan, Ming Yan, Bo An

AgentOCR: Reimagining Agent History via Optical Self-Compression

Recent advances in large language models (LLMs) enable agentic systems trained with reinforcement learning (RL) over multi-turn interaction trajectories, but practical deployment is bottlenecked by rapidly growing textual histories that inflate token budgets and memory usage. We introduce AgentOCR, a framework that exploits the...

💬 0 commentsarXiv:2601.04786v2PDF
0

Posted in cs.CV · 2026-01-08 · Xihe Qiu, Yang Dai, Xiaoyu Tan, Sijia Li, Fenghao Sun, Lu Gan, Liang Liu

SRU-Pix2Pix: A Fusion-Driven Generator Network for Medical Image Translation with Few-Shot Learning

Magnetic Resonance Imaging (MRI) provides detailed tissue information, but its clinical application is limited by long acquisition time, high cost, and restricted resolution. Image translation has recently gained attention as a strategy to address these limitations. Although Pix2Pix has been widely applied in medical image...

💬 0 commentsarXiv:2601.04785v1PDF
0

Posted in cs.HC · 2026-01-08 · Sophie Villenave, Pierre Raimbaud, Guillaume Lavoué

Dynamic Thermal Feedback in Highly Immersive VR Scenarios: a Multimodal Analysis of User Experience

Thermal feedback is critical to a range of Virtual Reality (VR) applications, such as firefighting training or thermal comfort simulation. Previous studies showed that adding congruent thermal feedback positively influences User eXperience (UX). However, existing work did not compare different levels of thermal feedback quality and...

💬 0 commentsarXiv:2601.04781v1PDF
0

Posted in cs.CV · 2026-01-08 · Akbar Saadat

Defocus Aberration Theory Confirms Gaussian Model in Most Imaging Devices

Over the past three decades, defocus has consistently provided groundbreaking depth information in scene images. However, accurately estimating depth from 2D images continues to be a persistent and fundamental challenge in the field of 3D recovery. Heuristic approaches involve with the ill-posed problem for inferring the spatial...

💬 0 commentsarXiv:2601.04779v1PDF
0

Posted in cs.CV · 2026-01-08 · Tobia Poppi, Burak Uzkent, Amanmeet Garg, Lucas Porto, Garin Kessler, Yezhou Yang, Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara, Florian Schiffers

CounterVid: Counterfactual Video Generation for Mitigating Action and Temporal Hallucinations in Video-Language Models

Video-language models (VLMs) achieve strong multimodal understanding but remain prone to hallucinations, especially when reasoning about actions and temporal order. Existing mitigation strategies, such as textual filtering or random video perturbations, often fail to address the root cause: over-reliance on language priors rather than...

💬 0 commentsarXiv:2601.04778v1PDF
0

Posted in cs.CV · 2026-01-08 · Shurong Zheng, Yousong Zhu, Hongyin Zhao, Fan Yang, Yufei Zhan, Ming Tang, Jinqiao Wang

GeM-VG: Towards Generalized Multi-image Visual Grounding with Multimodal Large Language Models

Multimodal Large Language Models (MLLMs) have demonstrated impressive progress in single-image grounding and general multi-image understanding. Recently, some methods begin to address multi-image grounding. However, they are constrained by single-target localization and limited types of practical tasks, due to the lack of unified...

💬 0 commentsarXiv:2601.04777v1PDF
0

Posted in cs.CV · 2026-01-08 · Jinyu Zhang, Xu Ma, Weili Chen

Segmentation-Driven Monocular Shape from Polarization based on Physical Model

Monocular shape-from-polarization (SfP) leverages the intrinsic relationship between light polarization properties and surface geometry to recover surface normals from single-view polarized images, providing a compact and robust approach for three-dimensional (3D) reconstruction. Despite its potential, existing monocular SfP methods...

💬 0 commentsarXiv:2601.04776v2PDF
0

Posted in cs.AI · 2026-01-08 · Encheng Su, Jianyu Wu, Chen Tang, Lintao Wang, Pengze Li, Aoran Wang, Jinouwen Zhang, Yizhou Wang, Yuan Meng, Xinzhu Ma, Shixiang Tang, Houqiang Li

SciIF: Benchmarking Scientific Instruction Following Towards Rigorous Scientific Intelligence

As large language models (LLMs) transition from general knowledge retrieval to complex scientific discovery, their evaluation standards must also incorporate the rigorous norms of scientific inquiry. Existing benchmarks exhibit a critical blind spot: general instruction-following metrics focus on superficial formatting, while...

💬 0 commentsarXiv:2601.04770v2PDF
0

Posted in cs.SE · 2026-01-08 · Yelena Mujibur Sheikh, Awez Akhtar Khatik, Luoxi Tang, Yuqiao Meng, Zhaohan Xi

RiskBridge: Turning CVEs into Business-Aligned Patch Priorities

Enterprises are confronted with an unprecedented escalation in cybersecurity vulnerabilities, with thousands of new CVEs disclosed each month. Conventional prioritization frameworks such as CVSS offer static severity metrics that fail to account for exploit probability, compliance urgency, and operational impact, resulting in...

💬 0 commentsarXiv:2601.06201v2PDF
0

Posted in cs.CL · 2026-01-08 · Dongjun Kim, Jeongho Yoon, Chanjun Park, Heuiseok Lim

LANGSAE EDITING: Improving Multilingual Information Retrieval via Post-hoc Language Identity Removal

Dense retrieval in multilingual settings often searches over mixed-language collections, yet multilingual embeddings encode language identity alongside semantics. This language signal can inflate similarity for same-language pairs and crowd out relevant evidence written in other languages. We propose LANGSAE EDITING, a post-hoc sparse...

💬 0 commentsarXiv:2601.04768v1PDF
0

Posted in cs.AI · 2026-01-08 · Zefang Zong, Dingwei Chen, Yang Li, Qi Yi, Bo Zhou, Chengming Li, Bo Qian, Peng Chen, Jie Jiang

AT$^2$PO: Agentic Turn-based Policy Optimization via Tree Search

LLM agents have emerged as powerful systems for tackling multi-turn tasks by interleaving internal reasoning and external tool interactions. Agentic Reinforcement Learning has recently drawn significant research attention as a critical post-training paradigm to further refine these capabilities. In this paper, we present AT$^2$PO...

💬 0 commentsarXiv:2601.04767v1PDF
0

Posted in cs.CL · 2026-01-08 · Shengyin Sun, Yiming Li, Renxi Liu, Weizhe Lin, Hui-Ling Zhen, Xianzhi Yu, Mingxuan Yuan, Chen Ma

Revisiting Judge Decoding from First Principles via Training-Free Distributional Divergence

Judge Decoding accelerates LLM inference by relaxing the strict verification of Speculative Decoding, yet it typically relies on expensive and noisy supervision. In this work, we revisit this paradigm from first principles, revealing that the ``criticality'' scores learned via costly supervision are intrinsically encoded in the...

💬 0 commentsarXiv:2601.04766v1PDF
0

Posted in cs.CL · 2026-01-08 · Santiago Acevedo, Alessandro Laio, Marco Baroni

Differential syntactic and semantic encoding in LLMs

We study how syntactic and semantic information is encoded in inner layer representations of Large Language Models (LLMs), focusing on the very large DeepSeek-V3. We find that, by averaging hidden-representation vectors of sentences sharing syntactic structure or meaning, we obtain vectors that capture a significant proportion of the...

💬 0 commentsarXiv:2601.04765v5PDF
0

Posted in cs.AI · 2026-01-08 · Zhen Chen, Weihao Xie, Peilin Chen, Shiqi Wang, Jianping Wang

Orion-RAG: Path-Aligned Hybrid Retrieval for Graphless Data

Retrieval-Augmented Generation (RAG) has proven effective for knowledge synthesis, yet it encounters significant challenges in practical scenarios where data is inherently discrete and fragmented. In most environments, information is distributed across isolated files like reports and logs that lack explicit links. Standard search...

💬 0 commentsarXiv:2601.04764v1PDF
0

Posted in cs.LG · 2026-01-08 · Rupsa Rani Mishra, D. Chandrasekhar Rao, Ajaya Kumar Tripathy

Smart IoT-Based Wearable Device for Detection and Monitoring of Common Cow Diseases Using a Novel Machine Learning Technique

Manual observation and monitoring of individual cows for disease detection present significant challenges in large-scale farming operations, as the process is labor-intensive, time-consuming, and prone to reduced accuracy. The reliance on human observation often leads to delays in identifying symptoms, as the sheer number of animals...

💬 0 commentsarXiv:2601.04761v1PDF
0

Posted in cs.CL · 2026-01-08 · Yehoon Jang, Chaewon Lee, Hyun-seok Min, Sungchul Choi

PILOT-Bench: A Benchmark for Legal Reasoning in the Patent Domain with IRAC-Aligned Classification Tasks

The Patent Trial and Appeal Board (PTAB) of the USPTO adjudicates thousands of ex parte appeals each year, requiring the integration of technical understanding and legal reasoning. While large language models (LLMs) are increasingly applied in patent and legal practice, their use has remained limited to lightweight tasks, with no...

💬 0 commentsarXiv:2601.04758v1PDF
0

Posted in cs.DB · 2026-01-08 · Cristian Riveros, Benjamin Scheidt, Nicole Schweikardt

Structural Indexing of Relational Databases for the Evaluation of Free-Connex Acyclic Conjunctive Queries

We present an index structure to boost the evaluation of free-connex acyclic conjunctive queries (fc-ACQs) over relational databases. The main ingredient of the index associated with a given database $D$ is an auxiliary database $D_{col}$. Our main result states that for any fc-ACQ $Q$ over $D$, we can count the number of answers of...

💬 0 commentsarXiv:2601.04757v1PDF
0

Posted in cs.DS · 2026-01-08 · Tuukka Korhonen, Sang-il Oum

Branch-width of connectivity functions is fixed-parameter tractable

A connectivity function on a finite set $V$ is a symmetric submodular function $f \colon 2^V \to \mathbb{Z}$ with $f(\emptyset)=0$. We prove that finding a branch-decomposition of width at most $k$ for a connectivity function given by an oracle is fixed-parameter tractable (FPT), by providing an algorithm of running time $2^{O(k^2)}...

💬 0 commentsarXiv:2601.04756v2PDF
0

Posted in cs.CV · 2026-01-08 · Yen-Jen Chiou, Wei-Tse Cheng, Yuan-Fu Yang

ProFuse: Efficient Cross-View Context Fusion for Open-Vocabulary 3D Gaussian Splatting

We present ProFuse, an efficient context-aware framework for open-vocabulary 3D scene understanding with 3D Gaussian Splatting (3DGS). The pipeline enhances cross-view consistency and intra-mask cohesion within a direct registration setup, adding minimal overhead and requiring no render-supervised fine-tuning. Instead of relying on a...

💬 0 commentsarXiv:2601.04754v2PDF