Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through July 28, 2026 — 06:48:49 EST

0

Posted in cs.CV · 2026-01-08 · Runze He, Yiji Cheng, Tiankai Hang, Zhimin Li, Yu Xu, Zijin Yin, Shiyi Zhang, Wenxun Dai, Penghui Du, Ao Ma, Chunyu Wang, Qinglin Lu, Jizhong Han, Jiao Dai

Re-Align: Structured Reasoning-guided Alignment for In-Context Image Generation and Editing

In-context image generation and editing (ICGE) enables users to specify visual concepts through interleaved image-text prompts, demanding precise understanding and faithful execution of user intent. Although recent unified multimodal models exhibit promising understanding capabilities, these strengths often fail to transfer...

💬 0 commentsarXiv:2601.05124v1PDF
0

Posted in cs.CV · 2026-01-08 · Zirui Wu, Zeren Jiang, Martin R. Oswald, Jie Song

From Rays to Projections: Better Inputs for Feed-Forward View Synthesis

Feed-forward view synthesis models predict a novel view in a single pass with minimal 3D inductive bias. Existing works encode cameras as Plücker ray maps, which tie predictions to the arbitrary world coordinate gauge and make them sensitive to small camera transformations, thereby undermining geometric consistency. In this paper, we...

💬 0 commentsarXiv:2601.05116v1PDF
0

Posted in cs.AI · 2026-01-08 · Wajid Nasser

Evaluative Fingerprints: Stable and Systematic Differences in LLM Evaluator Behavior

LLM-as-judge systems promise scalable, consistent evaluation. We find the opposite: judges are consistent, but not with each other; they are consistent with themselves. Across 3,240 evaluations (9 judges x 120 unique video x pack items x 3 independent runs), inter-judge agreement is near-zero (Krippendorff's α = 0.042). On two...

💬 0 commentsarXiv:2601.05114v1PDF
0

Posted in cs.CL · 2026-01-08 · Runyang You, Hongru Cai, Caiqi Zhang, Qiancheng Xu, Meng Liu, Tiezheng Yu, Yongqi Li, Wenjie Li

Agent-as-a-Judge

LLM-as-a-Judge has revolutionized AI evaluation by leveraging large language models for scalable assessments. However, as evaluands become increasingly complex, specialized, and multi-step, the reliability of LLM-as-a-Judge has become constrained by inherent biases, shallow single-pass reasoning, and the inability to verify...

💬 0 commentsarXiv:2601.05111v1PDF
0

Posted in cs.AI · 2026-01-08 · Wenhao Zeng, Xuteng Zhang, Yuling Shi, Chao Hu, Yuting Chen, Beijun Shen, Xiaodong Gu

GlimpRouter: Efficient Collaborative Inference by Glimpsing One Token of Thoughts

Large Reasoning Models (LRMs) achieve remarkable performance by explicitly generating multi-step chains of thought, but this capability incurs substantial inference latency and computational cost. Collaborative inference offers a promising solution by selectively allocating work between lightweight and large models, yet a fundamental...

💬 0 commentsarXiv:2601.05110v3PDF
0

Posted in cs.DC · 2026-01-08 · Marco Laju, Donghyun Son, Saurabh Agarwal, Nitin Kedia, Myungjin Lee, Jayanth Srinivasa, Aditya Akella

Nalar: An agent serving framework

LLM-driven agentic applications increasingly automate complex, multi-step tasks, but serving them efficiently remains challenging due to heterogeneous components, dynamic and model-driven control flow, long-running state, and unpredictable latencies. Nalar is a ground-up agent-serving framework that cleanly separates workflow...

💬 0 commentsarXiv:2601.05109v1PDF
0

Posted in cs.DB · 2026-01-08 · Philipp Hanisch, Markus Krötzsch

Rule Rewriting Revisited: A Fresh Look at Static Filtering for Datalog and ASP

Static filtering is a data-independent optimisation method for Datalog, which generalises algebraic query rewriting techniques from relational databases. In spite of its early discovery by Kifer and Lozinskii in 1986, the method has been overlooked in recent research and system development, and special cases are being rediscovered...

💬 0 commentsarXiv:2601.05108v2PDF
0

Posted in cs.AI · 2026-01-08 · Muzhao Tian, Zisu Huang, Xiaohua Wang, Jingwen Xu, Zhengkang Guo, Qi Qian, Yuanzhe Shen, Kaitao Song, Jiakang Yuan, Changze Lv, Xiaoqing Zheng

Controllable Memory Usage: Balancing Anchoring and Innovation in Long-Term Human-Agent Interaction

As LLM-based agents are increasingly used in long-term interactions, cumulative memory is critical for enabling personalization and maintaining stylistic consistency. However, most existing systems adopt an ``all-or-nothing'' approach to memory usage: incorporating all relevant past information can lead to \textit{Memory Anchoring},...

💬 0 commentsarXiv:2601.05107v1PDF
0

Posted in cs.AI · 2026-01-08 · Nuoya Xiong, Yuhang Zhou, Hanqing Zeng, Zhaorun Chen, Furong Huang, Shuchao Bi, Lizhu Zhang, Zhuokai Zhao

Token-Level LLM Collaboration via FusionRoute

Large language models (LLMs) exhibit strengths across diverse domains. However, achieving strong performance across these domains with a single general-purpose model typically requires scaling to sizes that are prohibitively expensive to train and deploy. On the other hand, while smaller domain-specialized models are much more...

💬 0 commentsarXiv:2601.05106v5PDF
0

Posted in cs.RO · 2026-01-08 · Oumaima Barhoumi, Mohamed H Zaki, Sofiène Tahar

Formal Safety Guarantees for Autonomous Vehicles using Barrier Certificates

Modern AI technologies enable autonomous vehicles to perceive complex scenes, predict human behavior, and make real-time driving decisions. However, these data-driven components often operate as black boxes, lacking interpretability and rigorous safety guarantees. Autonomous vehicles operate in dynamic, mixed-traffic environments...

💬 0 commentsarXiv:2601.09740v1PDF
0

Posted in cs.IT · 2026-01-08 · Roxana Smarandache, David G. M. Mitchell

The Number of Cycles of Bi-regular Tanner Graphs in Terms of the Eigenvalues of the Adjacency Matrix

In this paper, we explore new connections between the cycles in the graph of low-density parity-check (LDPC) codes and the eigenvalues of the corresponding adjacency matrix. The resulting observations are used to derive fast, simple, recursive formulas for the number of cycles $N_{2k}$ of length $2k$, $k<g$, in a bi-regular graph of...

💬 0 commentsarXiv:2601.05340v1PDF
0

Posted in cs.CR · 2026-01-08 · Badhan Chandra Das, Md Tasnim Jawad, Joaquin Molto, M. Hadi Amini, Yanzhao Wu

Multi-turn Jailbreaking Attack in Multi-Modal Large Language Models

In recent years, the security vulnerabilities of Multi-modal Large Language Models (MLLMs) have become a serious concern in the Generative Artificial Intelligence (GenAI) research. These highly intelligent models, capable of performing multi-modal tasks with high accuracy, are also severely susceptible to carefully launched security...

💬 0 commentsarXiv:2601.05339v1PDF
0

Posted in cs.AR · 2026-01-08 · Yuval Harary, Almog Sharoni, Esteban Garzón, Marco Lanuzza, Adam Teman, Leonid Yavits

PiC-BNN: A 128-kbit 65 nm Processing-in-CAM-Based End-to-End Binary Neural Network Accelerator

Binary Neural Networks (BNNs), where weights and activations are constrained to binary values (+1, -1), are a highly efficient alternative to traditional neural networks. Unfortunately, typical BNNs, while binarizing linear layers (matrix-vector multiplication), still implement other network layers (batch normalization, softmax,...

💬 0 commentsarXiv:2601.19920v1PDF
0

Posted in cs.RO · 2026-01-08 · Tracey Yee Hsin Tay, Xu Yan, Jonathan Ouyang, Daniel Wu, William Jiang, Jonathan Kao, Yuchen Cui

Intent at a Glance: Gaze-Guided Robotic Manipulation via Foundation Models

Designing intuitive interfaces for robotic control remains a central challenge in enabling effective human-robot interaction, particularly in assistive care settings. Eye gaze offers a fast, non-intrusive, and intent-rich input modality, making it an attractive channel for conveying user goals. In this work, we present GAMMA (Gaze...

💬 0 commentsarXiv:2601.05336v1PDF
0

Posted in cs.LG · 2026-01-08 · Fang Wu, Stan Z. Li

Dynamics-inspired Structure Hallucination for Protein-protein Interaction Modeling

Protein-protein interaction (PPI) represents a central challenge within the biology field, and accurately predicting the consequences of mutations in this context is crucial for drug design and protein engineering. Deep learning (DL) has shown promise in forecasting the effects of such mutations, but is hindered by two primary...

💬 0 commentsarXiv:2601.06214v1PDF
0

Posted in cs.CR · 2026-01-08 · Keerthi Kumar. M, Swarun Kumar Joginpelly, Sunil Khemka, Lakshmi. S R, Navin Chhibber

Cyber Threat Detection and Vulnerability Assessment System using Generative AI and Large Language Model

Background: Cyber-attacks have evolved rapidly in recent years, many individuals and business owners have been affected by cyber-attacks in various ways. Cyber-attacks include various threats such as ransomware, malware, phishing, and Denial of Service (DoS)-related attacks. Challenges: Traditional models such as Generative Artificial...

💬 0 commentsarXiv:2601.06213v1PDF
0

Posted in cs.AI · 2026-01-08 · Tengwei Song, Long Yin, Zhen Han, Zhiqiang Xu

Improving Enzyme Prediction with Chemical Reaction Equations by Hypergraph-Enhanced Knowledge Graph Embeddings

Predicting enzyme-substrate interactions has long been a fundamental problem in biochemistry and metabolic engineering. While existing methods could leverage databases of expert-curated enzyme-substrate pairs for models to learn from known pair interactions, the databases are often sparse, i.e., there are only limited and incomplete...

💬 0 commentsarXiv:2601.05330v1PDF
0

Posted in cs.SD · 2026-01-08 · Junyang Chen, Yuhang Jia, Hui Wang, Jiaming Zhou, Yong Qin

CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models

Automatic speech editing aims to modify spoken content based on textual instructions, yet traditional cascade systems rely on explicit temporal alignment and complex preprocessing. To address these limitations, we propose CosyEdit, an end-to-end speech editing model adapted from CosyVoice through task-specific post-training and a...

💬 0 commentsarXiv:2601.05329v2PDF
0

Posted in cs.CV · 2026-01-08 · Fenil R. Doshi, Thomas Fel, Talia Konkle, George Alvarez

Bi-Orthogonal Factor Decomposition for Vision Transformers

Self-attention is the central computational primitive of Vision Transformers, yet we lack a principled understanding of what information attention mechanisms exchange between tokens. Attention maps describe where weight mass concentrates; they do not reveal whether queries and keys trade position, content, or both. We introduce...

💬 0 commentsarXiv:2601.05328v1PDF
0

Posted in cs.CV · 2026-01-08 · Zeren Jiang, Chuanxia Zheng, Iro Laina, Diane Larlus, Andrea Vedaldi

Mesh4D: 4D Mesh Reconstruction and Tracking from Monocular Video

We propose Mesh4D, a feed-forward model for monocular 4D mesh reconstruction. Given a monocular video of a dynamic object, our model reconstructs the object's complete 3D shape and motion, represented as a deformation field. Our key contribution is a compact latent space that encodes the entire animation sequence in a single pass....

💬 0 commentsarXiv:2601.05251v1PDF
0

Posted in cs.CV · 2026-01-08 · Yuan-Kang Lee, Kuan-Lin Chen, Chia-Che Chang, Yu-Lun Liu

RL-AWB: Deep Reinforcement Learning for Auto White Balance Correction in Low-Light Night-time Scenes

Nighttime color constancy still remains a challenging problem in computational photography due to low-light noise and complex illumination conditions. We present RL-AWB, a novel framework combining statistical methods with deep reinforcement learning for nighttime white balance. Our method begins with a statistical algorithm tailored...

💬 0 commentsarXiv:2601.05249v4PDF
0

Posted in cs.CV · 2026-01-08 · Daniele Lizzio Bosco, Shuteng Wang, Giuseppe Serra, Vladislav Golyanik

QNeRF: Neural Radiance Fields on a Simulated Gate-Based Quantum Computer

Recently, Quantum Visual Fields (QVFs) have shown promising improvements in model compactness and convergence speed for learning the provided 2D or 3D signals. Meanwhile, novel-view synthesis has seen major advances with Neural Radiance Fields (NeRFs), where models learn a compact representation from 2D images to render 3D scenes,...

💬 0 commentsarXiv:2601.05250v1PDF
0

Posted in cs.RO · 2026-01-08 · Zhuoyang Liu, Jiaming Liu, Hao Chen, Jiale Yu, Ziyu Guo, Chengkai Hou, Chenyang Gu, Xiangju Mi, Renrui Zhang, Kun Wu, Zhengping Che, Jian Tang, Pheng-Ann Heng, Shanghang Zhang

LaST$_{0}$: Latent Spatio-Temporal Chain-of-Thought for Robotic Vision-Language-Action Model

Vision-Language-Action (VLA) models have recently shown strong generalization, with some approaches seeking to explicitly generate linguistic reasoning traces or predict future observations prior to execution. However, explicit reasoning typically incurs non-negligible inference latency, which constrains the temporal resolution...

💬 0 commentsarXiv:2601.05248v4PDF
0

Posted in cs.LO · 2026-01-08 · Oskar Fiuk

Random Models and the Guarded Fragment

Building on ideas of Gurevich and Shelah for the Gödel Class, we present a new probabilistic proof of the finite model property for the Guarded Fragment of First-Order Logic. Our proof is conceptually simple and yields the optimal doubly-exponential upper bound on the size of minimal models. We precisely analyse the obtained bound, up...

💬 0 commentsarXiv:2601.05247v2PDF
0

Posted in cs.CV · 2026-01-08 · Gangwei Xu, Haotong Lin, Hongcheng Luo, Haiyang Sun, Bing Wang, Guang Chen, Sida Peng, Hangjun Ye, Xin Yang

Pixel-Perfect Visual Geometry Estimation

Recovering clean and accurate geometry from images is essential for robotics and augmented reality. However, existing geometry foundation models still suffer severely from flying pixels and the loss of fine details. In this paper, we present pixel-perfect visual geometry models that can predict high-quality, flying-pixel-free point...

💬 0 commentsarXiv:2601.05246v1PDF