Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through July 28, 2026 — 13:34:58 EST

0

Posted in cs.AI · 2026-01-07 · Danchun Chen, Qiyao Yan, Liangming Pan

Towards a Mechanistic Understanding of Propositional Logical Reasoning in Large Language Models

Understanding how Large Language Models (LLMs) perform logical reasoning internally remains a fundamental challenge. While prior mechanistic studies focus on identifying taskspecific circuits, they leave open the question of what computational strategies LLMs employ for propositional reasoning. We address this gap through...

💬 0 commentsarXiv:2601.04260v1PDF
0

Posted in cs.NE · 2026-01-07 · Urmzd Mukhammadnaim

Reinforced Linear Genetic Programming

Linear Genetic Programming (LGP) is a powerful technique that allows for a variety of problems to be solved using a linear representation of programs. However, there still exists some limitations to the technique, such as the need for humans to explicitly map registers to actions. This thesis proposes a novel approach that uses...

💬 0 commentsarXiv:2601.09736v1PDF
0

Posted in cs.RO · 2026-01-07 · Samantha Sudhoff, Pranesh Velmurugan, Jiashu Liu, Vincent Zhao, Yung-Hsiang Lu, Kristen Yeon-Ji Yun

From Score to Sound: An End-to-End MIDI-to-Motion Pipeline for Robotic Cello Performance

Robot musicians require precise control to obtain proper note accuracy, sound quality, and musical expression. Performance of string instruments, such as violin and cello, presents a significant challenge due to the precise control required over bow angle and pressure to produce the desired sound. While prior robotic cellists focus on...

💬 0 commentsarXiv:2601.03562v1PDF
0

Posted in cs.CL · 2026-01-07 · Paul Tarau

Modeling Next-Token Prediction as Left-Nested Intuitionistic Implication

We introduce the \emph{Arrow Language Model}, a neural architecture derived from an intuitionistic-logic interpretation of next-token prediction. Instead of representing tokens as additive embeddings mixed by attention, we encode a prefix as a \emph{left-nested implication chain} whose structure preserves order through non-commutative...

💬 0 commentsarXiv:2601.19915v1PDF
0

Posted in cs.LG · 2026-01-07 · Hao Tang, Hao Chen, Hao Li, Chao Li

Neural Operators for Biomedical Spherical Heterogeneity

Spherical deep learning has been widely applied to a broad range of real-world problems. Existing approaches often face challenges in balancing strong spherical geometric inductive biases with the need to model real-world heterogeneity. To solve this while retaining spherical geometry, we first introduce a designable Green's function...

💬 0 commentsarXiv:2601.03561v3PDF
0

Posted in cs.CL · 2026-01-07 · Shidong Cao, Hongzhan Lin, Yuxuan Gu, Ziyang Luo, Jing Ma

DiffCoT: Diffusion-styled Chain-of-Thought Reasoning in LLMs

Chain-of-Thought (CoT) reasoning improves multi-step mathematical problem solving in large language models but remains vulnerable to exposure bias and error accumulation, as early mistakes propagate irreversibly through autoregressive decoding. In this work, we propose DiffCoT, a diffusion-styled CoT framework that reformulates CoT...

💬 0 commentsarXiv:2601.03559v2PDF
0

Posted in cs.SE · 2026-01-07 · Sabrina Haque, Sarvesh Ingale, Christoph Csallner

Do Autonomous Agents Contribute Test Code? A Study of Tests in Agentic Pull Requests

Testing is a critical practice for ensuring software correctness and long-term maintainability. As agentic coding tools increasingly submit pull requests (PRs), it becomes essential to understand how testing appears in these agent-driven workflows. Using the AIDev dataset, we present an empirical study of test inclusion in agentic...

💬 0 commentsarXiv:2601.03556v1PDF
0

Posted in cs.AI · 2026-01-07 · Yuxuan Jiang, Francis Ferraro

SCRIBE: Structured Mid-Level Supervision for Tool-Using Language Models

Training reliable tool-augmented agents remains a significant challenge, largely due to the difficulty of credit assignment in multi-step reasoning. While process-level reward models offer a promising direction, existing LLM-based judges often produce noisy and inconsistent signals because they lack fine-grained, task-specific rubrics...

💬 0 commentsarXiv:2601.03555v3PDF
0

Posted in cs.CL · 2026-01-07 · Sangyub Lee, Heedou Kim, Hyeoncheol Kim

Evaluating LLMs for Police Decision-Making: A Framework Based on Police Action Scenarios

The use of Large Language Models (LLMs) in police operations is growing, yet an evaluation framework tailored to police operations remains absent. While LLM's responses may not always be legally incorrect, their unverified use still can lead to severe issues such as unlawful arrests and improper evidence collection. To address this,...

💬 0 commentsarXiv:2601.03553v1PDF
0

Posted in cs.SI · 2026-01-07 · Lujia Bo, Mingxuan Chen, Youduo Chen, Xiaofan Gui, Jiang Bian, Chunyan Wang, Yi Liu

From Risk Perception to Behavior Large Language Models-Based Simulation of Pandemic Prevention Behaviors

Individual prevention behaviors are a primary line of defense during the early stages of novel infectious disease outbreaks, yet their adoption is heterogeneous and difficult to forecast-especially when empirical data are scarce and epidemic-policy contexts evolve rapidly. To address this gap, we develop an LLM-based...

💬 0 commentsarXiv:2601.03552v1PDF
0

Posted in cs.HC · 2026-01-07 · Michael Yin, Angela Chiang, Robert Xiao

Dissolving a Digital Relationship: A Critical Examination of Digital Severance Behaviours in Close Relationships

Fulfilling social connections are crucial for human well-being and belonging, but not all relationships last forever. As interactions increasingly move online, the act of digitally severing a relationship - e.g. through blocking or unfriending - has become progressively more common as well. This study considers actions of "digital...

💬 0 commentsarXiv:2601.03551v2PDF
0

Posted in cs.AI · 2026-01-07 · Zhizhang Fu, Yuancheng Gu, Chenkai Hu, Hanmeng Liu, Yue Zhang

ReEfBench: Quantifying the Reasoning Efficiency of LLMs

Test-time scaling has enabled Large Language Models (LLMs) to tackle complex reasoning, yet the limitations of current Chain-of-Thought (CoT) evaluation obscures whether performance gains stem from genuine reasoning or mere verbosity. To address this, (1) we propose a novel neuro-symbolic framework for the non-intrusive, comprehensive...

💬 0 commentsarXiv:2601.03550v1PDF
0

Posted in cs.CV · 2026-01-07 · Guobin Tu, Di Weng

FEA-SLT: A Gloss-Free End-to-End Framework for Facial-Expression-Aware Sign Language Translation

Sign Language Translation (SLT) is a challenging cross-modal task requiring joint modeling of manual articulations and non-manual signals. Existing gloss-free SLT methods effectively capture gestural dynamics but often underutilize facial expressions, which play crucial grammatical and disambiguating roles. This limitation can cause...

💬 0 commentsarXiv:2601.03549v2PDF
0

Posted in cs.CL · 2026-01-07 · Guanyu Chen, Chenxiao Yu, Xiyang Hu

Value-Action Alignment in Large Language Models under Privacy-Prosocial Conflict

Large language models (LLMs) are increasingly used to simulate decision-making tasks involving personal data sharing, where privacy concerns and prosocial motivations can push choices in opposite directions. Existing evaluations often measure privacy-related attitudes or sharing intentions in isolation, which makes it difficult to...

💬 0 commentsarXiv:2601.03546v2PDF
0

Posted in cs.CV · 2026-01-07 · Di Xu, Hengjie Liu, Yang Yang, Mary Feng, Jin Ning, Xin Miao, Jessica E. Scholey, Alexandra E. Hotca-cho, William C. Chen, Michael Ohliger, Martina Descovich, Huiming Dong, Wensha Yang, Ke Sheng

B-FIRE: Binning-Free Diffusion Implicit Neural Representation for Hyper-Accelerated Motion-Resolved MRI

Accelerated dynamic volumetric magnetic resonance imaging (4DMRI) is essential for applications relying on motion resolution. Existing 4DMRI produces acceptable artifacts of averaged breathing phases, which can blur and misrepresent instantaneous dynamic information. Recovery of such information requires a new paradigm to reconstruct...

💬 0 commentsarXiv:2601.06166v2PDF
0

Posted in cs.CL · 2026-01-07 · Ye Shen, Dun Pei, Yiqiu Guo, Junying Wang, Yijin Guo, Zicheng Zhang, Qi Jia, Jun Zhou, Guangtao Zhai

EvolMem: A Cognitive-Driven Benchmark for Multi-Session Dialogue Memory

Despite recent advances in understanding and leveraging long-range conversational memory, existing benchmarks still lack systematic evaluation of large language models(LLMs) across diverse memory dimensions, particularly in multi-session settings. In this work, we propose EvolMem, a new benchmark for assessing multi-session memory...

💬 0 commentsarXiv:2601.03543v1PDF
0

Posted in cs.CL · 2026-01-07 · Xukai Liu, Ye Liu, Jipeng Zhang, Yanghai Zhang, Kai Zhang, Qi Liu

Layer-Order Inversion: Rethinking Latent Multi-Hop Reasoning in Large Language Models

Large language models (LLMs) perform well on multi-hop reasoning, yet how they internally compose multiple facts remains unclear. Recent work proposes \emph{hop-aligned circuit hypothesis}, suggesting that bridge entities are computed sequentially across layers before later-hop answers. Through systematic analyses on real-world...

💬 0 commentsarXiv:2601.03542v1PDF
0

Posted in cs.CL · 2026-01-07 · Jin Cui, Jiaqi Guo, Jiepeng Zhou, Ruixuan Yang, Jiayi Lu, Jiajun Xu, Jiangcheng Song, Boran Zhao, Pengju Ren

MIND: From Passive Mimicry to Active Reasoning through Capability-Aware Multi-Perspective CoT Distillation

While Large Language Models (LLMs) have emerged with remarkable capabilities in complex tasks through Chain-of-Thought reasoning, practical resource constraints have sparked interest in transferring these abilities to smaller models. However, achieving both domain performance and cross-domain generalization remains challenging....

💬 0 commentsarXiv:2601.03717v1PDF
0

Posted in cs.CY · 2026-01-07 · François Rottenberg, Thomas Feys, Liesbet Van der Perre

The environmental impact of ICT in the era of data and artificial intelligence

The technology industry promotes artificial intelligence (AI) as a key enabler to solve a vast number of problems, including the environmental crisis. However, when looking at the emissions of datacenters from worldwide service providers, we observe a rapid increase aligned with the advent of AI. Some actors justify it by claiming...

💬 0 commentsarXiv:2601.06174v1PDF
0

Posted in cs.LG · 2026-01-07 · Weijie Shi, Yanxi Chen, Zexi Li, Xuchen Pan, Yuchang Sun, Jiajie Xu, Xiaofang Zhou, Yaliang Li

R$^3$L: Reflect-then-Retry Reinforcement Learning with Language-Guided Exploration, Pivotal Credit, and Positive Amplification

Reinforcement learning drives recent advances in LLM reasoning and agentic capabilities, yet current approaches struggle with both exploration and exploitation. Exploration suffers from low success rates on difficult tasks and high costs of repeated rollouts from scratch. Exploitation suffers from coarse credit assignment and training...

💬 0 commentsarXiv:2601.03715v2PDF
0

Posted in cs.CL · 2026-01-07 · Yunhao Liang, Ruixuan Ying, Bo Li, Hong Li, Kai Yan, Qingwen Li, Min Yang, Okamoto Satoshi, Zhe Cui, Shiwen Ni

Visual Merit or Linguistic Crutch? A Close Look at DeepSeek-OCR

DeepSeek-OCR utilizes an optical 2D mapping approach to achieve high-ratio vision-text compression, claiming to decode text tokens exceeding ten times the input visual tokens. While this suggests a promising solution for the LLM long-context bottleneck, we investigate a critical question: "Visual merit or linguistic crutch - which...

💬 0 commentsarXiv:2601.03714v2PDF
0

Posted in cs.CV · 2026-01-07 · Qingyao Tian, Bingyu Yang, Huai Liao, Xinyan Huang, Junyong Li, Dong Yi, Hongbin Liu

BREATH-VL: Vision-Language-Guided 6-DoF Bronchoscopy Localization via Semantic-Geometric Fusion

Vision-language models (VLMs) have recently shown remarkable performance in navigation and localization tasks by leveraging large-scale pretraining for semantic understanding. However, applying VLMs to 6-DoF endoscopic camera localization presents several challenges: 1) the lack of large-scale, high-quality, densely annotated, and...

💬 0 commentsarXiv:2601.03713v1PDF
0

Posted in cs.CR · 2026-01-07 · Ji Guo, Wenbo Jiang, Yansong Lin, Yijing Liu, Ruichen Zhang, Guomin Lu, Aiguo Chen, Xinshuo Han, Hongwei Li

State Backdoor: Towards Stealthy Real-world Poisoning Attack on Vision-Language-Action Model in State Space

Vision-Language-Action (VLA) models are widely deployed in safety-critical embodied AI applications such as robotics. However, their complex multimodal interactions also expose new security vulnerabilities. In this paper, we investigate a backdoor threat in VLA models, where malicious inputs cause targeted misbehavior while preserving...

💬 0 commentsarXiv:2601.04266v2PDF
0

Posted in cs.CY · 2026-01-07 · Sarah Spiekermann-Hoff, Marc Langheinrich, Johannes Hoff, Christiane Wendehorst, Jürgen Pfeffer, Thomas Fuchs, Armin Grunwald

The Power of 10: New Rules for the Digital World

As artificial intelligence rapidly advances, society is increasingly captivated by promises of superhuman machines and seamless digital futures. Yet these visions often obscure mounting social, ethical, and psychological concerns tied to pervasive digital technologies - from surveillance to mental health crises. This article argues...

💬 0 commentsarXiv:2601.03709v1PDF
0

Posted in cs.PL · 2026-01-07 · Qingyun Zou, Jiahao Cui, Nuo Chen, Bingsheng He, Weng-Fai Wong

MHRC-Bench: A Multilingual Hardware Repository-Level Code Completion benchmark

Large language models (LLMs) have achieved strong performance on code completion tasks in general-purpose programming languages. However, existing repository-level code completion benchmarks focus almost exclusively on software code and largely overlook hardware description languages. In this work, we present \textbf{MHRC-Bench},...

💬 0 commentsarXiv:2601.03708v2PDF