Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through July 20, 2026 — 16:37:04 EST

0

Posted in cs.SE · 2026-01-21 · Niful Islam, Ragib Shahriar Ayon, Deepak George Thomas, Shibbir Ahmed, Mohammad Wardat

When Agents Fail: A Comprehensive Study of Bugs in LLM Agents with Automated Labeling

Large Language Models (LLMs) have revolutionized intelligent application development. While standalone LLMs cannot perform any actions, LLM agents address the limitation by integrating tools. However, debugging LLM agents is difficult and costly as the field is still in it's early stage and the community is underdeveloped. To...

💬 0 commentsarXiv:2601.15232v2PDF
0

Posted in cs.LO · 2026-01-21 · Edgar F. A. Lederer

How to Verify a Turing Machine with Dafny

This paper describes the formal verification of two Turing machines using the program verifier Dafny. Both machines are deciders, so we prove total correctness. They are typical first examples of Turing machines used in any course of Theoretical Computer Science; in fact, the second machine is literally taken from a relevant textbook....

💬 0 commentsarXiv:2601.15230v1PDF
0

Posted in cs.CL · 2026-01-21 · Warren Johnson

The Perplexity Paradox: Why Code Compresses Better Than Math in LLM Prompts

In "Compress or Route?" (Johnson, 2026), we found that code generation tolerates aggressive prompt compression (r >= 0.6) while chain-of-thought reasoning degrades gradually. That study was limited to HumanEval (164 problems), left the "perplexity paradox" mechanism unvalidated, and provided no adaptive algorithm. This paper addresses...

💬 0 commentsarXiv:2602.15843v1PDF
0

Posted in cs.CV · 2026-01-21 · Yikai Wang, Junqiu Yu, Chenjie Cao, Xiangyang Xue, Yanwei Fu

Aligned Stable Inpainting: Mitigating Unwanted Object Insertion and Preserving Color Consistency

Generative image inpainting can produce realistic, high-fidelity results even with large, irregular masks. However, existing methods still face key issues that make inpainted images look unnatural. In this paper, we identify two main problems: (1) Unwanted object insertion: generative models may hallucinate arbitrary objects in the...

💬 0 commentsarXiv:2601.15368v2PDF
0

Posted in cs.CV · 2026-01-21 · Jianshu Zhang, Chengxuan Qian, Haosen Sun, Haoran Lu, Dingcheng Wang, Letian Xue, Han Liu

PROGRESSLM: Towards Progress Reasoning in Vision-Language Models

Estimating task progress requires reasoning over long-horizon dynamics rather than recognizing static visual content. While modern Vision-Language Models (VLMs) excel at describing what is visible, it remains unclear whether they can infer how far a task has progressed from partial observations. To this end, we introduce...

💬 0 commentsarXiv:2601.15224v2PDF
0

Posted in cs.RO · 2026-01-21 · Stavrow A. Bahnam, Robin Ferede, Till M. Blaha, Anton E. Lang, Erin Lucassen, Quentin Missinne, Aderik E. C. Verraest, Christophe De Wagter, Guido C. H. E. de Croon

MonoRace: Winning Champion-Level Drone Racing with Robust Monocular AI

Autonomous drone racing represents a major frontier in robotics research. It requires an Artificial Intelligence (AI) that can run on board light-weight flying robots under tight resource and time constraints, while pushing the physical system to its limits. The state of the art in this area consists of a system with a stereo camera...

💬 0 commentsarXiv:2601.15222v1PDF
0

Posted in cs.CV · 2026-01-21 · Hanlei Guo, Jiahao Shao, Xinya Chen, Xiyang Tan, Sheng Miao, Yujun Shen, Yiyi Liao

ScenDi: 3D-to-2D Scene Diffusion Cascades for Urban Generation

Recent advancements in 3D object generation using diffusion models have achieved remarkable success, but generating realistic 3D urban scenes remains challenging. Existing methods relying solely on 3D diffusion models tend to suffer a degradation in appearance details, while those utilizing only 2D diffusion models typically...

💬 0 commentsarXiv:2601.15221v1PDF
0

Posted in cs.CL · 2026-01-21 · Anmol Goel, Cornelius Emde, Sangdoo Yun, Seong Joon Oh, Martin Gubri

Privacy Collapse: Benign Fine-Tuning Can Break Contextual Privacy in Language Models

We identify a novel phenomenon in language models: benign fine-tuning of frontier models can lead to privacy collapse. We find that diverse, subtle patterns in training data can degrade contextual privacy, including optimisation for helpfulness, exposure to user information, emotional and subjective dialogue, and debugging code...

💬 0 commentsarXiv:2601.15220v2PDF
0

Posted in cs.HC · 2026-01-21 · Mason Kadem, Sarah Masri, Anthea Innes, Rong Zheng

Human-Centered Ambient and Wearable Sensing for Automated Monitoring in Dementia Care: A Scoping Review

We conducted a scoping review to map the rapidly evolving landscape of wearable and ambient sensing technologies for monitoring people with dementia across home and institutional settings. We analyzed empirical sensing studies (2015-2025) to identify and inform future technical and human-centered design requirements. Five key...

💬 0 commentsarXiv:2603.05516v1PDF
0

Posted in cs.SI · 2026-01-21 · G. Exarchakos, R. van der Hofstad, O. Nagy, M. Pandey

Bringing order to network centrality measures

We introduce a quantitative method to compare arbitrary pairs of graph centrality measures, based on the ordering of vertices induced by them. The proposed method is conceptually simple, mathematically elegant, and allows for a quantitative restatement of many conjectures that were previously cumbersome to formalize. Moreover, it...

💬 0 commentsarXiv:2601.16236v1PDF
0

Posted in cs.LO · 2026-01-21 · Yoshiki Nakamura

A Complete Propositional Dynamic Logic for Regular Expressions with Lookahead

We consider (logical) reasoning for regular expressions with lookahead (REwLA). In this paper, we give an axiomatic characterization for both the (match-)language equivalence and the largest substitution-closed equivalence that is sound for the (match-)language equivalence. To achieve this, we introduce a variant of propositional...

💬 0 commentsarXiv:2601.15214v2PDF
0

Posted in cs.GT · 2026-01-21 · Mayada Oudah, John Wooders

Real-time Facial Communication Restores Cooperation After Defection in Social Dilemmas

Facial expressions are central to human interaction, yet their role in strategic decision-making has received limited attention. We investigate how real-time facial communication influences cooperation in repeated social dilemmas. In a laboratory experiment, participants play a repeated Prisoner's Dilemma game under two conditions: in...

💬 0 commentsarXiv:2601.15211v1PDF
0

Posted in cs.HC · 2026-01-21 · Paige S. DeVries, Michaela Okosi, Ming Li, Nora Dunphy, Gidey Gezae, Dante Conway, Abraham Glasser, Raja Kushalnagar, Christian Vogler

Deaf and Hard of Hearing Access to Intelligent Personal Assistants: Comparison of Voice-Based Options with an LLM-Powered Touch Interface

We investigate intelligent personal assistants (IPAs) accessibility for deaf and hard of hearing (DHH) people who can use their voice in everyday communication. The inability of IPAs to understand diverse accents including deaf speech renders them largely inaccessible to non-signing and speaking DHH individuals. Using an Echo Show, we...

💬 0 commentsarXiv:2601.15209v2PDF
0

Posted in cs.IR · 2026-01-21 · Sangeet Sharma

Beyond the Geometric Curse: High-Dimensional N-Gram Hashing for Dense Retrieval

Why do even the most powerful 7B-parameter embedding models struggle with simple retrieval tasks that the decades old BM25 handles with ease? Recent theory suggests that this happens because of a dimensionality bottleneck. This occurs when we force infinite linguistic nuances into small, fixed-length learned vectors. We developed...

💬 0 commentsarXiv:2601.15205v1PDF
0

Posted in cs.CV · 2026-01-21 · Md Mahmudul Hoque, Shuvo Karmaker, Md. Hadi Al-Amin, Md Modabberul Islam, Jisun Junayed, Farha Ulfat Mahi

A Computer Vision Hybrid Approach: CNN and Transformer Models for Accurate Alzheimer's Detection from Brain MRI Scans

Early and accurate classification of Alzheimers disease (AD) from brain MRI scans is essential for timely clinical intervention and improved patient outcomes. This study presents a comprehensive comparative analysis of five CNN architectures (EfficientNetB0, ResNet50, DenseNet201, MobileNetV3, VGG16), five Transformer-based models...

💬 0 commentsarXiv:2601.15202v1PDF
0

Posted in cs.HC · 2026-01-21 · Bijean Ghafouri, Emilio Ferrara

Lost Before Translation: Social Information Transmission and Survival in AI-AI Communication

When AI systems summarize and relay information, they inevitably transform it. But how? We introduce an experimental paradigm based on the telephone game to study what happens when AI talks to AI. Across five studies tracking content through AI transmission chains, we find three consistent patterns. The first is convergence, where...

💬 0 commentsarXiv:2602.17674v1PDF
0

Posted in cs.CV · 2026-01-21 · Miroslav Purkrabek, Constantin Kolomiiets, Jiri Matas

BBoxMaskPose v2: Expanding Mutual Conditioning to 3D

Most 2D human pose estimation benchmarks are nearly saturated, with the exception of crowded scenes. We introduce PMPose, a top-down 2D pose estimator that incorporates the probabilistic formulation and the mask-conditioning. PMPose improves crowded pose estimation without sacrificing performance on standard scenes. Building on this,...

💬 0 commentsarXiv:2601.15200v1PDF
0

Posted in cs.AI · 2026-01-21 · Shijie Lian, Bin Yu, Xiaopeng Lin, Laurence T. Yang, Zhaolong Shen, Changti Wu, Yuzhuo Miao, Cong Huang, Kai Chen

LangForce: Bayesian Decomposition of Vision Language Action Models via Latent Action Queries

Vision-Language-Action (VLA) models have shown promise in robot manipulation but often struggle to generalize to new instructions or complex multi-task scenarios. We identify a critical pathology in current training paradigms where goal-driven data collection creates a dataset bias. In such datasets, language instructions are highly...

💬 0 commentsarXiv:2601.15197v7PDF
0

Posted in cs.SE · 2026-01-21 · Ramtin Ehsani, Sakshi Pathak, Shriya Rawal, Abdullah Al Mujahid, Mia Mohammad Imran, Preetha Chatterjee

Where Do AI Coding Agents Fail? An Empirical Study of Failed Agentic Pull Requests in GitHub

AI coding agents are now submitting pull requests (PRs) to software projects, acting not just as assistants but as autonomous contributors. As these agentic contributions are rapidly increasing across real repositories, little is known about how they behave in practice and why many of them fail to be merged. In this paper, we conduct...

💬 0 commentsarXiv:2601.15195v1PDF
0

Posted in cs.SE · 2026-01-21 · Stephan Wallraven, Tim Köhne, Hartmut Westenberger, Andreas Moser

Benchmarking Large Language Models for ABAP Code Generation: An Empirical Study on Iterative Improvement by Compiler Feedback

This work investigates the performance of Large Language Models (LLMs) in generating ABAP code. Despite successful applications of generative AI in many programming languages, there are hardly any systematic analyses of ABAP code generation to date. The aim of the study is to empirically analyze to what extent various LLMs can...

💬 0 commentsarXiv:2601.15188v1PDF
0

Posted in cs.CL · 2026-01-21 · Naghmeh Farzi, Laura Dietz, Dave D. Lewis

Supporting Humans in Evaluating AI Summaries of Legal Depositions

While large language models (LLMs) are increasingly used to summarize long documents, this trend poses significant challenges in the legal domain, where the factual accuracy of deposition summaries is crucial. Nugget-based methods have been shown to be extremely helpful for the automated evaluation of summarization approaches. In this...

💬 0 commentsarXiv:2601.15182v1PDF
0

Posted in cs.PL · 2026-01-21 · Pedro Ângelo, Atsushi Igarashi, Yuito Murase, Vasco T. Vasconcelos

Contextual Metaprogramming for Session Types

We propose the integration of staged metaprogramming into a session-typed message passing functional language. We build on a model of contextual modal type theory with multi-level contexts, where contextual values, closing arbitrary terms over a series of variables, may be boxed and transmitted in messages. Once received, one such...

💬 0 commentsarXiv:2601.15180v1PDF
0

Posted in cs.HC · 2026-01-21 · Runlong Ye, Oliver Huang, Patrick Yung Kang Lee, Michael Liut, Carolina Nobre, Ha-Kyung Kong

Reflexis: Supporting Reflexivity and Rigor in Collaborative Qualitative Analysis through Design for Deliberation

Reflexive Thematic Analysis (RTA) is a critical method for generating deep interpretive insights. Yet its core tenets, including researcher reflexivity, tangible analytical evolution, and productive disagreement, are often poorly supported by software tools that prioritize speed and consensus over interpretive depth. To address this...

💬 0 commentsarXiv:2601.15445v2PDF
0

Posted in cs.AI · 2026-01-21 · Alex Goessmann, Janina Schütte, Maximilian Fröhlich, Martin Eigel

A tensor network formalism for neuro-symbolic AI

The unification of neural and symbolic approaches to artificial intelligence remains a central open challenge. In this work, we introduce a tensor network formalism, which captures sparsity principles originating in the different approaches in tensor decompositions. In particular, we describe a basis encoding scheme for functions and...

💬 0 commentsarXiv:2601.15442v1PDF