Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through September 21, 2026 — 05:16:57 EST

0

Posted in cs.CL · 2026-08-26 · Haitong Luo, Xuying Meng, Weiyao Zhang, Wenji Zou, Shengfeng Lou, Xuefeng Jiang, Chungang Lin, Yujun Zhang

Unveiling Spectral Mechanisms in Training-Free LLM Text Detection

The rapid advancement of Large Language Models (LLMs) makes it increasingly difficult to distinguish human writing from machine-generated text. Training-free detection offers a scalable solution, yet common confidence-based metrics mainly measure average token probabilities and often miss the signal fluctuations that characterize...

💬 0 commentsarXiv:2608.25944v1PDF
0

Posted in cs.HC · 2026-08-26 · Elena Koung, Xinning Gui, Yubo Kou

Gaming Together on Discord: Teen Gamer's Cross-Platform Practices

Discord is one of the most popular communication platforms among gamers. While prior research has highlighted its role in community building, relatively little attention has been paid to its original gaming context-how it shapes gameplay and social experiences. To address this gap, we conducted semi-structured interviews with 16...

💬 0 commentsarXiv:2608.25942v1PDF
0

Posted in cs.LG · 2026-08-26 · Suchit Gupte, Xueru Zhang, Mohammad Mahdi Khalili

When Pruning Meets Interpretability: Preserving Sparse Autoencoder Robustness in LLMs

Sparse autoencoders (SAEs) are widely used to interpret the internal representations of large language models (LLMs), yet their reliability under post-hoc model compression remains poorly understood. We present a systematic study of how pruning affects SAE behavior and theoretically show that, for a fixed SAE, its impact is governed...

💬 0 commentsarXiv:2608.25941v1PDF
0

Posted in cs.RO · 2026-08-26 · Zaruhi Navasardyan, Hrant Davtyan

A Statistical Audit of Physical AI Benchmark Redundancy

Physical AI models are evaluated on suites of benchmarks that differ across model reports, leaving the model-by-benchmark matrix sparse and the relationship between benchmarks unmeasured. We construct a matrix of 51 models on 12 physical AI benchmarks, selected from a registry of 51 benchmarks and 152 models by reporting density,...

💬 0 commentsarXiv:2608.25940v1PDF
0

Posted in cs.SE · 2026-08-26 · Dung Le Quang, Dong Cao Van, Nam Le Hai, Linh Ngo Van, Anh M. T. Bui, Phuong T. Nguyen

XREPOTEST: Benchmarking Multilingual Repository-Level Unit Test Generation for Large Language Models

Large language models (LLMs) have shown promise for automated unit test generation, but existing evaluations largely rely on standalone settings and a narrow set of programming languages, overestimating real-world readiness. We introduce XREPOTEST, a multilingual repository-level benchmark for unit test generation spanning five...

💬 0 commentsarXiv:2608.25939v1PDF
0

Posted in cs.AI · 2026-08-26 · Jia-Hao Ji, Sijie Li, Jiabei Cheng, Zixi She, Jin-Tai Yu, Zhiyuan Yuan

Candidate supply and answer selection shape the value of LLM judging in multi-agent systems

Multi-agent systems (MAS) sometimes already have the potential to answer correctly, but still report a wrong answer. Explaining this outcome is difficult because generation, communication and final answer-selection rules usually change simultaneously. We conceptualize multi-agent reasoning as an evolutionary pipeline of candidate...

💬 0 commentsarXiv:2608.25937v1PDF
0

Posted in cs.LG · 2026-08-26 · Justin Robert, Raheel Qader

One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation

On-policy distillation trains a language model on its own generations while a teacher scores them token by token. It combines the dense supervision of imitation learning with the on-policy sampling of reinforcement learning. But it requires a second, larger model to act as teacher. On-Policy Self-Distillation (OPSD) removes that cost....

💬 0 commentsarXiv:2608.25936v1PDF
0

Posted in cs.CV · 2026-08-26 · Yuqiang Lin, Yan Shi, Sam Lockyer, Harish Tayyar Madabushi, Adrian Evans, Wenbin Li, Yinhai Wang, Nic Zhang

TAU-Agent: An Agentic Retrieval-Augmented Framework for Traffic Anomaly Understanding

Traffic Anomaly Understanding (TAU) requires models and systems to detect, reason about, and explain anomalous events in transportation videos. To address this challenge, we propose TAU-Agent, an agentic retrieval-augmented framework for traffic anomaly understanding. Given a task query, a central retrieval agent orchestrates two...

💬 0 commentsarXiv:2608.25935v1PDF
0

Posted in cs.AI · 2026-08-26 · Aida Usmanova, Zangir Iklassov, Markus Leippold, Ricardo Usbeck

How Robust Are Automated Fact-Checking Systems? A Cross-Benchmark Evaluation

Automated fact-checking (AFC) systems retrieve evidence and predict claim veracity, yet evaluations omit simple baselines, systems are developed for a single benchmark and cannot be trusted to generalise across domains. No prior work cross-evaluates the full two-stage retrieve-then-verify pipeline across diverse datasets,...

💬 0 commentsarXiv:2608.25934v1PDF
0

Posted in cs.CV · 2026-08-26 · Ruoqi Hu, Chulin Zhao, Jiashuo Chang, Ramon Ruiz-Dolz, Hanhe Lin

When Composition Doesn't Add Up: Humans Identifying Defects in AI-Generated Images

*Chulin Zhao and Ruoqi Hu contributed equally to this work. State-of-the-art text-to-image (T2I) models exhibit pronounced and systematic defects when prompts involve intricate compositional factors such as multiple entities and multiple attributes. In this paper, we investigate how humans identify such defects. Specifically, we...

💬 0 commentsarXiv:2608.25933v1PDF
0

Posted in cs.MA · 2026-08-26 · Peter Pak, Victor Alvarado, Amir Barati Farimani

AI Agentic Selective Laser Sintering Process Optimization

Agentic systems enable the intelligent automation of complex workflows, specific to additive manufacturing this is applicable for complex tasks such as process parameter optimization for mechanical properties. This work investigates the AI enabled agentic process optimization within Selective Laser Sintering (SLS) to iteratively...

💬 0 commentsarXiv:2608.25928v1PDF
0

Posted in cs.CV · 2026-08-26 · Yiwen Chen, Guosheng Lin, Chi Zhang

Code World Model: Coding Agent as World Brain

World models aim to simulate how complex environments evolve under actions and events, yet existing video-based world models primarily learn dynamics from visual observations, which reveal outcomes rather than the underlying knowledge, rules, and mechanisms governing world evolution. This makes it difficult to maintain persistent...

💬 0 commentsarXiv:2608.25927v1PDF
0

Posted in cs.AI · 2026-08-26 · Roberto Luvini, Giacomo Longo, Alessandro Armando, Enrico Russo

Formal, Executable and Explainable Runtime Monitoring of Spoken Air Traffic Control Operational Procedures

Air traffic control procedures are executed through spoken exchanges between controllers and pilots. These interactions are essential to the safety of air transportation: failures in their execution can create severe operational hazards, as evidenced by past fatal accidents. Assessing whether an instruction has been followed requires...

💬 0 commentsarXiv:2608.25926v1PDF
0

Posted in cs.LG · 2026-08-25 · Skye Goodman, Roussel Desmond Nzoyem, Leandro Junges, Peter Kissack, Yasser Qureshi, Amberly Brigden, Jeff Clark, Nawid Keshtmand

Evaluating Deep Multivariate Imputation Models on Wearable Device Data

Wearable device data enables continuous health monitoring, but suffers from structured missingness: features sharing a physical sensor drop out together. Deep imputation methods such as BRITS and SAITS have seen limited evaluation on multimodal physiological data under realistic missingness, and existing benchmarks use random-point...

💬 0 commentsarXiv:2608.24436v1PDF
0

Posted in cs.CV · 2026-08-25 · Francisco M. López, Jochen Triesch

Beauty is in the ELBO of the Beholder: A Variational Account of Processing Fluency in Face Perception

Facial attractiveness has been linked to statistical regularities such as symmetry and averageness, suggesting that beauty may depend on the ease with which a face is perceived. We empirically test this hypothesis by training variational autoencoders on four face datasets without attractiveness supervision and evaluating their...

💬 0 commentsarXiv:2608.24219v1PDF
0

Posted in cs.CV · 2026-08-24 · Matteo Dunnhofer, Christian Micheloni, Kohitij Kar

Primate vision reveals a missing principle for robust dynamic AI

How does an intelligent visual system combine what objects look like with how they move while remaining robust as appearance changes? We addressed this question by comparing human perception and neural activity in macaque inferior temporal cortex with representations from image- and video-based neural networks spanning recognition,...

💬 0 commentsarXiv:2608.23790v1PDF
0

Posted in cs.GT · 2026-08-25 · Uriel Feige, Yotam Gafni

Fair Allocation with Optional Selling

We consider fair allocation of indivisible goods in a setting in which agents have subjective valuation functions over the set of goods, and in addition, goods may be sold at given market prices. In this setting, a fair allocation involves {deciding which goods to sell, how to allocate the unsold goods, and how to divide the money...

💬 0 commentsarXiv:2608.24600v1PDF
0

Posted in cs.CY · 2026-08-25 · Jacy Reese Anthis, Erik Brynjolfsson, James Evans

Method, Mind, and Morality: How People Make Sense of Artificial Intelligence

How can humans make sense of the rapid takeoff of artificial intelligence (AI)? We studied the sensemaking dynamics of AI through an open-ended, mixed-methods study with computational text analysis of millions of AI-related newspaper articles and social media posts grounded in 57 semi-structured interviews with AI professionals in...

💬 0 commentsarXiv:2608.24748v1PDF
0

Posted in cs.LG · 2026-08-25 · Yixin Tao, Weiqiang Zheng

Optimal Alternating Regret for Online Learning and Games

We settle the minimax-optimal alternating regret, a regret notion motivated by alternating learning dynamics in games, for both online linear optimization (OLO) and online convex optimization (OCO). For OLO over the probability simplex $Δ_d$, we give an algorithm with $O(\log d)$ alternating regret that remains a constant for any...

💬 0 commentsarXiv:2608.24731v1PDF
0

Posted in cs.CV · 2026-08-25 · Xiaoyan Li, Shixin Xu, Arvind Gupta, Huaxiong Huang

Interpretable Fundus Image Classification via Ring-Based Retinal Vasculature Features

Retinal fundus photography is widely used for screening and monitoring ocular diseases, but many modern classification pipelines rely on deep latent representations and provide limited interpretability. This study develops an interpretable fundus image classification framework based on a ring-structured representation of the retinal...

💬 0 commentsarXiv:2608.24723v1PDF
0

Posted in cs.LG · 2026-08-25 · Claire Chen, Shuze Daniel Liu, Licheng Luo, Rohan Chandra, Nan Jiang, Shangtong Zhang

Robust Data-Collection Policy Learning for Low-Variance Online Policy Evaluation

In reinforcement learning policy evaluation, classic on-policy methods often suffer from high variance when estimating policy performance. To mitigate this issue, behavior policy search has been proposed to learn data-collecting policies tailored to reduce online evaluation variance. However, these approaches do not account for...

💬 0 commentsarXiv:2608.24146v1PDF
0

Posted in cs.LG · 2026-08-25 · Nadeem Shaikh

Knowing When to Ask for Help: Bayesian Self-Escalation in Hierarchical LLM Agents

Current LLM agent systems decide delegation before reasoning begins (a router picks a model) or after a response is complete (a verifier scores it and may retry). We study a third regime: an agent that recognises, during its own reasoning, that it is unlikely to succeed and transfers control to a stronger model. We formulate...

💬 0 commentsarXiv:2608.24087v1PDF
0

Posted in cs.LG · 2026-08-25 · Juntao Fang, Shifeng Xie, Ruichu Cai, Shengji Zheng, Zijian Li, Keli Zhang, Lujia Pan, Themis Palpanas, Zhifeng Hao

ChorusTIC: Training-Free Multivariate Time Series Classification via Chorus In-Context Learning

Time series classification underpins applications in healthcare, sensing, and industrial monitoring. Although time series foundation models support forecasting and transferable representation learning, classification still typically requires fitting a task-specific classifier on each target dataset, while individual channels of...

💬 0 commentsarXiv:2608.24033v1PDF
0

Posted in cs.LG · 2026-08-25 · Amirhesam Abedsoltan, Enric Boix-Adsera, Fivos Kalogiannis, Mikhail Belkin

Revenge of Monosemanticity: Specialized Neurons Improve Data Efficiency in MLPs

Understanding how neural networks learn and organize features is central to understanding their behavior. Much existing theory of feature learning has focused on the emergence of a global low-dimensional predictive geometry. We show that this picture is incomplete. In regression problems with clustered data, we demonstrate that...

💬 0 commentsarXiv:2608.24007v1PDF
0

Posted in cs.LG · 2026-08-25 · Hanna Jiamei Zhang, Alan Papalia, Michael Everett, David M. Rosen

$(\text{DNN})^2$: Doubly Non-Negative Relaxations for Deep Neural Networks

Existing linear program (LP) and semidefinite program (SDP) relaxations for rectified linear unit (ReLU) neural network (NN) verification yield overly-conservative safety guarantees due to significant relaxation gaps. While the completely positive program (CPP) formulation closes this gap, it is NP-hard to solve. Its cheapest...

💬 0 commentsarXiv:2608.24743v1PDF