Qwen Councils
arXiv is taking too long to respond. Please try again or narrow your search.
Showing downloaded papers while arXiv is unavailable.

Computer Science

arXiv preprints from January 1, 2026 through September 21, 2026 — 22:06:06 EST

0

Posted in cs.CL · 2026-09-01 · Himil Vasava, Ming Jiang

Beyond Scores: Understanding LLM-as-a-Judge Mechanisms in Summarization Evaluation

LLM-based evaluators of natural language generation (NLG) quality are widely deployed as scoring tools and as automated training signals, yet the internal procedure by which they assign a rating remains poorly understood. We investigate this procedure mechanistically through an eight-attack perturbation taxonomy across the Readability...

💬 0 commentsarXiv:2609.01604v1PDF
0

Posted in cs.SE · 2026-09-01 · Kefeng Duan, Dewu Zheng, Yanlin Wang, Xiwen Wang, Ensheng Shi, Xilin Liu, Yuchi Ma, Jiachi Chen, Mingwei Liu, Zibin Zheng

Efficient SWE Agent Benchmarking via Trajectory-Aware Evaluation

Evaluating software engineering agents on realistic benchmarks is costly, since each task may require multi-step code exploration, modification, and test execution. Existing efficient evaluation methods select representative subsets to estimate full-benchmark performance, but are largely result-only: they fit historical pass/fail...

💬 0 commentsarXiv:2609.01603v1PDF
0

Posted in cs.SE · 2026-09-01 · Kefeng Duan, Dewu Zheng, Yanlin Wang, Terry Yue Zhuo, Mingwei Liu, Jianxing Yu, Jiachi Chen, Ensheng Shi, Xilin Liu, Yuchi Ma, Zibin Zheng

Adaptive Critical Token-Aware Retrieval for Repository-Level Code Generation

The repository-level code generation task requires synthesizing code that satisfies task requirements while remaining consistent with the target repository context. Since real-world repositories often exceed the input length limits of LLMs, existing approaches commonly adopt retrieval-augmented generation (RAG) to provide...

💬 0 commentsarXiv:2609.01601v1PDF
0

Posted in cs.CL · 2026-09-01 · Damien Sileo, Dimitri Kachler

CordisBench: Can Language Models Reason About Component Lifecycles in Dynamic Agent Harnesses?

Dynamic agent harnesses let language models change the software that shapes their own execution. This flexibility brings a new reasoning burden: a local plugin change can propagate through dependencies and cleanup. We introduce CordisBench, a 1,200-question benchmark of this lifecycle reasoning. It combines a controlled formal setting...

💬 0 commentsarXiv:2609.01600v1PDF
0

Posted in cs.CV · 2026-09-01 · Asees Kaur, Suzanne S. Sindi, Erica M. Rutter

UI-VISA: U-Net Initialized Vascular Image Segmentation Architecture

Accurate segmentation of vascular structures in digital subtraction angiography (DSA) images remains challenging due to the thin, elongated, and branching nature of blood vessels. Pixel-wise deep learning approaches such as U-Net achieve strong general-purpose segmentation performance but often produce fragmented or discontinuous...

💬 0 commentsarXiv:2609.01598v1PDF
0

Posted in cs.CL · 2026-09-01 · Kshitij Tayal, Arun Sharma, Genta Indra Winata, Anirban Das, Sambit Sahu

The Rise of Verbal Reinforcement Learning

Natural language is emerging as a primary feedback channel for improving language agents, capable of conveying intent, preferences, and causal structure in forms interpretable by both humans and modern language models. We call this paradigm Verbal Reinforcement Learning (VRL) and offer the first unified account of it. We organize the...

💬 0 commentsarXiv:2609.01597v1PDF
0

Posted in cs.RO · 2026-09-01 · Haoyuan Deng, Haichao Liu, Wenkai Guo, Yuan Ling, Zaijia Yang, Yuanjiang Xue, Haosheng Sun, Liangzi Wang, Ziwei Wang

Facet-0: A Robotic Foundation Model for Contact-Rich Precise Manipulation

Real-world robotic assembly at sub-millimeter tolerances demands spatial precision, compliant interaction, and robustness to contact failures. We present Facet-0, a robotic foundation model that predicts and values the contact consequences of its actions. Facet-0 unifies multimodal representation learning and reinforcement learning...

💬 0 commentsarXiv:2609.01596v1PDF
0

Posted in cs.CL · 2026-09-01 · Ke Yang, Chenglong Wang, Michel Galley, Chandan Singh, Jeevana Priya Inala, ChengXiang Zhai, Jianfeng Gao

StudentSim: Training LLM-based Student Simulators

AI tutors are most useful when they adapt to each student's strengths, weaknesses, and preferred guidance, but evidence about which guidance works for which student is sparse, slow, and costly to collect from real learners. Student simulators can provide this signal as a proxy, yet existing approaches are limited: state-tracking...

💬 0 commentsarXiv:2609.01591v1PDF
0

Posted in cs.HC · 2026-09-01 · Chao Zhang, Abe Davis, Chih-Wei Chen, Chin-Chia Hsu

Designing Proactive Thought Partners for Writing

Writing involves diverse cognitive activities, from ideation to revision, and writers' needs vary across individuals and moments. Proactive AI promises to provide the right support at the right time, yet existing proactive tools largely focus on generic textual assistance, such as autocomplete. This paper studies the design space of...

💬 0 commentsarXiv:2609.01588v1PDF
0

Posted in cs.LG · 2026-09-01 · Jundong Hu, Shekar Ramachandran

The Structure of Quantization Damage in LLMs: Why the Next Bit Should Be Spent Globally

Post-training quantization (PTQ) is widely used to reduce the cost of serving large language models (LLMs), but its accuracy cost is uneven and is often tuned per model. We study where quantization damage occurs and how to allocate a small additional precision budget. Using causal mixed-precision intervention as ground truth (raise...

💬 0 commentsarXiv:2609.01587v1PDF
0

Posted in cs.CV · 2026-09-01 · Sergio M. Silva, Otavio T. Remer, Gabriel E. Lima, Lucas Wojcik, Rayson Laroca, David Menotti

A Benchmark for Vehicle Attribute Classification in Cross-Domain Surveillance Scenarios

Vehicle attribute analysis is a key component of Intelligent Transportation Systems (ITS), supporting applications such as vehicle identification, traffic monitoring, and forensic investigation. However, models trained under controlled conditions often degrade in real surveillance scenarios due to changes in viewpoint, occlusion,...

💬 0 commentsarXiv:2609.01584v1PDF
0

Posted in cs.AI · 2026-09-01 · Yingjian Pan, Xiaowei Ding, Kay Giesecke

Agentic Empirical Asset Pricing: Methodological Foundations

Recent advances in LLM agents enable a new paradigm for asset pricing, which we call Agentic Empirical Asset Pricing (AEAP): systems that autonomously conduct the scientific discovery process itself. We define AEAP and identify its core building blocks. Existing evaluation practices backtest only the outputs (factors or trades), not...

💬 0 commentsarXiv:2609.00731v1PDF
0

Posted in cs.SE · 2026-09-01 · Yue Sun, Tong Liu, Yipu Liao, Jingde Chen, Ke Li

Reliable LLM-Generated Programs for High-Energy Physics Experiments through Graph-Grounded Software Knowledge

Extracting physics information from modern particle-physics experiments requires multistage analyses implemented on top of large and highly interconnected software ecosystems. General-purpose large language models (LLMs) often produce unreliable programs for such tasks because a user request alone rarely specifies the required APIs,...

💬 0 commentsarXiv:2609.01095v1PDF
0

Posted in cs.CV · 2026-09-01 · Chujie Qin, Zilong Zhang, Zewei Chang, Chunle Guo, Ruixing Wang, Tao Hu, Ming-Ming Cheng, Chongyi Li

Dotting the Eye: An Intent-Driven Image Retouching Agent for Visual Focus Enhancement

Image retouching is commonly formulated as enhancing overall visual quality through color adjustment, but in practice, it also serves to emphasize visual focus by guiding viewers' attention toward a specific subject or region. Achieving such focus-oriented retouching is inherently challenging, as it requires well-coordinated global...

💬 0 commentsarXiv:2609.01148v1PDF
0

Posted in cs.CV · 2026-09-01 · Chaohao Yuan, Ruifeng Yuan, Zhuoxu Huang, Yu Rong, Hong Cheng, Hou Pong Chan, Chenghao Xiao

On the Design Fundamentals of Pixel Text Representation Learning

Text-rich visual inputs require models that can read, retrieve, and compress language directly in pixel space, yet existing pixel-text encoders struggle with fixed resolution pretraining, visual shortcut learning, weak visual grounding, and multilingual visual text understanding. In this work, we investigate the fundamental design...

💬 0 commentsarXiv:2609.01147v1PDF
0

Posted in cs.CV · 2026-09-01 · Hongtao Kang, Die Luo, Li Chen, Jing Cai, Junbo Hu, Xiuli Liu, Shenghua Cheng

StainPresetNet: Stain Preset Network for Fast Multi-to-Multi Stain Normalization

Stain normalization reduces color variations caused by variations in staining protocols and imaging conditions, thereby enhancing computer-aided diagnostic system performance. Traditional methods derive mapping relationships from individual or limited reference images through pixel-wise transformation, offering style flexibility but...

💬 0 commentsarXiv:2609.01146v1PDF
0

Posted in cs.CV · 2026-09-01 · Michael Zang, Haiyu Wu, Mrinal Sharma, Kevin W. Bowyer

Revisiting Face Recognition for Monozygotic Twins: The Celeb Twins Test Set

Past literature on face recognition for monozygotic (("identical") twins points to facial marks and mirror asymmetry as possible directions for improved accuracy of twins recognition. The Celeb Twins Test Set (CTTS) contains web-scraped image pairs for 80 sets of celebrity twins. It is the only twins test set with meta-data for twins...

💬 0 commentsarXiv:2609.01141v1PDF
0

Posted in cs.CL · 2026-09-01 · Sebastian Steindl, Nikos Voskarides, Alberto Gasparin, Diego Marcheggiani

Does task decomposition improve automatic NLG evaluation?

The LLM-as-a-judge (LLMaJ) framework has emerged as a promising solution for cheap, reproducible, reference-free Natural Language Generation (NLG) evaluation. Prior work seeks to improve LLMaJ by decomposing evaluation tasks into simpler sub-tasks. In this work, we systematically compare LLMaJ methods with and without decomposition on...

💬 0 commentsarXiv:2609.01139v1PDF
0

Posted in cs.CV · 2026-09-01 · Jiyoung Park, InJae Oh, Jung Uk Kim

Different Changes Require Different Reasoning: Change-Type-Specialized Experts for Robust Change Captioning

Change captioning is the task of generating natural language descriptions that explain the changes between a pair of images. Although different change types (e.g., color shifts, object additions) exhibit distinct visual cues and require specialized reasoning processes, existing methods often overlook these distinctions. To address...

💬 0 commentsarXiv:2609.01136v1PDF
0

Posted in cs.CL · 2026-09-01 · Riza Setiawan Soetedjo, Yusuke Sakai, Hidetaka Kamigaito, Katsuhiko Hayashi, Taro Watanabe

Overfitting Mitigation via Singular Value Decomposition in Minimum Bayes Risk Decoding

Minimum Bayes Risk (MBR) decoding enables high-quality text generation by selecting the hypothesis that maximizes a utility metric over sampled pseudo-references. However, it is highly susceptible to metric overfitting: it can irregularly inflate the chosen utility metric at the direct expense of other unoptimized evaluation metrics....

💬 0 commentsarXiv:2609.01135v1PDF
0

Posted in cs.IT · 2026-09-01 · Yi Wang, Linglong Dai

From Source Reconstruction to Predictive State Preservation: An Information-Theoretic Framework for AI-Native Communication

AI-native communication increasingly aims to support prediction rather than reproduce every detail of the source. This shift raises a basic question left implicit by conventional source coding: what should be preserved when the terminal goal is prediction? We take the source-induced predictive state as the fidelity object. It is the...

💬 0 commentsarXiv:2609.01131v1PDF
0

Posted in cs.LG · 2026-09-01 · Jiming Feng, Junliang Li

Scaled Idempotence in Transformer Attention: Paired OV Geometry and Shared-Value Algebras

We identify a recurrent algebraic regularity in Transformer attention: a sparse subset of effective OV operators $T=OV^\top$ nearly closes under composition, $T^2\approxαT$. Across six pretrained endpoints spanning 2.8B--235B parameters, 3.98--8.00% of heads reach squared closure alignment $\mathcal{P}\geq0.9$, while no matched...

💬 0 commentsarXiv:2609.01129v1PDF
0

Posted in cs.LG · 2026-09-01 · Takumi Fujimoto, Hiroaki Nishi

When Does Online Adaptation Pay on the Edge? A Leakage-Free Evaluation of Warmup, Learning-Rate Selection, and Resource Trade-offs for Time-Series Forecasting

Online adaptation can help edge time-series forecasting under distribution drift, but its measured benefit is sensitive to evaluation choices. We study six public multivariate streams, including building-sensor and smart-meter data, under a leakage-free streaming protocol. We identify two additional sources of comparison bias. First,...

💬 0 commentsarXiv:2609.01126v1PDF
0

Posted in cs.SD · 2026-09-01 · Saanvi Raghavendran, Abhishek Bhattacharjee

Artificial Rosetta Stone: Constrained Maximum A Posteriori (MAP) Reconstruction of Symbolic Raga Sequences via Order-k Markov Models

Reconstructing a damaged musical fragment is an inverse problem: the observed sequence contains partial information, while a raga encodes constraints limiting allowable completions. This paper formalizes a mathematical framework for this, proposing the Artificial Rosetta Stone (ARS). We separate three claims often conflated: a...

💬 0 commentsarXiv:2609.01064v1PDF
0

Posted in cs.LG · 2026-09-01 · Satoshi Hayakawa

From Truncation to Commitment: Persistent Context in Uniform Discrete Diffusion

Uniform-state discrete diffusion models update all tokens in parallel while keeping every position revisable. Even when the commonly used top-$p$ rule leaves only one candidate at a position, that choice affects only the current reverse step and can be revised at the next sampling step. We ask what changes when selected hypotheses...

💬 0 commentsarXiv:2609.01043v1PDF