Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through September 21, 2026 — 21:19:18 EST

0

Posted in cs.RO · 2026-09-01 · Rohit Menon, Shiva Rudra Lolla, Niklas Mueller-Goldingen, Gokul Chenchani, Ribana Roscher, Maren Bennewitz

SG-AMP: Scene-Graph-Guided Active Perception and Semantics-Aware Motion Planning for Pepper Plants

We present SG-AMP, integrating robust depth completion with input-conditioned uncertainty, persistent panoptic mapping, plant scene-graph reasoning, and semantics-aware active view-motion planning. Beyond inspecting uncertain observed regions, the scene graph explicitly hypothesizes unobserved pepper--peduncle attachments and directs...

💬 0 commentsarXiv:2609.01579v1PDF
0

Posted in cs.CL · 2026-09-01 · Maksim Evdokimov, Matvey Ivanov, Dmitrii Tsiupin, Olga Tsymboi, Anatolii Potapov, Aleksandr Ivanov

Closing Cost-Quality Gap in Document VLMs: Difficulty-Aware Data Curation and Quality-Adjusted Deployment Economics

Extracting structured fields from hundreds of millions of documents annually remains costly in regulated industries: bespoke OCR cascades cover only a fraction of workflows, privacy rules preclude external models, and existing open-source VLMs that clear quality thresholds cost more to serve than human annotation. We present a...

💬 0 commentsarXiv:2609.01575v1PDF
0

Posted in cs.CL · 2026-09-01 · Jingtan Wang, Arun Verma, Xiaoqiang Lin, Zhengyuan Liu, Nancy F. Chen, Daniela Rus, Bryan Kian Hsiang Low

Scaling Near-Optimal SFT-RL Annotation Budget Allocation from Small to Large LLMs

How to divide a fixed annotation budget between supervised fine-tuning (SFT) and reinforcement learning (RL) during LLM post-training remains an open problem. Existing work characterizes only broad trends (e.g., SFT dominates in low-data regimes), lacks a principled allocation framework, and does not examine whether the optimal ratio...

💬 0 commentsarXiv:2609.01573v1PDF
0

Posted in cs.CL · 2026-09-01 · Olga Tsymboi, Dmitrii Stoianov, Ramil Latypov, Danil Taranets, Daniil Dryabin, Mikhail Gashkov, Viktor Zelenkovskiy, Aleksandr Fida, Gleb Alektorov, Nikita Gulyakov, Arthur Babkin, Aleksandr Medvedev, Pavel Gein, Anatolii Potapov

From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request Mix

Data-residency constraints force enterprises to self-host LLMs, but continuous adoption of newer models without decommissioning their predecessors expands the serving fleet, fragmenting a finite GPU pool. We consolidate traffic from over 200 internal applications onto a single model by closing quality gaps identified through...

💬 0 commentsarXiv:2609.01572v1PDF
0

Posted in cs.HC · 2026-09-01 · Yulia A. Levites Strekalova, Rachel Liu Galvin, Jessica M. Ray, Samuel P. Border, Mishal Khan, Samantha Hoffman, Christina D. Beharry, Katie Kloss, Philipp Haessner, David Manthey, Sanjay Jain, Michael T. Eadon, Laura Barisoni, Pinaki Sarder

Evaluating Usability in Biomedical Visualization: Rethinking Heuristic Evaluation for Spatial Omics and Multidisciplinary Research Platforms

Introduction: Clinical research informatics (CRI) platforms support biomedical discovery by integrating advanced computational tools into research workflows. Emerging technologies such as spatial omics and AI-enabled imaging expand research capabilities but introduce complex interfaces that increase cognitive burden and alter...

💬 0 commentsarXiv:2609.01569v1PDF
0

Posted in cs.AI · 2026-09-01 · Matteo Merler, Giovanni Bonetta, Davide Zago, Rossella Cancelliere, Bernardo Magnini

Selective Agent Guidance via Entropy: Learning Autonomous Policies from Imperfect VLM Teachers

Vision-Language Models (VLMs) provide useful priors for interactive decision-making, but using them directly as policies is expensive and brittle: they must be queried at every step, do not improve from environment interaction, and can repeat systematic errors. We study how to learn a cheap autonomous policy from an online, expensive,...

💬 0 commentsarXiv:2609.01567v1PDF
0

Posted in cs.CL · 2026-09-01 · Manish Gupta, Chaitanya Giri, Jayasimha Talur

From Confusion to Clarity: Confusion-Aware Retrieval and Knowledge Injection for Text Classification

Large language models (LLMs) struggle to classify text into taxonomies with many semantically similar labels, as the distinctions are domain-specific and not captured by pre-training. To handle large label spaces, a common approach retrieves top-$K$ candidate labels by embedding similarity and prompt the LLM to choose among them....

💬 0 commentsarXiv:2609.01564v1PDF
0

Posted in cs.CL · 2026-09-01 · Ema Salkić, Alexander Fichtl, Philipp Ulrich, Hans Ehm, Marta Bonik, Georg Groh

A systematic Approach to constructing a Chance-and-Risk Matrix for Semiconductor Supply Chains

Semiconductor supply chains face escalating risks from geopolitical tensions, geographic concentration, and rapid technological shifts, yet no scalable system continuously extracts, structures, and prioritizes risk intelligence from public corporate disclosures. We present an end-to-end pipeline that retrieves corporate documents for...

💬 0 commentsarXiv:2609.01563v1PDF
0

Posted in cs.CV · 2026-09-01 · Danze Chen, Zeqing Wang, Ziyue Lin, Xingyi Yang, Yeying Jin

H3-World: Turning Language Understanding into World Control

We present H3-World, an efficient framework that turns the 33B MiniMax-H3 video generator into an interactive world model. Our key finding is that, as large video generators become more capable, language is emerging as a natural interface for control. MiniMax-H3, for example, already supports zero-shot control of character behavior...

💬 0 commentsarXiv:2609.01560v1PDF
0

Posted in cs.LG · 2026-08-31 · Arkadiusz Lipiecki, Rafał Weron

Foundation models for electricity price forecasting and battery arbitrage: Can they replace market-specific forecasting models?

Foundation models promise accurate forecasts with little or no task-specific training, but whether they can replace models designed specifically for electricity price forecasting remains unclear. We compare nine variants from five foundation model families, evaluated in zero-shot mode, with two state-of-the-art electricity price...

💬 0 commentsarXiv:2609.00089v1PDF
0

Posted in cs.NI · 2026-09-01 · Vittorio Todisco, Mattia Andreani, Maria Luisa Merani, Alessandro Bazzi

The Role of Collective Perception and 5G NR-V2X Sidelink in Road Safety

Vehicles and roadside infrastructure are increasingly equipped with sensors capable of perceiving their surroundings. Sharing this information through vehicle-to-everything (V2X) communications is a key enabler of Day-2 applications and is supported by the ETSI collective perception service (CPS). While CPS is expected to play a...

💬 0 commentsarXiv:2609.01478v1PDF
0

Posted in cs.CV · 2026-09-01 · Fatemeh Javadian, Zhu Chen, Zahra Aminparast, Johannes Stegmaier

Semantic-Guided Multimodal Preprocessing for Vision Transformer-Based Clear Cell Renal Cell Carcinoma Grading

Clear cell renal cell carcinoma (CCRCC) grading is essential for treatment planning, yet existing approaches either analyze patch-level images directly or focus solely on nuclei-level classification, without linking to final tumor grading. We propose a semantic-guided multimodal preprocessing method that integrates nuclei...

💬 0 commentsarXiv:2609.01426v1PDF
0

Posted in cs.AI · 2026-09-01 · Danial Noori Zadeh, Mohamed B. Elamien

Analog-DB: An Agent-First Analog Integrated Circuit Database, From Blocks to Systems

Sharing analog integrated circuit designs remains difficult: foundry non-disclosure agreements restrict the process details a design depends on, and the testbenches behind published results are rarely released. We present analog-db, an open-source, versioned database built on a shareable design representation. A domain-specific...

💬 0 commentsarXiv:2609.01286v1PDF
0

Posted in cs.RO · 2026-09-01 · Cheng Zhao, Jingru Zhu, Lei Guo

On Global Regulatability of Robot Manipulators by Classical PID

This paper studies a class of uncertain multi-input multi-output (MIMO) nonlinear systems using extended PID (EPID) control. We focus on systems possessing a well-defined vector relative degree whose components may vary across channels, a setting that received limited attention in the existing literature on PID-type control. We...

💬 0 commentsarXiv:2609.01207v1PDF
0

Posted in cs.CV · 2026-09-01 · Reza Heidari, Hamed R. Tavakoli, Juho Kannala

Compressing AI Traffic: Standardized Neural Network Coding of Visual-Token Representations in Split Vision-Language Inference

When the visual encoder and the language decoder of a vision-language model (VLM) run on different compute nodes, the intermediate visual-token embeddings become a communicated payload rather than an internal activation. We call such machine-consumed intermediate tensors AI traffic and ask how far they can be compressed with a...

💬 0 commentsarXiv:2609.01200v1PDF
0

Posted in cs.IT · 2026-09-01 · Shibsankar Das

Generalized Tan-Arlery-Rabaste-Lehmann-Ovarlez Lower Bound on Ambiguity Function of a Set of Sequences With Mismatched Filters

In this paper, a lower bound on the maximum ambiguity function (AF) sidelobes of a set of unimodular sequences is formulated for the desired low-ambiguity-zone (LAZ). Our main idea is to introduce a set of mismatched filters associated to a set of unimodular sequences and two weight vectors for the delay and Doppler shifts,...

💬 0 commentsarXiv:2609.01112v1PDF
0

Posted in cs.AI · 2026-09-01 · Jierui Zhang, Jianhao Huang, Zhanwei Wang, Kaibin Huang

Space Generative AI with Solar Energy Harvesting

Satellites are emerging as promising platforms to extend generative \emph{artificial intelligence} (AI) services to remote areas lacking terrestrial infrastructure. However, deploying space generative AI is fundamentally constrained by the limited, time-varying onboard energy supplied by solar \emph{energy harvesting} (EH). This paper...

💬 0 commentsarXiv:2609.01062v1PDF
0

Posted in cs.LG · 2026-09-01 · Skanda Athreya, Yutong Wang

One-Layer Transformer Provably Learns Multiclass One-Nearest Neighbor in Context

We extend recent work establishing an equivalence between one-layer transformers and nearest-neighbor classifiers in the binary setting to the multiclass case. By leveraging the simplex encoding, we show that one-layer transformers with an argmax classification head behave identically to a one-nearest-neighbor classifier in the...

💬 0 commentsarXiv:2609.01311v1PDF
0

Posted in cs.LG · 2026-09-01 · W. Ross Morrow

Multi-Head Self Attention is a Parameter Identification Mechanism

We prove that a multi-head scaled dot product attention can be viewed as a parameter identification strategy. The ratio of unidentified parameters to the total number of parameters scales like the reciprocal of the number of heads ($1/2 \to 1/(2H)$), meaning models with more heads are structurally more identified. A subtle side effect...

💬 0 commentsarXiv:2609.01231v1PDF
0

Posted in cs.CV · 2026-09-01 · Penghao Wu, Haiwen Diao, Weichen Fan, Lewei Lu, Dahua Lin, Ziwei Liu

Uncovering Understanding-Generation Synergy in Native Unified Multimodal Models: From Representation, Task to System

While unified multimodal models (UMMs) jointly perform visual understanding and generation within a single model, functional unification does not guarantee learning synergy: the two objectives may reinforce each other, compete for capacity, or merely coexist. We investigate their relationship at the representation, task, and system...

💬 0 commentsarXiv:2609.01607v1PDF
0

Posted in cs.CL · 2026-09-01 · Himil Vasava, Ming Jiang

Beyond Scores: Understanding LLM-as-a-Judge Mechanisms in Summarization Evaluation

LLM-based evaluators of natural language generation (NLG) quality are widely deployed as scoring tools and as automated training signals, yet the internal procedure by which they assign a rating remains poorly understood. We investigate this procedure mechanistically through an eight-attack perturbation taxonomy across the Readability...

💬 0 commentsarXiv:2609.01604v1PDF
0

Posted in cs.SE · 2026-09-01 · Kefeng Duan, Dewu Zheng, Yanlin Wang, Xiwen Wang, Ensheng Shi, Xilin Liu, Yuchi Ma, Jiachi Chen, Mingwei Liu, Zibin Zheng

Efficient SWE Agent Benchmarking via Trajectory-Aware Evaluation

Evaluating software engineering agents on realistic benchmarks is costly, since each task may require multi-step code exploration, modification, and test execution. Existing efficient evaluation methods select representative subsets to estimate full-benchmark performance, but are largely result-only: they fit historical pass/fail...

💬 0 commentsarXiv:2609.01603v1PDF
0

Posted in cs.SE · 2026-09-01 · Kefeng Duan, Dewu Zheng, Yanlin Wang, Terry Yue Zhuo, Mingwei Liu, Jianxing Yu, Jiachi Chen, Ensheng Shi, Xilin Liu, Yuchi Ma, Zibin Zheng

Adaptive Critical Token-Aware Retrieval for Repository-Level Code Generation

The repository-level code generation task requires synthesizing code that satisfies task requirements while remaining consistent with the target repository context. Since real-world repositories often exceed the input length limits of LLMs, existing approaches commonly adopt retrieval-augmented generation (RAG) to provide...

💬 0 commentsarXiv:2609.01601v1PDF
0

Posted in cs.CL · 2026-09-01 · Damien Sileo, Dimitri Kachler

CordisBench: Can Language Models Reason About Component Lifecycles in Dynamic Agent Harnesses?

Dynamic agent harnesses let language models change the software that shapes their own execution. This flexibility brings a new reasoning burden: a local plugin change can propagate through dependencies and cleanup. We introduce CordisBench, a 1,200-question benchmark of this lifecycle reasoning. It combines a controlled formal setting...

💬 0 commentsarXiv:2609.01600v1PDF
0

Posted in cs.CV · 2026-09-01 · Asees Kaur, Suzanne S. Sindi, Erica M. Rutter

UI-VISA: U-Net Initialized Vascular Image Segmentation Architecture

Accurate segmentation of vascular structures in digital subtraction angiography (DSA) images remains challenging due to the thin, elongated, and branching nature of blood vessels. Pixel-wise deep learning approaches such as U-Net achieve strong general-purpose segmentation performance but often produce fragmented or discontinuous...

💬 0 commentsarXiv:2609.01598v1PDF