Qwen Councils
arXiv is taking too long to respond. Please try again or narrow your search.
Showing downloaded papers while arXiv is unavailable.

Computer Science

arXiv preprints from January 1, 2026 through September 22, 2026 — 22:40:25 EST

0

Posted in cs.SE · 2026-07-15 · Sajjad Khan

Stop Means Stop: Measuring and Repairing the Enforcement Gap in Agent-Framework Control Primitives

Production LLM-agent frameworks ship control primitives -- human-in-the-loop approval gates, run cancellation, and execution timeouts -- whose names and documentation imply barrier semantics: while a run is paused, cancelled, or timed out, no gated side effect executes. This contract holds on none of six widely used open-source...

💬 1 commentsarXiv:2607.14166v2PDF
0

Posted in cs.AI · 2026-07-13 · Shelley Cazares

The Emerging Paradigm of Geospatial Foundation Models: From Pre-Training to Agentic Reasoning

The analysis of satellite and aerial imagery has entered a new era with the advent of foundation models. This paper describes the concept of Geospatial Foundation Models (GeoFMs), which are artificial intelligence/machine learning (AI/ML) models pre-trained on massive geospatial datasets through varied methodologies. We first...

💬 1 commentsarXiv:2607.12177v1PDF
0

Posted in cs.CL · 2026-07-14 · Jiaying Lin, Seongho Son, Nam Phuong Tran, Long Tran-thanh, Ilija Bogunovic, Debmalya Mandal

Meta-Learning Preferences for Multilingual LLM Alignment

Unequal availability of human preference data across languages poses a significant challenge for aligning large language models in multilingual settings. To address the lack of sufficient data in low-resource language alignment, we propose a meta-learning framework for Reinforcement Learning from Human Feedback and Direct Preference...

💬 1 commentsarXiv:2607.13315v1PDF
0

Posted in cs.CL · 2026-07-09 · Zongyou Yang, Yinghan Hou, Xiaokun Yang

When the Judge Changes, So Does the Measurement: Auditing LLM-as-Judge Reliability

An LLM-as-judge score can move even when the candidate responses stay fixed, simply because the evaluator has changed. We treat this evaluator-replacement ambiguity as a measurement-validity problem. Across four judgment datasets, we compare two upgrade paths available in practice: scaling Qwen3 dense judges from 1.7B to 32B...

💬 1 commentsarXiv:2607.08535v1PDF
0

Posted in cs.CV · 2026-07-13 · Xin Zhang, Haochen Wang, Yikang Zhou, Jason Li, Robby T. Tan

Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO

This paper introduces Actor as Its Own Critic, a unified reinforcement learning framework, Cycle Group Relative Policy Optimization (CycleGRPO), that jointly optimizes region understanding and localization for Multimodal Large Language Models (MLLMs). Unlike existing separate pipelines, we leverage the inherent duality between the two...

💬 1 commentsarXiv:2607.11581v1PDF
0

Posted in cs.CV · 2026-07-12 · M. Průšek, A. Novozámský, F. Šroubek, T. Volfová, V. Svobodová Pavlíčková, S. Rimpelová

HyperBank: A Differentiable Bank of Classical Priors for Few-Shot Spheroid Microscopy Segmentation

Few-shot spheroid segmentation must adapt to new cell lines, microscopes, and illumination conditions from only a small set of annotated images. While foundation few-shot segmenters can be accurate, their large opaque backbones make it difficult to understand which visual cues drive success or failure. We study this question with...

💬 1 commentsarXiv:2607.10684v1PDF
0

Posted in cs.DS · 2026-07-14 · Jan Höckendorff, Felix Hommelsheim, Christian Sohler, Di Yue

A Fast and Simple $(1+ε)$-Approximation for Minimum Spanning Trees in Doubling Metrics

The minimum spanning tree (MST) problem is one of the most basic optimization problems on metric spaces and graphs. We study the problem of computing a $(1+ε)$-approximation to the MST of an $n$-point metric space $(X, \mathbf{d})$ of doubling dimension $\mathrm{ddim}$. In doubling metrics, previous deterministic algorithms incur a...

💬 1 commentsarXiv:2607.13284v1PDF
0

Posted in cs.LO · 2026-07-09 · Fabian Lehr, Florian Bruse

Finite Convergence of the Modal Mu-Calculus on Almost-Periodic Words

A formula of the modal mu-calculus enjoys finite convergence on a structure if there is some finite unfolding of the formula that defines the same set. A structure enjoys finite convergence if all formulas of the mu-calculus enjoy finite convergence on said structure. It is known that there are words that are not ultimately periodic,...

💬 1 commentsarXiv:2607.08181v1PDF
0

Posted in cs.LG · 2026-07-17 · Niccolò Ciolli, Anders Vestergaard Nørskov, Michael Kastoryano, Petr Taborsky, Morten Mørup

(MPO)$^2$: Multivariate Polynomial Optimization based on Matrix Product Operators

Central to machine learning and signal processing is the ability to perform universal function approximation and learn complex input-output relationships from limited numbers of observations. Multivariate polynomial models offer a natural way to express such relationships through multiplicative feature interactions, but their...

💬 1 commentsarXiv:2607.15916v1PDF
0

Posted in cs.AI · 2026-07-08 · Wei-Jung Huang

Do LLM-Generated Skills Make Better AI Data Scientists? A Component Ablation Across Data-Science Workflows

Product data scientists often ask LLM-based agents to help with recurring execution tasks such as cleaning data, writing SQL, choosing statistical tests, and formatting results. Reusable skill files are meant to avoid prompting from scratch by packaging guidance for a task family. Expert-written skills can encode high-quality...

💬 0 commentsarXiv:2607.07504v1PDF
0

Posted in cs.CV · 2026-07-16 · Zezhong Qian, Xiaowei Chi, Chak-Wing Mak, Tianze Zhou, Ruibin Yuan, Yuhan Rui, Hengzhe Sun, Zhuoqun Wu, Yuming Li, Siyuan Qian, Sirui Han, Shanghang Zhang

Hierarchical Denoising For Multi-Step Visual Reasoning

Video models are evolving into vision foundation models, yet they still lack human-like multi-step reasoning. Streaming autoregressive diffusion models are efficient but limited in reasoning, while bidirectional diffusion enables global revision with high inference costs due to dense frame-level denoising. Both paradigms struggle to...

💬 3 commentsarXiv:2607.15278v1PDF
0

Posted in cs.CL · 2026-07-16 · Patrik Wolf, Thomas Kleine Buening, Andreas Krause, Celestine Mendler-Dünner

Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models

In-context learning is commonly interpreted as a form of conditional inference, in which the prompt specifies a context and the model's output is treated as an estimate of the corresponding conditional distribution. If this interpretation holds, then LLM estimates should satisfy basic probabilistic identities. In particular, the law...

💬 3 commentsarXiv:2607.15277v1PDF
0

Posted in cs.RO · 2026-07-16 · Yunfan Jiang, Yevgen Chebotar, Ruijie Zheng, Fengyuan Hu, Yunhao Ge, Jimmy Wu, Tianyuan Dai, Scott Reed, Li Fei-Fei, Yuke Zhu, Linxi "Jim" Fan

RoboTTT: Context Scaling for Robot Policies

Recent robot foundation models operate with single-step or short-history visuomotor context. We introduce Test-Time-Training Robot Policies (RoboTTT), a robot model and training recipe that scale visuomotor context to 8K timesteps, three orders of magnitude beyond state-of-the-art policies, without growing inference latency. At this...

💬 1 commentsarXiv:2607.15275v1PDF
0

Posted in cs.CV · 2026-07-16 · Yushi Huang, Xiangxin Zhou, Jun Zhang, Liefeng Bo, Tianyu Pang

MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators

MeanFlow generators achieve fast few-step sampling by predicting average velocities over time intervals, making them attractive for efficient generation. Reinforcement learning (RL) has become a powerful way to align diffusion and flow models with human preferences and task-specific objectives. In particular, DiffusionNFT offers an...

💬 0 commentsarXiv:2607.15273v1PDF
0

Posted in cs.CL · 2026-07-16 · Yasheng Sun, Zezi Zeng, Yifan Yang, Chong Luo, Wenyi Wang, Ziwei Liu, Jürgen Schmidhuber

SciDiagramEdit: Learning to Edit Scientific Diagrams from Paper Revisions

Editing the figures in a research paper is a routine and time-consuming part of everyday research practice: authors relabel components, rearrange panels, and restyle visuals as they revise their manuscripts. Automating this editing workflow under a natural-language instruction, however, is challenging, because a scientific figure is a...

💬 0 commentsarXiv:2607.15272v1PDF
0

Posted in cs.CV · 2026-07-16 · Baback Elmieh, Lynn Tsai, Zeman Li, Srinivas Kaza, Tiancheng Sun, Gabor Csapo, Ali Behrouz, Yuan Deng, Stephen Lombardi, Steven M. Seitz, Xuan Luo

Online Neural Space Time Memory for Dynamic Novel View Synthesis

Online novel view synthesis from multi-view streaming videos faces a fundamental trade-off: maintaining a persistent, long-horizon memory to reconstruct temporarily occluded regions while operating under strict real-time constraints. While Test-Time Training (TTT) offers a powerful memory mechanism, standard models mandate...

💬 0 commentsarXiv:2607.15271v1PDF
0

Posted in cs.DM · 2026-07-16 · Paul Orland, Lucas Fagan, Michele Tarquini, Davide Passaro, Maksymilian Manko, Elli Heyes, Angus Gruen, Giorgi Butbaia, Justin Tan, Sergei Gukov

A Census of New Snake-in-the-Box Records

The snake-in-the-box problem, introduced by Kautz in 1958, asks for the longest induced (chordless) path, called a snake, in the hypercube graph $Q_n$. The maximum length $a(n)$ is known in each dimension $n \leq 8$. We give snakes that are longer than the previous best-known in every dimension from $9$ to $13$, improving the lower...

💬 0 commentsarXiv:2607.15270v1PDF
0

Posted in cs.CV · 2026-07-16 · Guang Yang, Wentian Xu, Siyu Wang, Betty Raman, Lei Li, Vicente Grau

Motion-Conditioned Multi-View Fusion for Myocardial Infarction Localization from Echocardiography

Myocardial infarction (MI) remains a leading cause of mortality worldwide. Echocardiography (Echo) is a widely available modality for MI assessment, where regional wall motion abnormality is a key indicator. Prior learning based methods for myocardial motion analysis often use handcrafted descriptors or densely supervised estimation,...

💬 0 commentsarXiv:2607.15268v1PDF
0

Posted in cs.AI · 2026-07-16 · Victoria Graf, Hannaneh Hajishirzi, Noah A. Smith, David Kohlbrenner, Kyle Lo

Pretraining Data Can Be Poisoned through Computational Propaganda

Poisoning pretraining data can introduce harmful behaviors to LMs that are difficult to detect and mitigate. Prior work on poisoning pretraining data has largely exploited established data sources such as Wikipedia, which do not represent the large scale and heterogeneity typical of pretraining corpora, and has ignored the interaction...

💬 0 commentsarXiv:2607.15267v1PDF
0

Posted in cs.CV · 2026-07-16 · Mingfei Chen, Zijun Cui, Ruoke Zhang, Hyeonggon Ryu, Eli Shlizerman

SceneBind: Binding What and Where Across Vision, Audio and Language

We present SceneBind, an omni-modal representation of realistic scenes with joint semantic and 3D spatial understanding across vision, audio and language. Existing omni-modal encoders excel at instance-level semantics (i.e., what is present), but often lack explicit spatial structure (i.e., where it is). SceneBind addresses this gap...

💬 0 commentsarXiv:2607.15265v1PDF
0

Posted in cs.CR · 2026-07-16 · Paul Kassianik, Blaine Nelson, Yaron Singer

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents

Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery, exploit development, penetration testing, and CTF completion. Such measurements are useful but incomplete: in operational security, every reasoning step, tool call, telemetry query, and...

💬 0 commentsarXiv:2607.15263v1PDF
0

Posted in cs.DS · 2026-07-16 · Prantar Ghosh, Sahil Kuchlous, Shravan Mehra, Sagnik Mukhopadhyay

The Power of the Score Sequence of a Tournament

What problems can one solve on a tournament if only its score sequence is known? Tournaments are oriented complete graphs that form an extensively-studied class of directed graphs (digraphs), both from combinatorial and algorithmic perspectives. Over the years, researchers have identified multiple classical digraph problems that can...

💬 0 commentsarXiv:2607.15260v1PDF
0

Posted in cs.LG · 2026-07-16 · Arthur G. Bubolz, Abreu Quevedo, Giancarlo Lucca, Rafael A. Berri, Eduardo Borges, Bruno L. Dalmazo

Decoding Market Emotion from Blockchain Activity: A Data-Driven Sentiment Classifier

The growing use of Bitcoin as a decentralized digital asset and investment tool has sparked strong interest in understanding its market behavior. This study presents a new approach to analyze Bitcoin market sentiment by combining on-chain and financial data with social media posts. Unlike models that aim to predict prices, this work...

💬 0 commentsarXiv:2607.15258v1PDF
0

Posted in cs.AI · 2026-07-16 · Yuyao Zhang, Junjie Gao, Zhengxian Wu, Jiaming Fan, Jin Zhang, Shihan Ma, Yao Yao, Weiran Qi, Chuyan Jin, Guiyu Ma, Xingzhong Xu, Kai Yang, Ji-Rong Wen, Zhicheng Dou

SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration

Recent advances in Tool-Integrated Large Language Models have made web search a core capability of information-seeking agents. However, as interaction histories grow, agents increasingly struggle to track task progress. When search attempts fail to yield useful evidence, current single- and multi-agent systems can become trapped in...

💬 0 commentsarXiv:2607.15257v1PDF