Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through September 19, 2026 — 02:17:32 EST

0

Posted in cs.AI · 2026-09-14 · Keertana Chidambaram, Andrew Ilyas, Vasilis Syrgkanis

Corrupt Plans, Clean Traces: Evading Chain-of-Thought Monitoring with Plan Injection

Chain-of-thought (CoT) monitoring is a safety strategy where the reasoning of a large language model "actor" is inspected by a "monitor" (often another language model) for signs of unsafe planning, deception, or misalignment. We find that planting harmful but benign-sounding reasoning in the actor's context can steer it to perform...

💬 0 commentsarXiv:2609.15989v1PDF
0

Posted in cs.RO · 2026-09-14 · Gechen Qu, Tong Zhang, Bike Zhang, Yen-Jen Wang, Koushil Sreenath, Claire Tomlin, Jason Jangho Choi

ResSafe: Learning Safety Filtering with Residual Reinforcement Learning for Humanoids

Safe control of humanoid robots remains challenging due to their high-dimensional dynamics, contact-rich interactions, and sensitivity to disturbances. Although reinforcement learning has enabled effective locomotion and motion tracking, learned policies can still generate unsafe actions that lead to instability or falls. In this...

💬 0 commentsarXiv:2609.15988v1PDF
0

Posted in cs.LG · 2026-09-14 · Zhuoqing Song, Haotian Xu, Xikun Zhang, Lidong Bing

Bellman Policy Optimization

Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models (LLMs). We introduce Bellman Policy Optimization (BPO), a critic-free method derived from Policy Mirror Descent (PMD). For autoregressive generation with terminal rewards, BPO uses the Bellman equations to reformulate PMD...

💬 0 commentsarXiv:2609.15987v1PDF
0

Posted in cs.AI · 2026-09-14 · Honghao Lin, David P. Woodruff, Yuan Deng, Jieming Mao, Song Zuo, Vahab Mirrokni

Stellar Colosseum: A Many-Agent Harness for Long-Horizon Research in Mathematics and Theoretical Computer Science

Language models can produce plausible short proofs, but may still be unreliable on long-horizon research problems, where progress depends on a sequence of uncertain and interdependent decisions. We introduce Stellar Colosseum, a model-agnostic harness for allocating inference across research in mathematics and theoretical computer...

💬 0 commentsarXiv:2609.15983v1PDF
0

Posted in cs.LG · 2026-09-14 · Ruishuo Chen, Xun Wang, Yu Chen, Zhuoran Li, Longbo Huang

The Router Within: Eliciting Native Skill Routing from a Frozen LLM

Skills extend an LLM agent beyond its parametric knowledge, and the gain they promise rests on picking the right one. Deployed harnesses route by preloading every skill's metadata into the context, which disperses the agent's attention and caps the library size. Retrieval pipelines move the selection out of the context, but also out...

💬 0 commentsarXiv:2609.15982v1PDF
0

Posted in cs.DS · 2026-09-14 · Christian Coester, Elias Koutsoupias, Marek Zbysiński

The $k$-server conjecture is true

The $k$-server conjecture states that a deterministic online algorithm can achieve competitive ratio $k$ on every metric space. We prove the conjecture. Specifically, we show that the work function algorithm satisfies it. Our proof uses a natural algebraic representation of the work function as a matrix, which encodes all feasible...

💬 0 commentsarXiv:2609.15979v1PDF
0

Posted in cs.LG · 2026-09-14 · Xingyun Wang, Haomin Zheng, Man Yuan, Leqian Yang, Ziming Liu

A Chosen Future Can Still Be Rewritten: Causal Writability in Video Models

When a video model generates physically incorrect motion, did it fail to learn the correct motion, or did it learn it but fail to use it? We show the latter: the correct motion remains available inside the model and can still be made to control the generated video. We train on videos where red masses oscillate slowly and blue masses...

💬 0 commentsarXiv:2609.15980v1PDF
0

Posted in cs.HC · 2026-09-14 · David Grüning, Jasper Doeninghaus, Zina Efchary, Yui Kondo, Kevin Dunnell, Lennart Fischer, Isabella Zimmermann, Linnea Körte, Leo Mehlig, Frederik Riedel, Paul Schmiedmayer

The CAST-framework: Measure and model social media use as a multi-level phenomenon through real-world applications

Designing social media experiences that support well-being requires understanding when, how, and for whom use matters. Screen-time totals omit content and context, and connecting these with behavior and experience requires coordinating measurements across timescales. We introduce the CAST framework to connect measurement choices with...

💬 0 commentsarXiv:2609.15978v1PDF
0

Posted in cs.IT · 2026-09-14 · Jun Su, Guangyue Han

A Hardy-Space Proof of the Filter-Only Gaussian Feedback-Capacity Formula

The feedback capacity of power-constrained channels with additive stationary Gaussian noise was formulated by Kim in \cite{kim2010feedback} as an infinite-dimensional optimization over the spectrum of an independent stationary Gaussian component and a strictly causal feedback filter. Kim further asserted that the independent...

💬 0 commentsarXiv:2609.15977v1PDF
0

Posted in cs.RO · 2026-09-14 · Anuva Banwasi, William Muckelroy, Priya Sundaresan, Linfeng Zhao, Jeannette Bohg, Cherie Ho

MessyMem: Learning-from-Doing Memory for Mobile Manipulation

Mobile manipulators deployed across many rooms and visits should improve with experience: after discovering that a cabinet is locked or finding an object in a drawer, the robot should reuse that knowledge rather than start each task from scratch. Yet today's robots often treat each task as new: compact scene representations omit...

💬 0 commentsarXiv:2609.15976v1PDF
0

Posted in cs.CL · 2026-09-14 · Shwai He, Haichao Zhang, Shen Yan

Disentangling Representation Evolution in Transformers through Directional Decomposition

Transformer representations evolve through learned additive transformations that either preserve their current direction or redirect it. We study this evolution as a functional geometry, decomposing learned updates into parallel and perpendicular components. Across pretrained models, we find substantial parallel components beyond the...

💬 0 commentsarXiv:2609.15975v1PDF
0

Posted in cs.CL · 2026-09-14 · Ling Yang, Zhenfei Yin, Yingcheng Wu

Discovery Foundation Models: Toward Open-Ended Discovery Intelligence

Foundation models have progressed from learning and reasoning over existing knowledge, to increasingly learning through action, tool use, and outcome feedback. We argue that the next frontier is a further transition: from solving and acting within problems specified by humans to participating in the process by which new problems,...

💬 0 commentsarXiv:2609.15973v1PDF
0

Posted in cs.CL · 2026-09-14 · Zixuan Wang, Yufan Zhou, Jinzhou Tang, Xinle Yu, Chengjun Wu, Lyumanshan Ye, Zhaoxiang Feng, Letian Peng, Adyasha Patra, Fan Bai, Enze Ma, Zhengding Hu, Jianyang Gu, Zhao Wang, Yufei Ding, Jingbo Shang, Tianmin Shu, Zhiting Hu, Zhen Wang

Mind2Dialogue: Training Human-Aware Language Models by Simulating User Mental States

As language models become more capable, long-term collaboration in learning, reasoning, and decision-making calls for a deeper understanding of the people they serve. Yet training such human-aware language models faces a fundamental supervision gap because current datasets for LLM assistant training contain few if any well-informed...

💬 0 commentsarXiv:2609.15972v1PDF
0

Posted in cs.RO · 2026-09-11 · Pietro Gori, Francesco Iotti, Eduard Zelenay, Rastislav Marko, Michele Pierallini, Franco Angelini, Gabriele Pannocchia, Manolo Garabini

Global Path Planner with Multi-Model Switching

This work enhances global path planning via a pure-pursuit controller with multi-model kinematic switching that sustains plan fidelity across diverse terrains. The system includes a traversability graph for terrain analysis, a Heading-Aware A* algorithm for generating feasible paths, and a multi-model Pure Pursuit controller for...

💬 0 commentsarXiv:2609.13015v1PDF
0

Posted in cs.RO · 2026-09-11 · Michel Albonico, Andreas Wortmann, Ivano Malavolta

Tuning ROS 2 for Energy-Efficient Navigation: Empirical Insights from Costmap 2D Configurations

Robots are increasingly used in diverse application areas, where autonomous navigation plays a central role. As these systems become more widespread, improving their energy efficiency is critical to extending operational time and reducing environmental impact. The Robot Operating System (ROS) is a widely adopted middleware for...

💬 0 commentsarXiv:2609.12971v1PDF
0

Posted in cs.SD · 2026-09-11 · Bin Lin, Bo Zhao, Boyang Wang, Boyang Zhang, Boyong Wu, Chao Yan, Chen Geng, Chen Wu, Cheng Yi, Chengli Feng, Chenglin Zhu, DanNi Wan, Daxin Jiang, Dongqing Pang, Fei Tian, Feng Tian, Future Li, Gang Yu, Guanglong Yang, Jia Peng, Jiahao Song, Jiamin Fan, Jiangjie Zhen, Jianzheng Gao, Jun Chen, Li Xie, Lifang Zhang, Lingli Ji, Liying Shi, Lun Cai, Min Xu, Na Wang, Peilin Li, Peng Yang, Pengfei Tan, Qingjian Lin, Ruijie Xiong, Runze Li, Shenghua Hu, Shi Qiu, Siqi Tu, Siyi Zhou, Tianjiao Deng, Wanying Lu, Weiming Niu, Wen Sun, WenWen Qu, Xiangyu Zhang, Xianwei Zhang, XiaoSu Su, Xing Chen, Xinyu Liu, Xuerui Yang, Yang Li, Yang Yang, Yechang Huang, Yibo Zhu, Yifan Zhang, Yiyang Xu, Yu Fu, Yu Luo, Yu Zhou, Yumang Wang, Yunzhou Ju, Yuxiang Yang, Zekai Liu, Zengwei Yao, Zhenwei Mou, Zheqi Dai, Zhiyue Wu, Zichao Zhou

StepAudio 3 Gen Technical Report

We introduce StepAudio 3 Gen, a general-purpose audio generation model that supports zero-shot text-to-speech (TTS), voice design, vocal generation, sound effects, music, vibe speech, and mixtures of multiple audio types within a unified framework. At its core, StepAudio 3 Gen is a discrete autoregressive generator that models audio...

💬 0 commentsarXiv:2609.12945v1PDF
0

Posted in cs.RO · 2026-09-11 · Bojan Derajić, Sebastian Bernhard, Wolfgang Hönig

VertexCBF: Improving Neural Control Barrier Functions via Vertex-Restricted Control Search

As the number of autonomous robots continues to grow, safety becomes increasingly important. Control barrier functions (CBFs) provide a theoretically grounded framework for ensuring safety, but existing design methods often face limitations in effectiveness, scalability, or interpretability, and may result in overly conservative safe...

💬 0 commentsarXiv:2609.12831v1PDF
0

Posted in cs.RO · 2026-09-11 · Nicola Musiu, Francesco Iacovacci, Fausto Lupo, Matteo Pini, Giovanni Scapicchi, Francesco Moretti, Eugenio Mascaro, Pietro Musso, Ayoub Raji, Marko Bertogna, Vincenzo Maria Arricale, Angelo Lo Sapio, Alessandro Piccarelli, Garron Fish

High-Fidelity Multi-Body Simulator for Autonomous Racing

We present a custom high-fidelity vehicle dynamics simulation environment for testing and validation of Autonomous Racing software. The digital twin of the autonomous vehicle is developed in Dymola, using racecar dynamics modeling libraries to build a complete multi-body model. A 3D road surface, including elevation profiles and...

💬 0 commentsarXiv:2609.12795v1PDF
0

Posted in cs.CE · 2026-09-11 · Nikolaos D. Tantaroudas, Ilias Karachalios, Andrew J. McCracken

Transducer Placement and the Limits of a Four-State Reduced Model in Post-Flutter Piezoelectric Energy Harvesting from a Pitch-Plunge-Flap Aerofoil

Aeroelastic ?utter is normally a failure mode to be designed against, yet the limit-cycle oscillations (LCOs) that follow it convert flow energy into sustained structural motion that a piezoelectric transducer can turn into electrical power. A transducer is embedded in a three-degree-of-freedom pitch-plunge aerofoil with a finite-mass...

💬 0 commentsarXiv:2609.12788v1PDF
0

Posted in cs.AI · 2026-09-11 · Jinting Wang, Chenxing Li, Dong Yu, Li Liu

CMA-OT: Hierarchical Expert Supervision for Dance-to-Music Generation

Dance-to-music (D2M) generation aims to synthesize music that is rhythmically and stylistically aligned with dance videos. A key challenge arises from the semantic mismatch between sparse dance cues, such as rhythm and style, and the dense information required for music composition, including structure, instrumentation, and expressive...

💬 0 commentsarXiv:2609.13118v1PDF
0

Posted in cs.CL · 2026-09-11 · Yunqi Lu, Tyler Baumgartner, Nikhil Johri, Brandon Tai, Candice Fan, Luc Debaupte, Ruben Aguilar, Bill Wang, Yi Zhong

Continue, Adapt, or Yield: In-Turn Adaptation to Overlapping Speech in Full-Duplex Agents

Full-duplex evaluation often emphasizes whether an agent keeps speaking or stops. That binary cannot express a third response humans use routinely: continuing to speak while incorporating what the listener just contributed. The contribution may be a missing word, a correction or a clarification. We introduce Duplex Cue, an evaluation...

💬 0 commentsarXiv:2609.13117v1PDF
0

Posted in cs.SE · 2026-09-11 · Michael Neumann, Darja Šmite

Beyond Establishing the Four-Day Workweek: Understanding Adaptation and Long-Term Survival in an Agile Software Organization

Context: Existing research on the four-day workweek (4DWW) has primarily examined its introduction and short-term effects, with limited understanding of its long-term survival or its interaction with agile software development. Objective: We study how a reduced-hour 4DWW is introduced, adapted, institutionalized, and sustained under...

💬 0 commentsarXiv:2609.13089v1PDF
0

Posted in cs.RO · 2026-09-11 · Zhenfeng Gan, Yanbo Chen, Lirong Che, Junbo Tan, Xueqian Wang

ASTRIL-MPC: Autonomous Traversal Framework of Articulated Tracked Robots with Language-Guided Neural-Kinematic MPC

In urban search and rescue, articulated tracked robots (ATRs) must traverse structured but contact-rich environments such as stairwells and cluttered building interiors. Reliable autonomy remains challenging because robot-terrain interaction (RTI) is hybrid and discontinuous, and effective flipper-track coordination is difficult to...

💬 0 commentsarXiv:2609.13083v1PDF
0

Posted in cs.AI · 2026-09-11 · Baoyang Jiang, Fengchun Zhang, Leyuan Wang, Haotian Li, Yida Wang, Zhe Ji, Jinshan Lai, Xi Ren, Danyang Li, Zheng Yang, Jianwei Hu, Qiang Ma

Embodied-BenchForge: A Closed-Loop Agentic Workflow for Embodied Benchmark Construction

Agentic systems offer a promising way to automate embodied benchmark construction, but existing approaches typically cover isolated stages or remain specialized to predefined environments and task families. More importantly, multi-step construction produces dependent intermediate artifacts that are often passed downstream without...

💬 0 commentsarXiv:2609.13082v1PDF
0

Posted in cs.AI · 2026-09-11 · Junghyun Min, Huseyin Uzunalioglu, Mohamed Trabelsi

Autonomous Research for Open-Ended Problems: A Case Study on Telecom Ticket Retrieval

Recent breakthroughs in LLM-based systems and their abilities in problem solving and coding have allowed progress in the AI for Science paradigm, potentially replacing human roles in machine learning (ML) research. However, while several frameworks of fully autonomous end-to-end ML research have been proposed, successful...

💬 0 commentsarXiv:2609.13073v1PDF