Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through September 22, 2026 — 00:01:07 EST

0

Posted in cs.CL · 2026-01-21 · Brian Christian, Matan Mazor

Self-Blinding and Counterfactual Self-Simulation Mitigate Biases and Sycophancy in Large Language Models

Fair decisions require ignoring irrelevant, potentially biasing, information. To achieve this, decision-makers need to approximate what decision they would have made had they not known certain facts, such as the gender or race of a job candidate. This counterfactual self-simulation is notoriously hard for humans, leading to biased...

💬 0 commentsarXiv:2601.14553v1PDF
0

Posted in cs.RO · 2026-01-21 · Tailai Cheng, Kejia Chen, Lingyun Chen, Liding Zhang, Yue Zhang, Yao Ling, Mahdi Hamad, Zhenshan Bing, Fan Wu, Karan Sharma, Alois Knoll

TacUMI: A Multi-Modal Universal Manipulation Interface for Contact-Rich Tasks

Task decomposition is critical for understanding and learning complex long-horizon manipulation tasks. Especially for tasks involving rich physical interactions, relying solely on visual observations and robot proprioceptive information often fails to reveal the underlying event transitions. This raises the requirement for efficient...

💬 0 commentsarXiv:2601.14550v1PDF
0

Posted in cs.LG · 2026-01-21 · Nilesh Prasad Pandey, Jangseon Park, Onat Gungor, Flavio Ponzina, Tajana Rosing

QMC: Efficient SLM Edge Inference via Outlier-Aware Quantization and Emergent Memories Co-Design

Deploying Small Language Models (SLMs) on edge platforms is critical for real-time, privacy-sensitive generative AI, yet constrained by memory, latency, and energy budgets. Quantization reduces model size and cost but suffers from device noise in emerging non-volatile memories, while conventional memory hierarchies further limit...

💬 0 commentsarXiv:2601.14549v1PDF
0

Posted in cs.IR · 2026-01-21 · Xinyuan Zhang, Lina Zhang, Lisung Chen, Guangyao Liu, Shuai Nie, Jiaming Xu, Runyu Shi, Ying Huang, Guoquan Zhang

Unified Multimodal and Multilingual Retrieval via Multi-Task Learning with NLU Integration

Multimodal retrieval systems typically employ Vision Language Models (VLMs) that encode images and text independently into vectors within a shared embedding space. Despite incorporating text encoders, VLMs consistently underperform specialized text models on text-only retrieval tasks. Moreover, introducing additional text encoders...

💬 0 commentsarXiv:2601.14714v1PDF
0

Posted in cs.AI · 2026-01-21 · Mingxuan Song, Yusen Huo, Bohan Zhou, Shenglin Yin, Zhen Xiao, Jieyi Long, Zhilin Zhang, Chuan Yu

DARA: Few-shot Budget Allocation in Online Advertising via In-Context Decision Making with RL-Finetuned LLMs

Optimizing the advertiser's cumulative value of winning impressions under budget constraints poses a complex challenge in online advertising, under the paradigm of AI-Generated Bidding (AIGB). Advertisers often have personalized objectives but limited historical interaction data, resulting in few-shot scenarios where traditional...

💬 0 commentsarXiv:2601.14711v1PDF
0

Posted in cs.LG · 2026-01-21 · Tianchi Chen, Jan Bima, Sean L. Wu, Otto Ritter, Bingjia Yang, Xiang Yu

Case-Guided Sequential Assay Planning in Drug Discovery

Optimally sequencing experimental assays in drug discovery is a high-stakes planning problem under severe uncertainty and resource constraints. A primary obstacle for standard reinforcement learning (RL) is the absence of an explicit environment simulator or transition data $(s, a, s')$; planning must rely solely on a static database...

💬 0 commentsarXiv:2601.14710v1PDF
0

Posted in cs.HC · 2026-01-21 · Nazar Ponochevnyi, Young-Ho Kim, Joseph Jay Williams, Anastasia Kuzminykh

Talk Me Through It: Developing Effective Systems for Chart Authoring

Recent chart-authoring systems increasingly focus on natural-language input, enabling users to form a mental image of the chart they wish to create and express this intent using spoken instructions (spoken imagined-chart data). Yet these systems are predominantly trained on typed instructions written while viewing the target chart...

💬 0 commentsarXiv:2601.14707v1PDF
0

Posted in cs.CV · 2026-01-21 · Gensmo. ai, Chao Gao, Siqiao Xue, Jiwen Fu, Tingyi Gu, Shanshan Li, Fan Zhou

LookBench: A Live and Holistic Open Benchmark for Fashion Image Retrieval

In this paper, we present LookBench (We use the term "look" to reflect retrieval that mirrors how people shop -- finding the exact item, a close substitute, or a visually consistent alternative.), a live, holistic and challenging benchmark for fashion image retrieval in real e-commerce settings. LookBench includes both recent product...

💬 0 commentsarXiv:2601.14706v3PDF
0

Posted in cs.NE · 2026-01-21 · Casimir Czworkowski, Stephen Hornish, Alhassan S. Yasin

Proximal Policy Optimization with Evolutionary Mutations

Proximal Policy Optimization (PPO) is a widely used reinforcement learning algorithm known for its stability and sample efficiency, but it often suffers from premature convergence due to limited exploration. In this paper, we propose POEM (Proximal Policy Optimization with Evolutionary Mutations), a novel modification to PPO that...

💬 0 commentsarXiv:2601.14705v1PDF
0

Posted in cs.CV · 2026-01-21 · Xinquan Yang, Xuguang Li, Mianjie Zheng, Xuefen Liu, Kun Tang, Kian Ming Lim, He Meng, Jianfeng Ren, Linlin Shen

RegFreeNet: A Registration-Free Network for CBCT-based 3D Dental Implant Planning

As the commercial surgical guide design software usually does not support the export of implant position for pre-implantation data, existing methods have to scan the post-implantation data and map the implant to pre-implantation space to get the label of implant position for training. Such a process is time-consuming and heavily...

💬 0 commentsarXiv:2601.14703v1PDF
0

Posted in cs.AI · 2026-01-21 · Zecong Tang, Zixu Wang, Yifei Wang, Weitong Lian, Tianjian Gao, Haoran Li, Tengju Ru, Lingyi Meng, Zhejun Cui, Yichen Zhu, Qi Kang, Kaixuan Wang, Yu Zhang

Drive-P2D: A Progressive Perception-to-Decision Benchmark for VLMs in Autonomous Driving

Autonomous driving requires reliable perception and safe decision-making in complex scenarios. Recent vision-language models (VLMs) demonstrate reasoning and generalization abilities, opening new possibilities for autonomous driving; however, existing benchmarks often evaluate perception and decision-making separately, limit failure...

💬 0 commentsarXiv:2601.14702v2PDF
0

Posted in cs.CL · 2026-01-21 · Chongxuan Huang, Lei Lin, Xiaodong Shi, Wenping Hu, Ruiming Tang

DARL: Encouraging Diverse Answers for General Reasoning without Verifiers

Reinforcement Learning with Verifiable Rewards (RLVR) has demonstrated promising gains in enhancing the reasoning capabilities of large language models. However, its dependence on domain-specific verifiers significantly restricts its applicability to open and general domains. Recent efforts such as RLPR have extended RLVR to general...

💬 0 commentsarXiv:2601.14700v1PDF
0

Posted in cs.NE · 2026-01-21 · Ziqing Li, Myung Cho, Qiutong Jin, Weiyu Xu

Repair Brain Damage: Real-Numbered Error Correction Code for Neural Network

We consider a neural network (NN) that may experience memory faults and computational errors. In this paper, we propose a novel real-number-based error correction code (ECC) capable of detecting and correcting both memory errors and computational errors. The proposed approach introduces structures in the form of real-number-based...

💬 0 commentsarXiv:2602.00076v1PDF
0

Posted in cs.CL · 2026-01-21 · Michael Theologitis, Preetam Prabhu Srikar Dammu, Chirag Shah, Dan Suciu

ClaimDB: A Fact Verification Benchmark over Large Structured Data

Real-world fact-checking often involves verifying claims grounded in structured data at scale. Despite substantial progress in fact-verification benchmarks, this setting remains largely underexplored. In this work, we introduce ClaimDB, a fact-verification benchmark where the evidence for claims is derived from compositions of...

💬 0 commentsarXiv:2601.14698v2PDF
0

Posted in cs.IR · 2026-01-21 · Shutong Qiao, Wei Yuan, Tong Chen, Xiangyu Zhao, Quoc Viet Hung Nguyen, Hongzhi Yin

When Text-as-Vision Meets Semantic IDs in Generative Recommendation: An Empirical Study

Semantic ID learning is a key interface in Generative Recommendation (GR) models, mapping items to discrete identifiers grounded in side information, most commonly via a pretrained text encoder. However, these text encoders are primarily optimized for well-formed natural language. In real-world recommendation data, item descriptions...

💬 0 commentsarXiv:2601.14697v1PDF
0

Posted in cs.CL · 2026-01-21 · Zhaiyu Fang, Ruipeng Sun

AdaTIR: Adaptive Tool-Integrated Reasoning via Difficulty-Aware Policy Optimization

Tool-Integrated Reasoning (TIR) has significantly enhanced the capabilities of Large Language Models (LLMs), yet current agents tend to exhibit cognitive offloading, redundantly invoking external tools even for simple tasks. In this paper, we suggest that true agentic intelligence requires not just tool invocation, but the adaptive...

💬 0 commentsarXiv:2601.14696v1PDF
0

Posted in cs.LG · 2026-01-21 · Yutong Chen, Jiandong Gao, Ji Wu

CoScale-RL: Efficient Post-Training by Co-Scaling Data and Computation

Training Large Reasoning Model (LRM) is usually unstable and unpredictable, especially on hard problems or weak foundation models. We found that the current post-training scaling strategy can still improve on these cases. We propose CoScale-RL, a novel scaling strategy with better data and computational efficiency. We first scale up...

💬 0 commentsarXiv:2601.14695v1PDF
0

Posted in cs.LG · 2026-01-21 · Pengfei Ding, Yan Wang, Guanfeng Liu

Re-understanding Graph Unlearning through Memorization

Graph unlearning (GU), which removes nodes, edges, or features from trained graph neural networks (GNNs), is crucial in Web applications where graph data may contain sensitive, mislabeled, or malicious information. However, existing GU methods lack a clear understanding of the key factors that determine unlearning effectiveness,...

💬 0 commentsarXiv:2601.14694v1PDF
0

Posted in cs.LG · 2026-01-21 · Jianwen Sun, Xinrui Li, Fuqing Li, Xiaoxuan Shen

Beyond Error-Based Optimization: Experience-Driven Symbolic Regression with Goal-Conditioned Reinforcement Learning

Symbolic Regression aims to automatically identify compact and interpretable mathematical expressions that model the functional relationship between input and output variables. Most existing search-based symbolic regression methods typically rely on the fitting error to inform the search process. However, in the vast expression space,...

💬 0 commentsarXiv:2601.14693v1PDF
0

Posted in cs.AI · 2026-01-21 · Muhammad Khalifa, Lajanugen Logeswaran, Jaekyeom Kim, Sungryull Sohn, Yunxiang Zhang, Moontae Lee, Hao Peng, Lu Wang, Honglak Lee

Gaming the Judge: Unfaithful Chain-of-Thought Can Undermine Agent Evaluation

Large language models (LLMs) are increasingly used as judges to evaluate agent performance, particularly in non-verifiable settings where judgments rely on agent trajectories including chain-of-thought (CoT) reasoning. This paradigm implicitly assumes that the agent's CoT faithfully reflects both its internal reasoning and the...

💬 0 commentsarXiv:2601.14691v2PDF
0

Posted in cs.CV · 2026-01-21 · Yian Huang, Qing Qin, Aji Mao, Xiangyu Qiu, Liang Xu, Xian Zhang, Zhenming Peng

FeedbackSTS-Det: Sparse Frames-Based Spatio-Temporal Semantic Feedback Network for Moving Infrared Small Target Detection

Infrared small target detection (ISTD) has been a critical technology in defense and civilian applications over the past several decades, such as missile warning, maritime surveillance, and disaster monitoring. Nevertheless, moving infrared small target detection still faces considerable challenges: existing models suffer from...

💬 0 commentsarXiv:2601.14690v2PDF
0

Posted in cs.LG · 2026-01-21 · Zhihao Chen, Zirui Gong, Jianting Ning, Yanjun Zhang, Leo Yu Zhang

Beyond Denial-of-Service: The Puppeteer's Attack for Fine-Grained Control in Ranking-Based Federated Learning

Federated Rank Learning (FRL) is a promising Federated Learning (FL) paradigm designed to be resilient against model poisoning attacks due to its discrete, ranking-based update mechanism. Unlike traditional FL methods that rely on model updates, FRL leverages discrete rankings as a communication parameter between clients and the...

💬 0 commentsarXiv:2601.14687v1PDF
0

Posted in cs.AI · 2026-01-21 · Shuai Wang, Yaoming Yang, Bingdong Li, Hao Hao, Aimin Zhou

IB-GRPO: Aligning LLM-based Learning Path Recommendation with Educational Objectives via Indicator-Based Group Relative Policy Optimization

Learning Path Recommendation (LPR) aims to generate personalized sequences of learning items that maximize long-term learning effect while respecting pedagogical principles and operational constraints. Although large language models (LLMs) offer rich semantic understanding for free-form recommendation, applying them to long-horizon...

💬 0 commentsarXiv:2601.14686v1PDF
0

Posted in cs.HC · 2026-01-21 · Zuoyu Zhang, Yancheng Zhu

Enhancing Tool Calling in LLMs with the International Tool Calling Dataset

Tool calling allows large language models (LLMs) to interact with external systems like APIs, enabling applications in customer support, data analysis, and dynamic content generation. While recent benchmarks have advanced tool-use research, they suffer from key limitations, including reliance on simulated or restricted APIs, limited...

💬 0 commentsarXiv:2603.05515v1PDF
0

Posted in cs.SD · 2026-01-21 · Kanami Imamura, Tomohiko Nakamura, Kohei Yatabe, Hiroshi Saruwatari

Dissecting Performance Degradation in Audio Source Separation under Sampling Frequency Mismatch

Audio processing methods based on deep neural networks are typically trained at a single sampling frequency (SF). To handle untrained SFs, signal resampling is commonly employed, but it can degrade performance, particularly when the input SF is lower than the trained SF. This paper investigates the causes of this degradation through...

💬 0 commentsarXiv:2601.14684v1PDF