Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through July 20, 2026 — 20:23:18 EST

0

Posted in cs.LG · 2026-01-12 · Yang Zhao, Hepeng Wang, Xiao Ding, Yangou Ouyang, Bibo Cai, Kai Xiong, Jinglong Gao, Zhouhao Sun, Li Du, Bing Qin, Ting Liu

MAESTRO: Meta-learning Adaptive Estimation of Scalarization Trade-offs for Reward Optimization

Group-Relative Policy Optimization (GRPO) has emerged as an efficient paradigm for aligning Large Language Models (LLMs), yet its efficacy is primarily confined to domains with verifiable ground truths. Extending GRPO to open-domain settings remains a critical challenge, as unconstrained generation entails multi-faceted and often...

💬 0 commentsarXiv:2601.07208v2PDF
0

Posted in cs.AI · 2026-01-12 · Hao Li, Yiqun Zhang, Zhaoyan Guo, Chenxu Wang, Shengji Tang, Qiaosheng Zhang, Yang Chen, Biqing Qi, Peng Ye, Lei Bai, Zhen Wang, Shuyue Hu

LLMRouterBench: A Massive Benchmark and Unified Framework for LLM Routing

Large language model (LLM) routing assigns each query to the most suitable model from an ensemble. We introduce LLMRouterBench, a large-scale benchmark and unified framework for LLM routing. It comprises over 400K instances from 21 datasets and 33 models. Moreover, it provides comprehensive metrics for both performance-oriented...

💬 0 commentsarXiv:2601.07206v1PDF
0

Posted in cs.SI · 2026-01-12 · Artem Novobritskii

Intercultural Communication Strategies of a Technology Brand: A Comparative Quantitative Analysis of Xiaomi's Digital Marketing in China and Russia

In the 21st century, the era of globalization, consumers are dispersed across the globe, and brands compete for their attention and loyalty, largely within the digital realm. This reality elevates the importance of effective communication and the transmission of product value across diverse cultural contexts. This study presents a...

💬 0 commentsarXiv:2601.07204v1PDF
0

Posted in cs.CL · 2026-01-12 · Ziao Yang, Zizhang Chen, Lei Zhang, Hongfu Liu

Recontextualizing Famous Quotes for Brand Slogan Generation

Slogans are concise and memorable catchphrases that play a crucial role in advertising by conveying brand identity and shaping public perception. However, advertising fatigue reduces the effectiveness of repeated slogans, creating a growing demand for novel, creative, and insightful slogan generation. While recent work leverages large...

💬 0 commentsarXiv:2602.06049v1PDF
0

Posted in cs.LG · 2026-01-12 · Ibne Farabi Shihab, Sanjeda Akter, Anuj Sharma

CalPro: Prior-Aware Evidential--Conformal Prediction with Structure-Aware Guarantees for Protein Structures

Deep protein structure predictors such as AlphaFold provide confidence estimates (e.g., pLDDT) that are often miscalibrated and degrade under distribution shifts across experimental modalities, temporal changes, and intrinsically disordered regions. We introduce CalPro, a prior-aware evidential-conformal framework for shift-robust...

💬 0 commentsarXiv:2601.07201v1PDF
0

Posted in cs.LG · 2026-01-12 · Haozhong Wang, Zhuo Li, Yibo Yang, He Zhao, Hongyuan Zha, Dandan Guo

Safeguarding LLM Fine-tuning via Push-Pull Distributional Alignment

The inherent safety alignment of Large Language Models (LLMs) is prone to erosion during fine-tuning, even when using seemingly innocuous datasets. While existing defenses attempt to mitigate this via data selection, they typically rely on heuristic, instance-level assessments that neglect the global geometry of the data distribution...

💬 0 commentsarXiv:2601.07200v1PDF
0

Posted in cs.LG · 2026-01-12 · Murtaza Nikzad, Raghuram Ramanujan

Forward versus Backward: Comparing Reasoning Objectives in Direct Preference Optimization

Large language models exhibit impressive reasoning capabilities yet frequently generate plausible but incorrect solutions, a phenomenon commonly termed hallucination. This paper investigates the effect of training objective composition on reasoning reliability through Direct Preference Optimization. Two complementary training signals...

💬 0 commentsarXiv:2601.07199v1PDF
0

Posted in cs.LG · 2026-01-12 · Ibne Farabi Shihab, Sanjeda Akter, Anuj Sharma

Beyond Variance: Knowledge-Aware LLM Compression via Fisher-Aligned Subspace Diagnostics

Post-training activation compression is essential for deploying Large Language Models (LLMs) on resource-constrained hardware. However, standard methods like Singular Value Decomposition (SVD) are gradient-blind: they preserve high-variance dimensions regardless of their impact on factual knowledge preservation. We introduce...

💬 0 commentsarXiv:2601.07197v1PDF
0

Posted in cs.CL · 2026-01-12 · Manzong Huang, Chenyang Bu, Yi He, Xingrui Zhuo, Xindong Wu

Relink: Constructing Query-Driven Evidence Graph On-the-Fly for GraphRAG

Graph-based Retrieval-Augmented Generation (GraphRAG) mitigates hallucinations in Large Language Models (LLMs) by grounding them in structured knowledge. However, current GraphRAG methods are constrained by a prevailing \textit{build-then-reason} paradigm, which relies on a static, pre-constructed Knowledge Graph (KG). This paradigm...

💬 0 commentsarXiv:2601.07192v1PDF
0

Posted in cs.AI · 2026-01-12 · Nikhil Verma

Active Context Compression: Autonomous Memory Management in LLM Agents

Large Language Model (LLM) agents struggle with long-horizon software engineering tasks due to "Context Bloat." As interaction history grows, computational costs explode, latency increases, and reasoning capabilities degrade due to distraction by irrelevant past errors. Existing solutions often rely on passive, external summarization...

💬 0 commentsarXiv:2601.07190v1PDF
0

Posted in cs.CV · 2026-01-12 · Hema Hariharan Samson

ForensicFormer: Hierarchical Multi-Scale Reasoning for Cross-Domain Image Forgery Detection

The proliferation of AI-generated imagery and sophisticated editing tools has rendered traditional forensic methods ineffective for cross-domain forgery detection. We present ForensicFormer, a hierarchical multi-scale framework that unifies low-level artifact detection, mid-level boundary analysis, and high-level semantic reasoning...

💬 0 commentsarXiv:2601.08873v1PDF
0

Posted in cs.LG · 2026-01-12 · Susana Lopez-Moreno, Eric Dolores-Cuenca, Sangil Kim

Standardization of Post-Publication Code Verification by Journals is Possible with the Support of the Community

Reproducibility remains a challenge in machine learning research. While code and data availability requirements have become increasingly common, post-publication verification in journals is still limited and unformalized. This position paper argues that it is plausible for journals and conference proceedings to implement...

💬 0 commentsarXiv:2601.07189v1PDF
0

Posted in cs.RO · 2026-01-12 · Zainab Altaweel, Mohaiminul Al Nahian, Jake Juettner, Adnan Siraj Rakin, Shiqi Zhang

PROTEA: Securing Robot Task Planning and Execution

Robots need task planning methods to generate action sequences for complex tasks. Recent work on adversarial attacks has revealed significant vulnerabilities in existing robot task planners, especially those built on foundation models. In this paper, we aim to address these security challenges by introducing PROTEA, an LLM-as-a-Judge...

💬 0 commentsarXiv:2601.07186v1PDF
0

Posted in cs.CR · 2026-01-12 · Shawn Li, Chenxiao Yu, Zhiyu Ni, Hao Li, Charith Peris, Chaowei Xiao, Yue Zhao

Defenses Against Prompt Attacks Learn Surface Heuristics

Large language models (LLMs) are increasingly deployed in security-sensitive applications, where they must follow system- or developer-specified instructions that define the intended task behavior, while completing benign user requests. When adversarial instructions appear in user queries or externally retrieved content, models may...

💬 0 commentsarXiv:2601.07185v1PDF
0

Posted in cs.CV · 2026-01-12 · Shezheng Song, Shasha Li, Jie Yu

Seeing Right but Saying Wrong: Inter- and Intra-Layer Refinement in MLLMs without Training

Multimodal Large Language Models (MLLMs) have demonstrated strong capabilities across a variety of vision-language tasks. However, their internal reasoning often exhibits a critical inconsistency: although deeper layers may attend to the correct visual regions, final predictions are frequently misled by noisy attention from earlier...

💬 0 commentsarXiv:2601.07359v1PDF
0

Posted in cs.IT · 2026-01-12 · Yichen Fu, Tianming Wang, Ke Wei

Fast and Provable Nonconvex Robust Matrix Completion

We study the robust matrix completion (RMC) problem subject to both sparse outliers and stochastic noise. A non-convex method termed Accelerated Robust Matrix Completion (ARMC) is proposed, which accelerates a prior non-convex approach by incorporating an explicit subspace projection step into the low-rank update, leading to...

💬 0 commentsarXiv:2601.07355v2PDF
0

Posted in cs.CL · 2026-01-12 · Tianyu Liu, Qitan Lv, Yuhao Shen, Xiao Sun, Xiaoyan Sun

TALON: Confidence-Aware Speculative Decoding with Adaptive Token Trees

Speculative decoding (SD) has become a standard technique for accelerating LLM inference without sacrificing output quality. Recent advances in speculative decoding have shifted from sequential chain-based drafting to tree-structured generation, where the draft model constructs a tree of candidate tokens to explore multiple possible...

💬 0 commentsarXiv:2601.07353v1PDF
0

Posted in cs.CL · 2026-01-12 · Linhao Zhong, Linyu Wu, Bozhen Fang, Tianjian Feng, Chenchen Jing, Wen Wang, Jiaheng Zhang, Hao Chen, Chunhua Shen

Beyond Hard Masks: Progressive Token Evolution for Diffusion Language Models

Diffusion Language Models (DLMs) offer a promising alternative for language modeling by enabling parallel decoding through iterative refinement. However, most DLMs rely on hard binary masking and discrete token assignments, which hinder the revision of early decisions and underutilize intermediate probabilistic representations. In...

💬 0 commentsarXiv:2601.07351v2PDF
0

Posted in cs.CL · 2026-01-12 · Zongqi Wang, Rui Wang, Yuchuan Wu, Yiyao Yu, Pinyi Zhang, Shaoning Sun, Yujiu Yang, Yongbin Li

Reward Modeling from Natural Language Human Feedback

Reinforcement Learning with Verifiable reward (RLVR) on preference data has become the mainstream approach for training Generative Reward Models (GRMs). Typically in pairwise rewarding tasks, GRMs generate reasoning chains ending with critiques and preference labels, and RLVR then relies on the correctness of the preference labels as...

💬 0 commentsarXiv:2601.07349v3PDF
0

Posted in cs.CL · 2026-01-12 · Tu Hu, Ronghao Chen, Shuo Zhang, Jianghao Yin, Mou Xiao Feng, Jingping Liu, Shaolei Zhang, Wenqi Jiang, Yuqi Fang, Sen Hu, Huacan Wang, Yi Xu

Controlled Self-Evolution for Algorithmic Code Optimization

Self-evolution methods enhance code generation through iterative "generate-verify-refine" cycles, yet existing approaches suffer from low exploration efficiency, failing to discover solutions with superior complexity within limited budgets. This inefficiency stems from initialization bias trapping evolution in poor solution regions,...

💬 0 commentsarXiv:2601.07348v5PDF
0

Posted in cs.CL · 2026-01-12 · Shaokai He, Kaiwen Wei, Xinyi Zeng, Xiang Chen, Xue Yang, Zhenyang Li, Jiang Zhong, Yu Tian

DiffER: Diffusion Entity-Relation Modeling for Reversal Curse in Diffusion Large Language Models

The "reversal curse" refers to the phenomenon where large language models (LLMs) exhibit predominantly unidirectional behavior when processing logically bidirectional relationships. Prior work attributed this to autoregressive training -- predicting the next token inherently favors left-to-right information flow over genuine...

💬 0 commentsarXiv:2601.07347v1PDF
0

Posted in cs.RO · 2026-01-12 · Mengyun Liu, Shanshan Huang, Jianan Jiang

EdgeNav-QE: QLoRA Quantization and Dynamic Early Exit for LAM-based Navigation on Edge Devices

Large Action Models (LAMs) have shown immense potential in autonomous navigation by bridging high-level reasoning with low-level control. However, deploying these multi-billion parameter models on edge devices remains a significant challenge due to memory constraints and latency requirements. In this paper, we propose EdgeNav-QE, a...

💬 0 commentsarXiv:2602.15836v1PDF
0

Posted in cs.CV · 2026-01-12 · Jiao Xu, Junwei Liu, Jiangwei Lao, Qi Zhu, Yunpeng Zhao, Congyun Jin, Shinan Liu, Zhihong Lu, Lihe Zhang, Xin Chen, Jian Wang, Ping Wang

PulseMind: A Multi-Modal Medical Model for Real-World Clinical Diagnosis

Recent advances in medical multi-modal models focus on specialized image analysis like dermatology, pathology, or radiology. However, they do not fully capture the complexity of real-world clinical diagnostics, which involve heterogeneous inputs and require ongoing contextual understanding during patient-physician interactions. To...

💬 0 commentsarXiv:2601.07344v1PDF
0

Posted in cs.AI · 2026-01-12 · Nicolas Tacheny

Agentic Diagnostic Reasoning over Telecom and Datacenter Infrastructure

Large-scale telecom and datacenter infrastructures rely on multi-layered service and resource models, where failures propagate across physical and logical components and affect multiple customers. Traditional approaches to root cause analysis(RCA) rely on hard-coded graph traversal algorithms or rule-based correlation engines, which...

💬 0 commentsarXiv:2601.07342v1PDF