Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through July 21, 2026 — 09:07:00 EST

0

Posted in cs.AI · 2026-01-11 · Shujian Gao, Yuan Wang, Jiangtao Yan, Zuxuan Wu, Yu-Gang Jiang

Thinking with Deltas: Incentivizing Reinforcement Learning via Differential Visual Reasoning Policy

Reinforcement Learning with Verifiable Rewards (RLVR) has significantly advanced reasoning capabilities in Large Language Models. However, adapting RLVR to multimodal domains suffers from a critical \textit{perception-reasoning decoupling}. Existing paradigms, driven by text-centric outcome rewards, reasoning in language medium,...

💬 0 commentsarXiv:2601.06801v1PDF
0

Posted in cs.CL · 2026-01-11 · Zili Wei, Xiaocui Yang, Yilin Wang, Zihan Wang, Weidong Bao, Shi Feng, Daling Wang, Yifei Zhang

CIRAG: Construction-Integration Retrieval and Adaptive Generation for Multi-hop Question Answering

Triple-based Iterative Retrieval-Augmented Generation (iRAG) mitigates document-level noise for multi-hop question answering. However, existing methods still face limitations: (i) greedy single-path expansion, which propagates early errors and fails to capture parallel evidence from different reasoning branches, and (ii)...

💬 0 commentsarXiv:2601.06799v1PDF
0

Posted in cs.IR · 2026-01-11 · Zhiyang Zhang, Junda She, Kuo Cai, Bo Chen, Shiyao Wang, Xinchen Luo, Qiang Luo, Ruiming Tang, Han Li, Kun Gai, Guorui Zhou

Unleashing the Native Recommendation Potential: LLM-Based Generative Recommendation via Structured Term Identifiers

Leveraging the vast open-world knowledge and understanding capabilities of Large Language Models (LLMs) to develop general-purpose, semantically-aware recommender systems has emerged as a pivotal research direction in generative recommendation. However, existing methods face bottlenecks in constructing item identifiers. Text-based...

💬 0 commentsarXiv:2601.06798v1PDF
0

Posted in cs.AI · 2026-01-11 · Zhengqing Yan, Xinyang Liu, Yi Zhang, Fan Guo, ChengXun Jia, Junchen Wan, Yao Liu, Qi Liu, Jihao Huang, Kang Song

GDEPO: Group Dual-dynamic and Equal-right Advantage Policy Optimization with Enhanced Training Data Utilization for Sample-Constrained Reinforcement Learning

Automated Theorem Proving (ATP) represents a fundamental challenge in Artificial Intelligence (AI), requiring the construction of machine-verifiable proofs in formal languages such as Lean to evaluate AI reasoning capabilities. Reinforcement learning (RL), particularly the high-performance Group Relative Policy Optimization (GRPO)...

💬 0 commentsarXiv:2601.06795v3PDF
0

Posted in cs.AI · 2026-01-11 · Zhicong Li, Lingjie Jiang, Yulan Hu, Xingchen Zeng, Yixia Li, Xiangwen Zhang, Guanhua Chen, Zheng Pan, Xin Li, Yong Liu

No More Stale Feedback: Co-Evolving Critics for Open-World Agent Learning

Critique-guided reinforcement learning (RL) has emerged as a powerful paradigm for training LLM agents by augmenting sparse outcome rewards with natural-language feedback. However, current methods often rely on static or offline critic models, which fail to adapt as the policy evolves. In on-policy RL, the agent's error patterns shift...

💬 0 commentsarXiv:2601.06794v2PDF
0

Posted in cs.CV · 2026-01-11 · Zhongping Ji

CliffordNet: All You Need is Geometric Algebra

Modern computer vision architectures, from CNNs to Transformers, predominantly rely on the stacking of heuristic modules: spatial mixers (Attention/Conv) followed by channel mixers (FFNs). In this work, we challenge this paradigm by returning to mathematical first principles. We propose the Clifford Algebra Network (CAN), also...

💬 0 commentsarXiv:2601.06793v2PDF
0

Posted in cs.LG · 2026-01-11 · Malavika Pradeep, Akshay Sasi, Nusaibah Farrukh, Rahul Venugopal, Elizabeth Sherly

Cross-Modal Computational Model of Brain-Heart Interactions via HRV and EEG Feature

The electroencephalogram (EEG) has been the gold standard for quantifying mental workload; however, due to its complexity and non-portability, it can be constraining. ECG signals, which are feasible on wearable equipment pieces such as headbands, present a promising method for cognitive state monitoring. This research explores whether...

💬 0 commentsarXiv:2601.06792v1PDF
0

Posted in cs.CR · 2026-01-11 · Bowen Shen, Yuyue Chen, Peng Yang, Bin Zhang, Xi Zhang, Zoe L. Jiang

SecMoE: Communication-Efficient Secure MoE Inference via Select-Then-Compute

Privacy-preserving Transformer inference has gained attention due to the potential leakage of private information. Despite recent progress, existing frameworks still fall short of practical model scales, with gaps up to a hundredfold. A possible way to close this gap is the Mixture of Experts (MoE) architecture, which has emerged as a...

💬 0 commentsarXiv:2601.06790v1PDF
0

Posted in cs.SE · 2026-01-11 · Qihao Wang, Ziming Cheng, Shuo Zhang, Fan Liu, Rui Xu, Heng Lian, Kunyi Wang, Xiaoming Yu, Jianghao Yin, Sen Hu, Yue Hu, Shaolei Zhang, Yanbing Liu, Ronghao Chen, Huacan Wang

MemGovern: Enhancing Code Agents through Learning from Governed Human Experiences

While autonomous software engineering (SWE) agents are reshaping programming paradigms, they currently suffer from a "closed-world" limitation: they attempt to fix bugs from scratch or solely using local context, ignoring the immense historical human experience available on platforms like GitHub. Accessing this open-world experience...

💬 0 commentsarXiv:2601.06789v2PDF
0

Posted in cs.LG · 2026-01-11 · Min Chen, Zihan Wang, Canyu Chen, Zeguan Wu, Manling Li, Junyu Liu

Artificial Entanglement in the Fine-Tuning of Large Language Models

Large language models (LLMs) can be adapted to new tasks using parameter-efficient fine-tuning (PEFT) methods that modify only a small number of trainable parameters, often through low-rank updates. In this work, we adopt a quantum-information-inspired perspective to understand their effectiveness. From this perspective, low-rank...

💬 0 commentsarXiv:2601.06788v1PDF
0

Posted in cs.CL · 2026-01-11 · Jaewon Sok, Jewon Yeom, Seonghyeon Park, Jeongjae Park, Taesup Kim

Garbage Attention in Large Language Models: BOS Sink Heads and Sink-aware Pruning

Large Language Models (LLMs) are known to contain significant redundancy, yet a systematic explanation for why certain components, particularly in higher layers, are more redundant has remained elusive. In this work, we identify the BOS sink phenomenon as a key mechanism driving this layer-wise sensitivity. We show that attention...

💬 0 commentsarXiv:2601.06787v1PDF
0

Posted in cs.CL · 2026-01-11 · Jewon Yeom, Jaewon Sok, Seonghyeon Park, Jeongjae Park, Taesup Kim

EpiCaR: Knowing What You Don't Know Matters for Better Reasoning in LLMs

Improving the reasoning abilities of large language models (LLMs) has largely relied on iterative self-training with model-generated data. While effective at boosting accuracy, existing approaches primarily reinforce successful reasoning paths, incurring a substantial calibration cost: models become overconfident and lose the ability...

💬 0 commentsarXiv:2601.06786v1PDF
0

Posted in cs.HC · 2026-01-11 · Huatao Xu, Zihe Liu, Zilin Zeng, Baichuan Li, Mo Li

AutoTour: Automatic Photo Tour Guide with Smartphones and LLMs

We present AutoTour, a system that enhances user exploration by automatically generating fine-grained landmark annotations and descriptive narratives for photos captured by users. The key idea of AutoTour is to fuse visual features extracted from photos with nearby geospatial features queried from open matching databases. Unlike...

💬 0 commentsarXiv:2601.06781v1PDF
0

Posted in cs.CL · 2026-01-11 · Keito Inoshita, Xiaokang Zhou, Akira Kawai

Multi-Stage Evolutionary Model Merging with Meta Data Driven Curriculum Learning for Sentiment-Specialized Large Language Modeling

The emergence of large language models (LLMs) has significantly transformed natural language processing (NLP), enabling more generalized models to perform various tasks with minimal training. However, traditional sentiment analysis methods, which focus on individual tasks such as sentiment classification or aspect-based analysis, are...

💬 0 commentsarXiv:2601.06780v1PDF
0

Posted in cs.CR · 2026-01-11 · Vasanth Iyer, Leonardo Bobadilla, S. S. Iyengar

CyberLLM-FINDS 2025: Instruction-Tuned Fine-tuning of Domain-Specific LLMs with Retrieval-Augmented Generation and Graph Integration for MITRE Evaluation

Large Language Models (LLMs) such as Gemma-2B have shown strong performance in various natural language processing tasks. However, general-purpose models often lack the domain expertise required for cybersecurity applications. This work presents a methodology to fine-tune the Gemma-2B model into a domain-specific cybersecurity LLM. We...

💬 0 commentsarXiv:2601.06779v1PDF
0

Posted in cs.CV · 2026-01-11 · Ali Lotfi, Adam Carter, Mohammad Meysami, Thuan Ha, Kwabena Nketia, Steve Shirtliffe

The Normalized Difference Layer: A Differentiable Spectral Index Formulation for Deep Learning

Normalized difference indices have been a staple in remote sensing for decades. They stay reliable under lighting changes produce bounded values and connect well to biophysical signals. Even so, they are usually treated as a fixed pre processing step with coefficients set to one, which limits how well they can adapt to a specific...

💬 0 commentsarXiv:2601.06777v1PDF
0

Posted in cs.AI · 2026-01-11 · Xufei Tian, Wenli Du, Shaoyi Yang, Han Hu, Hui Xin, Shifeng Qu, Ke Ye

From Text to Simulation: A Multi-Agent LLM Workflow for Automated Chemical Process Design

Process simulation is a critical cornerstone of chemical engineering design. Current automated chemical design methodologies focus mainly on various representations of process flow diagrams. However, transforming these diagrams into executable simulation flowsheets remains a time-consuming and labor-intensive endeavor, requiring...

💬 0 commentsarXiv:2601.06776v1PDF
0

Posted in cs.HC · 2026-01-11 · Xiangzhe Yuan, Jiajun Wang, Huanchen Wang, Qian Wan, Siying Hu

ImmuniFraug: A Metacognitive Intervention Anti-Fraud Approach to Enhance Undergraduate Students' Cyber Fraud Awareness

Cyber fraud now constitutes over half of criminal cases in China, with undergraduate students experiencing a disproportionate rise in victimization. Traditional anti-fraud training remains predominantly passive, yielding limited engagement and retention. This paper introduces ImmuniFraug, a Large Language Model (LLM)-based...

💬 0 commentsarXiv:2601.06774v1PDF
0

Posted in cs.LG · 2026-01-11 · Hao-Xiang Xu, Jun-Yu Ma, Ziqi Peng, Yuhao Sun, Zhen-Hua Ling, Jia-Chen Gu

Multiplicative Orthogonal Sequential Editing for Language Models

Knowledge editing aims to efficiently modify the internal knowledge of large language models (LLMs) without compromising their other capabilities. The prevailing editing paradigm, which appends an update matrix to the original parameter matrix, has been shown by some studies to damage key numerical stability indicators (such as...

💬 0 commentsarXiv:2601.07873v1PDF
0

Posted in cs.SI · 2026-01-11 · Shihui Feng, Baiyue He, Dragan Gasevic, Alec Kirkley

Heterogeneous Interaction Network Analysis (HINA): A New Learning Analytics Approach for Modelling, Analyzing, and Visualizing Complex Interactions in Learning Processes

Existing learning analytics approaches, which often model learning processes as sequences of learner actions or homogeneous relationships, are limited in capturing the distributed, multi-faceted nature of interactions in contemporary learning environments. To address this, we propose Heterogeneous Interaction Network Analysis (HINA),...

💬 0 commentsarXiv:2601.06771v2PDF
0

Posted in cs.LG · 2026-01-11 · Sofiia Huraka, Vakhtang Putkaradze

Structure-preserving learning and prediction in optimal control of collective motion

Wide-spread adoption of unmanned vehicle technologies requires the ability to predict the motion of the combined vehicle operation from observations. While the general prediction of such motion for an arbitrary control mechanism is difficult, for a particular choice of control, the dynamics reduces to the Lie-Poisson equations...

💬 0 commentsarXiv:2601.06770v1PDF
0

Posted in cs.CR · 2026-01-11 · Muhammad Wahid Akram, Keshav Sood, Muneeb Ul Hassan, Dhananjay Thiruvady

ALFA: A Safe-by-Design Approach to Mitigate Quishing Attacks Launched via Fancy QR Codes

Phishing with Quick Response (QR) codes is termed as Quishing. The attackers exploit this method to manipulate individuals into revealing their confidential data. Recently, we see the colorful and fancy representations of QR codes, the 2D matrix of QR codes which does not reflect a typical mixture of black-white modules anymore....

💬 0 commentsarXiv:2601.06768v1PDF
0

Posted in cs.CL · 2026-01-11 · Shubhashis Roy Dipta, Khairul Mahbub, Nadia Najjar

GanitLLM: Difficulty-Aware Bengali Mathematical Reasoning through Curriculum-GRPO

We present a Bengali mathematical reasoning model called GanitLLM (named after the Bangla word for mathematics, Ganit), together with a new difficulty-aware Bengali math corpus and a curriculum-based GRPO pipeline. Bengali is one of the world's most widely spoken languages, yet existing LLMs either reason in English and then...

💬 0 commentsarXiv:2601.06767v3PDF
0

Posted in cs.DB · 2026-01-11 · Jesse Comer, Val Tannen

The Complexity of Finding Missing Answer Repairs

We investigate the problem of identifying database repairs for missing tuples in query answers. We show that when the query is part of the input - the combined complexity setting - determining whether or not a repair exists is polynomial-time is equivalent to the satisfiability problem for classes of queries admitting a weak form of...

💬 0 commentsarXiv:2601.06764v1PDF