Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through July 28, 2026 — 05:27:38 EST

0

Posted in cs.IT · 2026-01-09 · Avraham Kreindel, Isaac Barouch Essayag, Aryeh Lev Zabokritskiy

Multiset Deletion Codes: Cyclic Constructions, Bounds, and Exact Results

We study deletion-correcting codes in the space of length-$n$ multisets over a $q$-ary alphabet. We present an explicit cyclic Sidon-type construction for arbitrary alphabet size $q$ and deletion radius $t$, defined by a single congruence modulo $t(t+1)^{q-2}+1$. The construction has redundancy at most $\log_q(t(t+1)^{q-2}+1)$ and...

💬 0 commentsarXiv:2601.05636v2PDF
0

Posted in cs.CR · 2026-01-09 · Honghao Liu, Xuhui Jiang, Chengjin Xu, Cehao Yang, Yiran Cheng, Lionel Ni, Jian Guo

Continual Pretraining on Encrypted Synthetic Data for Privacy-Preserving LLMs

Preserving privacy in sensitive data while pretraining large language models on small, domain-specific corpora presents a significant challenge. In this work, we take an exploratory step toward privacy-preserving continual pretraining by proposing an entity-based framework that synthesizes encrypted training data to protect personally...

💬 0 commentsarXiv:2601.05635v2PDF
0

Posted in cs.CL · 2026-01-09 · Nuoyan Lyu, Bingbing Xu, Xueyun Tian, Weihao Meng, Yige Yuan, Yang Zhang, Zhiyong Huang, Tat-Seng Chua, Huawei Shen

GIFT: Games as Informal Training for Generalizable LLMs

Recent LLMs excel at formal tasks such as mathematical reasoning and code generation, but still struggle with broader abilities such as planning, creativity, and social intelligence. Inspired by human learning, where formal instruction and informal experience jointly shape intelligence, we introduce informal learning into LLM training...

💬 0 commentsarXiv:2601.05633v2PDF
0

Posted in cs.AI · 2026-01-09 · Jiapu Wang, Xinghe Cheng, Zezheng Wu, Ruiqi Ma, Rui Wang, Zhichao Yan, Haoran Luo, Yuhao Jiang, Kai Sun

Cumulative Path-Level Semantic Reasoning for Inductive Knowledge Graph Completion

Conventional Knowledge Graph Completion (KGC) methods aim to infer missing information in incomplete Knowledge Graphs (KGs) by leveraging existing information, which struggle to perform effectively in scenarios involving emerging entities. Inductive KGC methods can handle the emerging entities and relations in KGs, offering greater...

💬 0 commentsarXiv:2601.05629v1PDF
0

Posted in cs.CL · 2026-01-09 · Abayomi O. Agbeyangi

Text Detoxification in isiXhosa and Yorùbá: A Cross-Lingual Machine Learning Approach for Low-Resource African Languages

Toxic language is one of the major barrier to safe online participation, yet robust mitigation tools are scarce for African languages. This study addresses this critical gap by investigating automatic text detoxification (toxic to neutral rewriting) for two low-resource African languages, isiXhosa and Yorùbá. The work contributes a...

💬 0 commentsarXiv:2601.05624v1PDF
0

Posted in cs.LG · 2026-01-09 · Zhi Wang, Zhongbin Wu, Yanni Li, Bing Liu, Guangxi Li, Yuping Wang

Continual Learning of Achieving Forgetting-free and Positive Knowledge Transfer

Existing research on continual learning (CL) of a sequence of tasks focuses mainly on dealing with catastrophic forgetting (CF) to balance the learning plasticity of new tasks and the memory stability of old tasks. However, an ideal CL agent should not only be able to overcome CF, but also encourage positive forward and backward...

💬 0 commentsarXiv:2601.05623v1PDF
0

Posted in cs.SE · 2026-01-09 · Chengjie Wang, Jingzheng Wu, Hao Lyu, Xiang Ling, Tianyue Luo, Yanjun Wu, Chen Zhao

A Large Scale Empirical Analysis on the Adherence Gap between Standards and Tools in SBOM

A Software Bill of Materials (SBOM) is a machine-readable artifact that systematically organizes software information, enhancing supply chain transparency and security. To facilitate the exchange and utilization of SBOMs, organizations such as the Linux Foundation and OWASP have proposed SBOM standards. Following standards,...

💬 0 commentsarXiv:2601.05622v1PDF
0

Posted in cs.SE · 2026-01-09 · Omar Abedelkader, Stéphane Ducasse, Oleksandr Zaitsev, Romain Robbes, Guillermo Polito

Package-Aware Approach for Repository-Level Code Completion in Pharo

Pharo offers a sophisticated completion engine based on semantic heuristics, which coordinates specific fetchers within a lazy architecture. These heuristics can be recomposed to support various activities (e.g., live programming or history usage navigation). While this system is powerful, it does not account for the repository...

💬 0 commentsarXiv:2601.05617v1PDF
0

Posted in cs.LG · 2026-01-09 · ShaoZhen Liu, Xinting Huang, Houwen Peng, Xin Chen, Xinyang Song, Qi Li, Zhenan Sun

Dual-Phase LLM Reasoning: Self-Evolved Mathematical Frameworks

In recent years, large language models (LLMs) have demonstrated significant potential in complex reasoning tasks like mathematical problem-solving. However, existing research predominantly relies on reinforcement learning (RL) frameworks while overlooking supervised fine-tuning (SFT) methods. This paper proposes a new two-stage...

💬 0 commentsarXiv:2601.05616v1PDF
0

Posted in cs.LG · 2026-01-09 · Yiming Zhou, Jiahao Wang, Mingyue Cheng, Hao Wang, Defu Lian, Enhong Chen

PiXTime: A Model for Federated Time Series Forecasting with Heterogeneous Data across Nodes

While collaborative forecasting on distributed time series is highly desirable, directly pooling localized datasets is often impractical due to data sharing constraints. Federated learning offers a promising alternative, yet conventional federated learning algorithms require homogeneous model architectures, which are incompatible with...

💬 0 commentsarXiv:2601.05613v2PDF
0

Posted in cs.CV · 2026-01-09 · Chengen Xie, Chonghao Sima, Tianyu Li, Bin Sun, Junjie Wu, Zhihui Hao, Hongyang Li

FLARE: Learning Future-Aware Latent Representations from Vision-Language Models for Autonomous Driving

While Vision-Language Models (VLMs) offer rich world knowledge for end-to-end autonomous driving, current approaches heavily rely on labor-intensive language annotations (e.g., VQA) to bridge perception and control. This paradigm suffers from a fundamental mismatch between discrete linguistic tokens and continuous driving...

💬 0 commentsarXiv:2601.05611v2PDF
0

Posted in cs.AI · 2026-01-09 · Percy Jardine

CTHA: Constrained Temporal Hierarchical Architecture for Stable Multi-Agent LLM Systems

Recently, multi-time-scale agent architectures have extended the ubiquitous single-loop paradigm by introducing temporal hierarchies with distinct cognitive layers. While yielding substantial performance gains, this diversification fundamentally compromises the coordination stability intrinsic to unified agent systems, which causes...

💬 0 commentsarXiv:2601.10738v1PDF
0

Posted in cs.CL · 2026-01-09 · Nguyen Minh Phuong, Ha-Thanh Nguyen, May Myo Zin, Ken Satoh

Data Augmented Pipeline for Legal Information Extraction and Reasoning

In this paper, we propose a pipeline leveraging Large Language Models (LLMs) for data augmentation in Information Extraction tasks within the legal domain. The proposed method is both simple and effective, significantly reducing the manual effort required for data annotation while enhancing the robustness of Information Extraction...

💬 0 commentsarXiv:2601.05609v1PDF
0

Posted in cs.CV · 2026-01-09 · Miao Pan, Wangjie Gan, Jintao Chen, Wenqi Zhang, Bing Sun, Jianwei Yin, Xuhong Zhang

Ground What You See: Hallucination-Resistant MLLMs via Caption Feedback, Diversity-Aware Sampling, and Conflict Regularization

While Multimodal Large Language Models (MLLMs) have achieved remarkable success across diverse tasks, their practical deployment is severely hindered by hallucination issues, which become particularly acute during Reinforcement Learning (RL) optimization. This paper systematically analyzes the root causes of hallucinations in MLLMs...

💬 0 commentsarXiv:2601.06224v2PDF
0

Posted in cs.LG · 2026-01-09 · Zijun Min, Bingshuai Liu, Ante Wang, Long Zhang, Anxiang Zeng, Haibo Zhang, Jinsong Su

Orchestrating Tokens and Sequences: Dynamic Hybrid Policy Optimization for RLVR

Reinforcement Learning with Verifiable Rewards (RLVR) offers a promising framework for optimizing large language models in reasoning tasks. However, existing RLVR algorithms focus on different granularities, and each has complementary strengths and limitations. Group Relative Policy Optimization (GRPO) updates the policy with...

💬 0 commentsarXiv:2601.05607v1PDF
0

Posted in cs.MA · 2026-01-09 · Chen Han, Jin Tan, Bohan Yu, Wenzhen Zheng, Xijin Tang

Conformity Dynamics in LLM Multi-Agent Systems: The Roles of Topology and Self-Social Weighting

Large Language Models (LLMs) are increasingly instantiated as interacting agents in multi-agent systems (MAS), where collective decisions emerge through social interaction rather than independent reasoning. A fundamental yet underexplored mechanism in this process is conformity, the tendency of agents to align their judgments with...

💬 0 commentsarXiv:2601.05606v1PDF
0

Posted in cs.CV · 2026-01-09 · Zengbin Wang, Junjie Li, Saihui Hou, Xu Liu, Chunshui Cao, Yongzhen Huang, Muyi Sun, Siye Wang, Man Zhang

Learning Geometric Invariance for Gait Recognition

The goal of gait recognition is to extract identity-invariant features of an individual under various gait conditions, e.g., cross-view and cross-clothing. Most gait models strive to implicitly learn the common traits across different gait conditions in a data-driven manner to pull different gait conditions closer for recognition....

💬 0 commentsarXiv:2601.05604v1PDF
0

Posted in cs.IR · 2026-01-09 · Watheq Mansour, J. Shane Culpepper, Joel Mackenzie, Andrew Yates

Revisiting Human-vs-LLM judgments using the TREC Podcast Track

Using large language models (LLMs) to annotate relevance is an increasingly important technique in the information retrieval community. While some studies demonstrate that LLMs can achieve high user agreement with ground truth (human) judgments, other studies have argued for the opposite conclusion. To the best of our knowledge, these...

💬 0 commentsarXiv:2601.05603v2PDF
0

Posted in cs.CV · 2026-01-09 · Chuhan Wang, Xintong Li, Jennifer Yuntong Zhang, Junda Wu, Chengkai Huang, Lina Yao, Julian McAuley, Jingbo Shang

SceneAlign: Aligning Multimodal Reasoning to Scene Graphs in Complex Visual Scenes

Multimodal large language models often struggle with faithful reasoning in complex visual scenes, where intricate entities and relations require precise visual grounding at each step. This reasoning unfaithfulness frequently manifests as hallucinated entities, mis-grounded relations, skipped steps, and over-specified reasoning....

💬 0 commentsarXiv:2601.05600v1PDF
0

Posted in cs.CL · 2026-01-09 · Boxiang Zhao, Qince Li, Zhonghao Wang, Zelin Cao, Yi Wang, Peng Cheng, Bo Lin

Structure-BiEval: A Self-Supervised, Dual-Track Framework for Decoupling Structure and Content in LLM Evaluation for Web Information Systems

As Large Language Models (LLMs) evolve into the core of Web-based autonomous agents and complex Web Information Systems, their ability to faithfully translate natural language into rigorous structured formats has become paramount, as this capability is critical for Web API invocation and data exchange. However, evaluating this...

💬 0 commentsarXiv:2601.19923v2PDF
0

Posted in cs.CV · 2026-01-09 · Takito Sawada, Akinori Iwata, Masahiro Okuda

Quantifying and Inducing Shape Bias in CNNs via Max-Pool Dilation

Convolutional Neural Networks (CNNs) exhibit a well-known texture bias, prioritizing local patterns over global shapes - a tendency inherent to their convolutional architecture. While this bias is beneficial for texture-rich natural images, it often degrades performance on shape-dominant data such as illustrations and sketches....

💬 0 commentsarXiv:2601.05599v2PDF
0

Posted in cs.LG · 2026-01-09 · Sílvia Casacuberta, Moritz Hardt

Good Allocations from Bad Estimates

Conditional average treatment effect (CATE) estimation is the de facto gold standard for targeting a treatment to a heterogeneous population. The method estimates treatment effects up to an error $ε> 0$ in each of $M$ different strata of the population, targeting individuals in decreasing order of estimated treatment effect until the...

💬 0 commentsarXiv:2601.05597v1PDF
0

Posted in cs.CY · 2026-01-09 · Edward C. Cheng, Jeshua Cheng, Alice Siu

Toward Safe and Responsible AI Agents: A Three-Pillar Model for Transparency, Accountability, and Trustworthiness

This paper presents a conceptual and operational framework for developing and operating safe and trustworthy AI agents based on a Three-Pillar Model grounded in transparency, accountability, and trustworthiness. Building on prior work in Human-in-the-Loop systems, reinforcement learning, and collaborative AI, the framework defines an...

💬 0 commentsarXiv:2601.06223v1PDF
0

Posted in cs.CV · 2026-01-09 · Xinghao Wang, Changtao Miao, Dianmo Sheng, Tao Gong, Qi Chu, Nenghai Yu, Quanchen Zou, Deyue Zhang, Xiangzheng Zhang

SAPL: Semantic-Agnostic Prompt Learning in CLIP for Weakly Supervised Image Manipulation Localization

Malicious image manipulation threatens public safety and requires efficient localization methods. Existing approaches depend on costly pixel-level annotations which make training expensive. Existing weakly supervised methods rely only on image-level binary labels and focus on global classification, often overlooking local edge cues...

💬 0 commentsarXiv:2601.06222v1PDF
0

Posted in cs.LG · 2026-01-09 · Jingcheng Hu, Yinmin Zhang, Shijie Shang, Xiaobo Yang, Yue Peng, Zhewei Huang, Hebin Zhou, Xin Wu, Jie Cheng, Fanqi Wan, Xiangwen Kong, Chengyuan Yao, Kaiwen Yan, Ailin Huang, Hongyu Zhou, Qi Han, Zheng Ge, Daxin Jiang, Xiangyu Zhang, Heung-Yeung Shum

PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning

We introduce Parallel Coordinated Reasoning (PaCoRe), a training-and-inference framework designed to overcome a central limitation of contemporary language models: their inability to scale test-time compute (TTC) far beyond sequential reasoning under a fixed context window. PaCoRe departs from the traditional sequential paradigm by...

💬 0 commentsarXiv:2601.05593v1PDF