Qwen Councils
arXiv could not process that search. Try a simpler keyword search or an arXiv field query such as all:quantum.
Showing downloaded papers while arXiv is unavailable.

Computer Science

arXiv preprints from January 1, 2026 through September 24, 2026 — 10:22:55 EST

0

Posted in cs.SE · 2026-01-09 · Omar Abedelkader, Stéphane Ducasse, Oleksandr Zaitsev, Romain Robbes, Guillermo Polito

Package-Aware Approach for Repository-Level Code Completion in Pharo

Pharo offers a sophisticated completion engine based on semantic heuristics, which coordinates specific fetchers within a lazy architecture. These heuristics can be recomposed to support various activities (e.g., live programming or history usage navigation). While this system is powerful, it does not account for the repository...

💬 0 commentsarXiv:2601.05617v1PDF
0

Posted in cs.LG · 2026-01-09 · ShaoZhen Liu, Xinting Huang, Houwen Peng, Xin Chen, Xinyang Song, Qi Li, Zhenan Sun

Dual-Phase LLM Reasoning: Self-Evolved Mathematical Frameworks

In recent years, large language models (LLMs) have demonstrated significant potential in complex reasoning tasks like mathematical problem-solving. However, existing research predominantly relies on reinforcement learning (RL) frameworks while overlooking supervised fine-tuning (SFT) methods. This paper proposes a new two-stage...

💬 0 commentsarXiv:2601.05616v1PDF
0

Posted in cs.LG · 2026-01-09 · Yiming Zhou, Jiahao Wang, Mingyue Cheng, Hao Wang, Defu Lian, Enhong Chen

PiXTime: A Model for Federated Time Series Forecasting with Heterogeneous Data across Nodes

While collaborative forecasting on distributed time series is highly desirable, directly pooling localized datasets is often impractical due to data sharing constraints. Federated learning offers a promising alternative, yet conventional federated learning algorithms require homogeneous model architectures, which are incompatible with...

💬 0 commentsarXiv:2601.05613v2PDF
0

Posted in cs.CV · 2026-01-09 · Chengen Xie, Chonghao Sima, Tianyu Li, Bin Sun, Junjie Wu, Zhihui Hao, Hongyang Li

FLARE: Learning Future-Aware Latent Representations from Vision-Language Models for Autonomous Driving

While Vision-Language Models (VLMs) offer rich world knowledge for end-to-end autonomous driving, current approaches heavily rely on labor-intensive language annotations (e.g., VQA) to bridge perception and control. This paradigm suffers from a fundamental mismatch between discrete linguistic tokens and continuous driving...

💬 0 commentsarXiv:2601.05611v2PDF
0

Posted in cs.AI · 2026-01-09 · Percy Jardine

CTHA: Constrained Temporal Hierarchical Architecture for Stable Multi-Agent LLM Systems

Recently, multi-time-scale agent architectures have extended the ubiquitous single-loop paradigm by introducing temporal hierarchies with distinct cognitive layers. While yielding substantial performance gains, this diversification fundamentally compromises the coordination stability intrinsic to unified agent systems, which causes...

💬 0 commentsarXiv:2601.10738v1PDF
0

Posted in cs.CL · 2026-01-09 · Nguyen Minh Phuong, Ha-Thanh Nguyen, May Myo Zin, Ken Satoh

Data Augmented Pipeline for Legal Information Extraction and Reasoning

In this paper, we propose a pipeline leveraging Large Language Models (LLMs) for data augmentation in Information Extraction tasks within the legal domain. The proposed method is both simple and effective, significantly reducing the manual effort required for data annotation while enhancing the robustness of Information Extraction...

💬 0 commentsarXiv:2601.05609v1PDF
0

Posted in cs.CV · 2026-01-09 · Miao Pan, Wangjie Gan, Jintao Chen, Wenqi Zhang, Bing Sun, Jianwei Yin, Xuhong Zhang

Ground What You See: Hallucination-Resistant MLLMs via Caption Feedback, Diversity-Aware Sampling, and Conflict Regularization

While Multimodal Large Language Models (MLLMs) have achieved remarkable success across diverse tasks, their practical deployment is severely hindered by hallucination issues, which become particularly acute during Reinforcement Learning (RL) optimization. This paper systematically analyzes the root causes of hallucinations in MLLMs...

💬 0 commentsarXiv:2601.06224v2PDF
0

Posted in cs.LG · 2026-01-09 · Zijun Min, Bingshuai Liu, Ante Wang, Long Zhang, Anxiang Zeng, Haibo Zhang, Jinsong Su

Orchestrating Tokens and Sequences: Dynamic Hybrid Policy Optimization for RLVR

Reinforcement Learning with Verifiable Rewards (RLVR) offers a promising framework for optimizing large language models in reasoning tasks. However, existing RLVR algorithms focus on different granularities, and each has complementary strengths and limitations. Group Relative Policy Optimization (GRPO) updates the policy with...

💬 0 commentsarXiv:2601.05607v1PDF
0

Posted in cs.MA · 2026-01-09 · Chen Han, Jin Tan, Bohan Yu, Wenzhen Zheng, Xijin Tang

Conformity Dynamics in LLM Multi-Agent Systems: The Roles of Topology and Self-Social Weighting

Large Language Models (LLMs) are increasingly instantiated as interacting agents in multi-agent systems (MAS), where collective decisions emerge through social interaction rather than independent reasoning. A fundamental yet underexplored mechanism in this process is conformity, the tendency of agents to align their judgments with...

💬 0 commentsarXiv:2601.05606v1PDF
0

Posted in cs.CV · 2026-01-09 · Zengbin Wang, Junjie Li, Saihui Hou, Xu Liu, Chunshui Cao, Yongzhen Huang, Muyi Sun, Siye Wang, Man Zhang

Learning Geometric Invariance for Gait Recognition

The goal of gait recognition is to extract identity-invariant features of an individual under various gait conditions, e.g., cross-view and cross-clothing. Most gait models strive to implicitly learn the common traits across different gait conditions in a data-driven manner to pull different gait conditions closer for recognition....

💬 0 commentsarXiv:2601.05604v1PDF
0

Posted in cs.IR · 2026-01-09 · Watheq Mansour, J. Shane Culpepper, Joel Mackenzie, Andrew Yates

Revisiting Human-vs-LLM judgments using the TREC Podcast Track

Using large language models (LLMs) to annotate relevance is an increasingly important technique in the information retrieval community. While some studies demonstrate that LLMs can achieve high user agreement with ground truth (human) judgments, other studies have argued for the opposite conclusion. To the best of our knowledge, these...

💬 0 commentsarXiv:2601.05603v2PDF
0

Posted in cs.CV · 2026-01-09 · Chuhan Wang, Xintong Li, Jennifer Yuntong Zhang, Junda Wu, Chengkai Huang, Lina Yao, Julian McAuley, Jingbo Shang

SceneAlign: Aligning Multimodal Reasoning to Scene Graphs in Complex Visual Scenes

Multimodal large language models often struggle with faithful reasoning in complex visual scenes, where intricate entities and relations require precise visual grounding at each step. This reasoning unfaithfulness frequently manifests as hallucinated entities, mis-grounded relations, skipped steps, and over-specified reasoning....

💬 0 commentsarXiv:2601.05600v1PDF
0

Posted in cs.CL · 2026-01-09 · Boxiang Zhao, Qince Li, Zhonghao Wang, Zelin Cao, Yi Wang, Peng Cheng, Bo Lin

Structure-BiEval: A Self-Supervised, Dual-Track Framework for Decoupling Structure and Content in LLM Evaluation for Web Information Systems

As Large Language Models (LLMs) evolve into the core of Web-based autonomous agents and complex Web Information Systems, their ability to faithfully translate natural language into rigorous structured formats has become paramount, as this capability is critical for Web API invocation and data exchange. However, evaluating this...

💬 0 commentsarXiv:2601.19923v2PDF
0

Posted in cs.CV · 2026-01-09 · Takito Sawada, Akinori Iwata, Masahiro Okuda

Quantifying and Inducing Shape Bias in CNNs via Max-Pool Dilation

Convolutional Neural Networks (CNNs) exhibit a well-known texture bias, prioritizing local patterns over global shapes - a tendency inherent to their convolutional architecture. While this bias is beneficial for texture-rich natural images, it often degrades performance on shape-dominant data such as illustrations and sketches....

💬 0 commentsarXiv:2601.05599v2PDF
0

Posted in cs.LG · 2026-01-09 · Sílvia Casacuberta, Moritz Hardt

Good Allocations from Bad Estimates

Conditional average treatment effect (CATE) estimation is the de facto gold standard for targeting a treatment to a heterogeneous population. The method estimates treatment effects up to an error $ε> 0$ in each of $M$ different strata of the population, targeting individuals in decreasing order of estimated treatment effect until the...

💬 0 commentsarXiv:2601.05597v1PDF
0

Posted in cs.CY · 2026-01-09 · Edward C. Cheng, Jeshua Cheng, Alice Siu

Toward Safe and Responsible AI Agents: A Three-Pillar Model for Transparency, Accountability, and Trustworthiness

This paper presents a conceptual and operational framework for developing and operating safe and trustworthy AI agents based on a Three-Pillar Model grounded in transparency, accountability, and trustworthiness. Building on prior work in Human-in-the-Loop systems, reinforcement learning, and collaborative AI, the framework defines an...

💬 0 commentsarXiv:2601.06223v1PDF
0

Posted in cs.CV · 2026-01-09 · Xinghao Wang, Changtao Miao, Dianmo Sheng, Tao Gong, Qi Chu, Nenghai Yu, Quanchen Zou, Deyue Zhang, Xiangzheng Zhang

SAPL: Semantic-Agnostic Prompt Learning in CLIP for Weakly Supervised Image Manipulation Localization

Malicious image manipulation threatens public safety and requires efficient localization methods. Existing approaches depend on costly pixel-level annotations which make training expensive. Existing weakly supervised methods rely only on image-level binary labels and focus on global classification, often overlooking local edge cues...

💬 0 commentsarXiv:2601.06222v1PDF
0

Posted in cs.LG · 2026-01-09 · Jingcheng Hu, Yinmin Zhang, Shijie Shang, Xiaobo Yang, Yue Peng, Zhewei Huang, Hebin Zhou, Xin Wu, Jie Cheng, Fanqi Wan, Xiangwen Kong, Chengyuan Yao, Kaiwen Yan, Ailin Huang, Hongyu Zhou, Qi Han, Zheng Ge, Daxin Jiang, Xiangyu Zhang, Heung-Yeung Shum

PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning

We introduce Parallel Coordinated Reasoning (PaCoRe), a training-and-inference framework designed to overcome a central limitation of contemporary language models: their inability to scale test-time compute (TTC) far beyond sequential reasoning under a fixed context window. PaCoRe departs from the traditional sequential paradigm by...

💬 0 commentsarXiv:2601.05593v1PDF
0

Posted in cs.AI · 2026-01-09 · Haoming Gong, Qingyao Ai, Zhihao Tao, Yongfeng Zhang

A Causal Information-Flow Framework for Unbiased Learning-to-Rank

In web search and recommendation systems, user clicks are widely used to train ranking models. However, click data is heavily biased, i.e., users tend to click higher-ranked items (position bias), choose only what was shown to them (selection bias), and trust top results more (trust bias). Without explicitly modeling these biases, the...

💬 0 commentsarXiv:2601.05590v1PDF
0

Posted in cs.CV · 2026-01-09 · Arnav S. Sonavane

Domain-Specific Self-Supervised Pre-training for Agricultural Disease Classification: A Hierarchical Vision Transformer Study

We investigate the impact of domain-specific self-supervised pre-training on agricultural disease classification using hierarchical vision transformers. Our key finding is that SimCLR pre-training on just 3,000 unlabeled agricultural images provides a +4.57% accuracy improvement--exceeding the +3.70% gain from hierarchical...

💬 0 commentsarXiv:2601.11612v1PDF
0

Posted in cs.GR · 2026-01-09 · Cyprien Plateau Holleville, Bruno Lévy

More Power to the Particles: Analytic Geometry for Partial Optimal Transport-based Fluid simulation

We propose unified data structures and algorithms for free-surface fluid simulations based on partial optimal transport, such as the Power Particles method or Gallouët-Mérigot's scheme. Such methods previously relied on a discretization of the cells by leveraging a classical convex cell clipping algorithm. However, this results in a...

💬 0 commentsarXiv:2601.05765v2PDF
0

Posted in cs.CR · 2026-01-09 · Haris Khan, Sadia Asif, Shumaila Asif

Multi-Agent Framework for Controllable and Protected Generative Content Creation: Addressing Copyright and Provenance in AI-Generated Media

The proliferation of generative AI systems creates unprecedented opportunities for content creation while raising critical concerns about controllability, copyright infringement, and content provenance. Current generative models operate as "black boxes" with limited user control and lack built-in mechanisms to protect intellectual...

💬 0 commentsarXiv:2601.06232v1PDF
0

Posted in cs.LG · 2026-01-09 · Turkan Simge Ispak, Salih Tileylioglu, Erdem Akagunduz

Variational Autoencoders for P-wave Detection on Strong Motion Earthquake Spectrograms

Accurate P-wave detection is critical for earthquake early warning, yet strong-motion records pose challenges due to high noise levels, limited labeled data, and complex waveform characteristics. This study reframes P-wave arrival detection as a self-supervised anomaly detection task to evaluate how architectural variations regulate...

💬 0 commentsarXiv:2601.05759v1PDF
0

Posted in cs.CR · 2026-01-09 · Junda Lin, Zhaomeng Zhou, Zhi Zheng, Shuochen Liu, Tong Xu, Yong Chen, Enhong Chen

VIGIL: Defending LLM Agents Against Tool Stream Injection via Verify-Before-Commit

LLM agents operating in open environments face escalating risks from indirect prompt injection, particularly within the tool stream where manipulated metadata and runtime feedback hijack execution flow. Existing defenses encounter a critical dilemma as advanced models prioritize injected rules due to strict alignment while static...

💬 0 commentsarXiv:2601.05755v2PDF
0

Posted in cs.CL · 2026-01-09 · Shu Yang, Jingyu Hu, Tong Li, Hanqi Yan, Wenxuan Wang, Di Wang

AutoMonitor-Bench: Evaluating the Reliability of LLM-Based Misbehavior Monitor

We introduce AutoMonitor-Bench, the first benchmark designed to systematically evaluate the reliability of LLM-based misbehavior monitors across diverse tasks and failure modes. AutoMonitor-Bench consists of 3,010 carefully annotated test samples spanning question answering, code generation, and reasoning, with paired misbehavior and...

💬 0 commentsarXiv:2601.05752v3PDF