Qwen Councils
arXiv is taking too long to respond. Please try again or narrow your search.
Showing downloaded papers while arXiv is unavailable.

Computer Science

arXiv preprints from January 1, 2026 through July 28, 2026 — 17:30:55 EST

0

Posted in cs.CL · 2026-01-07 · Pingjun Hong, Benjamin Roth

Do LLM Self-Explanations Help Users Predict Model Behavior? Evaluating Counterfactual Simulatability with Pragmatic Perturbations

Large Language Models (LLMs) can produce verbalized self-explanations, yet prior studies suggest that such rationales may not reliably reflect the model's true decision process. We ask whether these explanations nevertheless help users predict model behavior, operationalized as counterfactual simulatability. Using StrategyQA, we...

💬 0 commentsarXiv:2601.03775v1PDF
0

Posted in cs.AI · 2026-01-07 · Zihang Li, Yuhang Wang, Yikun Zong, Wenhan Yu, Xiaokun Yuan, Runhan Jiang, Zirui Liu, Tong Yang, Arthur Jiang

EntroCoT: Enhancing Chain-of-Thought via Adaptive Entropy-Guided Segmentation

Chain-of-Thought (CoT) prompting has significantly enhanced the mathematical reasoning capabilities of Large Language Models. We find existing fine-tuning datasets frequently suffer from the "answer right but reasoning wrong" probelm, where correct final answers are derived from hallucinated, redundant, or logically invalid...

💬 0 commentsarXiv:2601.03769v3PDF
0

Posted in cs.PL · 2026-01-07 · Yichen Xu, Martin Odersky

Agentic Proof Automation: A Case Study

Proof engineering is notoriously labor-intensive: proofs that are straightforward on paper often require lengthy scripts in theorem provers. Recent advances in large language models (LLMs) create new opportunities for proof automation: modern LLMs not only generate proof scripts, but also support agentic behavior, exploring codebases...

💬 0 commentsarXiv:2601.03768v1PDF
0

Posted in cs.LG · 2026-01-07 · Noam Levi

Learning Shrinks the Hard Tail: Training-Dependent Inference Scaling in a Solvable Linear Model

We analyze neural scaling laws in a solvable model of last-layer fine-tuning where targets have intrinsic, instance-heterogeneous difficulty. In our Latent Instance Difficulty (LID) model, each input's target variance is governed by a latent ``precision'' drawn from a heavy-tailed distribution. While generalization loss recovers...

💬 0 commentsarXiv:2601.03764v1PDF
0

Posted in cs.NI · 2026-01-07 · Nguyen Cong Luong, Zeping Sui, Duc Van Le, Jie Cao, Bo Ma, Nguyen Duc Hai, Ruichen Zhang, Vu Van Quang, Dusit Niyato, Shaohan Feng

Incentive Mechanism Design for Resource Management in Satellite Networks: A Comprehensive Survey

Resource management is one of the challenges in satellite networks due to their high mobility, wide coverage, long propagation distances, and stringent constraints on energy, communication, and computation resources. Traditional resource allocation approaches rely only on hard and rigid system performance metrics. Meanwhile, incentive...

💬 0 commentsarXiv:2601.03757v1PDF
0

Posted in cs.LG · 2026-01-07 · Paulius Rauba, Viktor Cikojevic, Fran Bartolic, Sam Levang, Ty Dickinson, Chase Dwelle

Probabilistic Transformers for Joint Modeling of Global Weather Dynamics and Decision-Centric Variables

Weather forecasts sit upstream of high-stakes decisions in domains such as grid operations, aviation, agriculture, and emergency response. Yet forecast users often face a difficult trade-off. Many decision-relevant targets are functionals of the atmospheric state variables, such as extrema, accumulations, and threshold exceedances,...

💬 0 commentsarXiv:2601.03753v1PDF
0

Posted in cs.CL · 2026-01-07 · Dominik Macko

Evaluation of Multilingual LLMs Personalized Text Generation Capabilities Targeting Groups and Social-Media Platforms

Capabilities of large language models to generate multilingual coherent text have continuously enhanced in recent years, which opens concerns about their potential misuse. Previous research has shown that they can be misused for generation of personalized disinformation in multiple languages. It has also been observed that...

💬 0 commentsarXiv:2601.03752v1PDF
0

Posted in cs.IR · 2026-01-07 · Dario Maio, Stefano Rizzi

Bridging OLAP and RAG: A Multidimensional Approach to the Design of Corpus Partitioning

Retrieval-Augmented Generation (RAG) systems are increasingly deployed on large-scale document collections, often comprising millions of documents and tens of millions of text chunks. In industrial-scale retrieval platforms, scalability is typically addressed through horizontal sharding and a combination of Approximate...

💬 0 commentsarXiv:2601.03748v1PDF
0

Posted in cs.CL · 2026-01-07 · Jakob Schuster, Vagrant Gautam, Katja Markert

Whose Facts Win? LLM Source Preferences under Knowledge Conflicts

As large language models (LLMs) are more frequently used in retrieval-augmented generation pipelines, it is increasingly relevant to study their behavior under knowledge conflicts. Thus far, the role of the source of the retrieved information has gone unexamined. We address this gap with a novel framework to investigate how source...

💬 0 commentsarXiv:2601.03746v3PDF
0

Posted in cs.CL · 2026-01-07 · Yi Yao, He Zhu, Piaohong Wang, Jincheng Ren, Xinlong Yang, Qianben Chen, Xiaowan Li, Dingfeng Shi, Jiaxian Li, Qiexiang Wang, Sinuo Wang, Xinpeng Liu, Jiaqi Wu, Minghao Liu, Wangchunshu Zhou

O-Researcher: An Open Ended Deep Research Model via Multi-Agent Distillation and Agentic RL

The performance gap between closed-source and open-source large language models (LLMs) is largely attributed to disparities in access to high-quality training data. To bridge this gap, we introduce a novel framework for the automated synthesis of sophisticated, research-grade instructional data. Our approach centers on a multi-agent...

💬 0 commentsarXiv:2601.03743v1PDF
0

Posted in cs.CV · 2026-01-07 · Jinghan Yu, Junhao Xiao, Chenyu Zhu, Jiaming Li, Jia Li, HanMing Deng, Xirui Wang, Guoli Jia, Jianjun Li, Xiang Bai, Bowen Zhou, Zhiyuan Ma

I2E: From Image Pixels to Actionable Interactive Environments for Text-Guided Image Editing

Existing text-guided image editing methods primarily rely on end-to-end pixel-level inpainting paradigm. Despite its success in simple scenarios, this paradigm still significantly struggles with compositional editing tasks that require precise local control and complex multi-object spatial reasoning. This paradigm is severely limited...

💬 0 commentsarXiv:2601.03741v2PDF
0

Posted in cs.CV · 2026-01-07 · Shuyan Bai, Tingfa Xu, Peifu Liu, Yuhao Qiu, Huiyan Bai, Huan Chen, Yanyan Peng, Jianan Li

HyperCOD: The First Challenging Benchmark and Baseline for Hyperspectral Camouflaged Object Detection

RGB-based camouflaged object detection struggles in real-world scenarios where color and texture cues are ambiguous. While hyperspectral image offers a powerful alternative by capturing fine-grained spectral signatures, progress in hyperspectral camouflaged object detection (HCOD) has been critically hampered by the absence of a...

💬 0 commentsarXiv:2601.03736v1PDF
0

Posted in cs.CV · 2026-01-07 · Xiaoxian Shen, Yuhui Zhang, Sahithi Ankireddy, Xiaohan Wang, Maya Varma, Henry Guo, Curtis Langlotz, Serena Yeung-Levy

RadDiff: Describing Differences in Radiology Image Sets with Natural Language

Understanding how two radiology image sets differ is critical for generating clinical insights and for interpreting medical AI systems. We introduce RadDiff, a multimodal agentic system that performs radiologist-style comparative reasoning to describe clinically meaningful differences between paired radiology studies. RadDiff builds...

💬 0 commentsarXiv:2601.03733v1PDF
0

Posted in cs.SE · 2026-01-07 · Jia Li, Yuxin Su, Michael R. Lyu

From Laboratory to Real-World Applications: Benchmarking Agentic Code Reasoning at the Repository Level

As large language models (LLMs) evolve into autonomous agents, evaluating repository-level reasoning, the ability to maintain logical consistency across massive, real-world, interdependent file systems, has become critical. Current benchmarks typically fluctuate between isolated code snippets and black-box evaluations. We present...

💬 0 commentsarXiv:2601.03731v3PDF
0

Posted in cs.IR · 2026-01-07 · Fabian Haak, Philipp Schaer

Perception-Aware Bias Detection for Query Suggestions

Bias in web search has been in the spotlight of bias detection research for quite a while. At the same time, little attention has been paid to query suggestions in this regard. Awareness of the problem of biased query suggestions has been raised. Likewise, there is a rising need for automatic bias detection approaches. This paper adds...

💬 0 commentsarXiv:2601.03730v1PDF
0

Posted in cs.CV · 2026-01-07 · Donghwan Lee, Byeongjin Kim, Geunhee Kim, Hyukjin Kwon, Nahyeon Maeng, Wooju Kim

MATANet: A Multi-context Attention and Taxonomy-Aware Network for Fine-Grained Underwater Recognition of Marine Species

Fine-grained recognition of marine organisms is important for ecological research, biodiversity monitoring, and habitat conservation. However, existing methods often focus on the target organism alone, which can overlook informative cues from the surrounding environment. Moreover, biological taxonomy is often underused during model...

💬 0 commentsarXiv:2601.03729v3PDF
0

Posted in cs.CV · 2026-01-07 · Zhipeng Qian, Zihan Liang, Yufei Ma, Ben Chen, Huangyu Dai, Yiwei Ma, Jiayi Ji, Chenyi Lei, Han Li, Xiaoshuai Sun

CSMCIR: CoT-Enhanced Symmetric Alignment with Memory Bank for Composed Image Retrieval

Composed Image Retrieval (CIR) enables users to search for target images using both a reference image and manipulation text, offering substantial advantages over single-modality retrieval systems. However, existing CIR methods suffer from representation space fragmentation: queries and targets comprise heterogeneous modalities and are...

💬 0 commentsarXiv:2601.03728v3PDF
0

Posted in cs.CL · 2026-01-07 · Fadhil Muhammad, Alwin Djuliansah, Adrian Aryaputra Hamzah, Kurniawati Azizah

Stuttering-Aware Automatic Speech Recognition for Indonesian Language

Automatic speech recognition systems have achieved remarkable performance on fluent speech but continue to degrade significantly when processing stuttered speech, a limitation that is particularly acute for low-resource languages like Indonesian where specialized datasets are virtually non-existent. To overcome this scarcity, we...

💬 0 commentsarXiv:2601.03727v2PDF
0

Posted in cs.LG · 2026-01-07 · Jing-Cheng Pang, Liu Sun, Chang Zhou, Xian Tang, Haichuan Ma, Kun Jiang, Jianlong Wang, Kai Zhang, Sijie Wu, Haoran Cai, Chenwei Wu, Xubin Li, Xin Chen

EDCO: Dynamic Curriculum Orchestration for Domain-specific Large Language Model Fine-tuning

Domain-specific large language models (LLMs), typically developed by fine-tuning a pre-trained general-purpose LLM on specialized datasets, represent a significant advancement in applied AI. A common strategy in LLM fine-tuning is curriculum learning, which pre-orders training samples based on metrics like difficulty to improve...

💬 0 commentsarXiv:2601.03725v1PDF
0

Posted in cs.LG · 2026-01-07 · Shijie Zhang, Kevin Zhang, Zheyuan Gu, Xiang Guo, Rujun Guo, Shaoyu Liu, Guanjun Jiang, Xiaozhao Wang

ETR: Outcome-Guided Elastic Trust Regions for Policy Optimization

Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as an important paradigm for unlocking reasoning capabilities in large language models, exemplified by the success of OpenAI o1 and DeepSeek-R1. Currently, Group Relative Policy Optimization (GRPO) stands as the dominant algorithm in this domain due to its stable...

💬 0 commentsarXiv:2601.03723v1PDF
0

Posted in cs.CV · 2026-01-07 · Wenyong Li, Qi Jiang, Weijian Hu, Kailun Yang, Zhanjun Zhang, Wenjun Tian, Kaiwei Wang, Jian Bai

Towards Real-world Lens Active Alignment with Unlabeled Data via Domain Adaptation

Active Alignment (AA) is a key technology for the large-scale automated assembly of high-precision optical systems. Compared with labor-intensive per-model on-device calibration, a digital-twin pipeline built on optical simulation offers a substantial advantage in generating large-scale labeled data. However, complex imaging...

💬 0 commentsarXiv:2601.03718v2PDF
0

Posted in cs.CY · 2026-01-07 · Mark Theby

A Mixed Methods Systematic Analysis of Issues and Factors Influencing Organizational Cloud Computing Adoption and Usage in the Public Sector: Initial Findings

Cloud computing has been shown to be an essential enabling technology for public sector organizations PSOs and offers numerous potential benefits, including reduced information technology infrastructure costs, increased innovation potential, and improved resource resilience and scalability. Despite governments' intensifying efforts to...

💬 0 commentsarXiv:2601.06175v1PDF
0

Posted in cs.SD · 2026-01-07 · Benedikt Mayrhofer, Franz Pernkopf, Philipp Aichinger, Martin Hagmüller

Lightweight and perceptually-guided voice conversion for electro-laryngeal speech

Electro-laryngeal (EL) speech is characterized by constant pitch, limited prosody, and mechanical noise, reducing naturalness and intelligibility. We propose a lightweight adaptation of the state-of-the-art StreamVC framework to this setting by removing pitch and energy modules and combining self-supervised pretraining with supervised...

💬 0 commentsarXiv:2601.03892v2PDF
0

Posted in cs.LG · 2026-01-07 · Ibrahim Delibasoglu

Spectral Manifold Regularization for Stable and Modular Routing in Deep MoE Architectures

Mixture of Experts (MoE) architectures enable efficient scaling of neural networks but suffer from expert collapse, where routing converges to a few dominant experts. This reduces model capacity and causes catastrophic interference during adaptation. We propose the Spectrally-Regularized Mixture of Experts (SR-MoE), which imposes...

💬 0 commentsarXiv:2601.03889v1PDF
0

Posted in cs.CR · 2026-01-07 · Aakash Singh, Kuldeep Singh Yadav, V. Anil Kumar, Samiran Ghosh, Pranita Baro, Basavala Bhanu Prasanth

A Longitudinal Measurement Study of Log4Shell Exploitation from a Reactive Network Telescope

The disclosure of the Log4Shell vulnerability in December 2021 led to an unprecedented wave of global scanning and exploitation activity. A recent study provided important initial insights, but was largely limited in duration and geography, focusing primarily on European and U.S. network telescope deployments and covering the...

💬 0 commentsarXiv:2601.04281v2PDF