Qwen Councils
arXiv could not process that search. Try a simpler keyword search or an arXiv field query such as all:quantum.
Showing downloaded papers while arXiv is unavailable.

Computer Science

arXiv preprints from January 1, 2026 through July 20, 2026 — 11:44:02 EST

0

Posted in cs.CL · 2026-01-11 · Davide Baldelli, Ali Parviz, Amal Zouaq, Sarath Chandar

LLMs Can't Play Hangman: On the Necessity of a Private Working Memory for Language Agents

As LLMs move from text completion toward autonomous agents, they remain constrained by the standard chat interface, which lacks private working memory. This raises a fundamental question: can agents reliably perform interactive tasks that depend on hidden state? We define Private State Interactive Tasks (PSITs), which require agents...

💬 0 commentsarXiv:2601.06973v1PDF
0

Posted in cs.CL · 2026-01-11 · Nathan Roll, Pranav Bhalerao, Martijn Bartelds, Arjun Pawar, Yuka Tatsumi, Tolulope Ogunremi, Chen Shani, Calbert Graham, Meghan Sumner, Dan Jurafsky

Categorize Early, Integrate Late: Divergent Processing Strategies in Automatic Speech Recognition

In speech language modeling, two architectures dominate the frontier: the Transformer and the Conformer. However, it remains unknown whether their comparable performance stems from convergent processing strategies or distinct architectural inductive biases. We introduce Architectural Fingerprinting, a probing framework that isolates...

💬 0 commentsarXiv:2601.06972v2PDF
0

Posted in cs.IT · 2026-01-11 · Qinshan Zhang, Bin Chen, Yong Jiang, Shu-Tao Xia

Generalization Bounds for Transformer Channel Decoders

Transformer channel decoders, such as the Error Correction Code Transformer (ECCT), have shown strong empirical performance in channel decoding, yet their generalization behavior remains theoretically unclear. This paper studies the generalization performance of ECCT from a learning-theoretic perspective. By establishing a connection...

💬 0 commentsarXiv:2601.06969v1PDF
0

Posted in cs.LG · 2026-01-11 · Jinduo Guo, Yinzhi Cao

A Robust Certified Machine Unlearning Method Under Distribution Shift

The Newton method has been widely adopted to achieve certified unlearning. A critical assumption in existing approaches is that the data requested for unlearning are selected i.i.d.(independent and identically distributed). However,the problem of certified unlearning under non-i.i.d. deletions remains largely unexplored. In practice,...

💬 0 commentsarXiv:2601.06967v1PDF
0

Posted in cs.CL · 2026-01-11 · Haonan Bian, Zhiyuan Yao, Sen Hu, Zishan Xu, Shaolei Zhang, Yifu Guo, Ziliang Yang, Xueran Han, Huacan Wang, Ronghao Chen

RealMem: Benchmarking LLMs in Real-World Memory-Driven Interaction

As Large Language Models (LLMs) evolve from static dialogue interfaces to autonomous general agents, effective memory is paramount to ensuring long-term consistency. However, existing benchmarks primarily focus on casual conversation or task-oriented dialogue, failing to capture **"long-term project-oriented"** interactions where...

💬 0 commentsarXiv:2601.06966v1PDF
0

Posted in cs.CV · 2026-01-11 · Yu Zhong, Tianwei Lin, Ruike Zhu, Yuqian Yuan, Haoyu Zheng, Liang Liang, Wenqiao Zhang, Feifei Shao, Haoyuan Li, Wanggui He, Hao Jiang, Yueting Zhuang

Unified Personalized Understanding, Generating and Editing

Unified large multimodal models (LMMs) have achieved remarkable progress in general-purpose multimodal understanding and generation. However, they still operate under a ``one-size-fits-all'' paradigm and struggle to model user-specific concepts (e.g., generate a photo of \texttt{<maeve>}) in a consistent and controllable manner....

💬 0 commentsarXiv:2601.06965v1PDF
0

Posted in cs.LG · 2026-01-11 · Vladimer Khasia

HAS-VQ: Hessian-Adaptive Sparse Vector Quantization for High-Fidelity LLM Compression

Post-training quantization is essential for deploying Large Language Models (LLMs) on resource-constrained devices. However, standard integer quantization (e.g., INT4) fundamentally degrades performance by imposing a uniform grid on the heavy-tailed distribution of weight parameters, particularly in smaller-scale models (e.g., <2B...

💬 0 commentsarXiv:2601.06959v1PDF
0

Posted in cs.CC · 2026-01-11 · Holger Boche, Volker Pohl, H. Vincent Poor

Arithmetic Complexity of Solutions of the Dirichlet Problem

The classical Dirichlet problem on the unit disk can be solved by different numerical approaches. The two most common and popular approaches are the integration of the associated Poisson integral and, by applying Dirichlet's principle, solving a particular minimization problem. For practical use, these procedures need to be...

💬 0 commentsarXiv:2601.06954v1PDF
0

Posted in cs.CL · 2026-01-11 · Jie Wu, Haoling Li, Xin Zhang, Jiani Guo, Jane Luo, Steven Liu, Yangyu Huang, Ruihang Chu, Scarlett Li, Yujiu Yang

X-Coder: Advancing Competitive Programming with Fully Synthetic Tasks, Solutions, and Tests

Competitive programming poses a significant challenge for Code LLMs. While recent models have shown promise, they heavily rely on finite real-world data, raising concerns about scalability and contamination. In this paper, we investigate a critical question: Can we elevate models to expert-level reasoning performance using fully...

💬 0 commentsarXiv:2601.06953v2PDF
0

Posted in cs.CR · 2026-01-11 · Zhuoran Tan, Ke Xiao, Jeremy Singer, Christos Anagnostopoulos

Operational Runtime Behavior Mining for Open-Source Supply Chain Security

Open-source software (OSS) is a critical component of modern software systems, yet supply chain security remains challenging in practice due to unavailable or obfuscated source code. Consequently, security teams often rely on runtime observations collected from sandboxed executions to investigate suspicious third-party components. We...

💬 0 commentsarXiv:2601.06948v2PDF
0

Posted in cs.LG · 2026-01-11 · Deyu Cao, Yixin Yin, Samin Aref

Sliced-Wasserstein Distribution Alignment Loss Improves the Ultra-Low-Bit Quantization of Large Language Models

The benefits of most large language models come with steep and often hidden economic and environmental costs due to their resource usage inefficiency during deployment. Model quantization improves energy and memory efficiency through representing model parameters by lower-precision values. However, compression below 4-bits often...

💬 0 commentsarXiv:2601.07878v1PDF
0

Posted in cs.DS · 2026-01-11 · Mateus de Oliveira Oliveira, Wim Van den Broeck

Optimal Extended Formulations from Optimal Dynamic Programming Algorithms

Vertex Subset Problems (VSPs) are a class of combinatorial optimization problems on graphs where the goal is to find a subset of vertices satisfying a predefined condition. Two prominent approaches for solving VSPs are dynamic programming over tree-like structures, such as tree decompositions or clique decompositions, and linear...

💬 0 commentsarXiv:2601.06947v2PDF
0

Posted in cs.CV · 2026-01-11 · Yuhang Su, Mei Wang, Yaoyao Zhong, Guozhang Li, Shixing Li, Yihan Feng, Hua Huang

SketchJudge: A Diagnostic Benchmark for Grading Hand-drawn Diagrams with Multimodal Large Language Models

While Multimodal Large Language Models (MLLMs) have achieved remarkable progress in visual understanding, they often struggle when faced with the unstructured and ambiguous nature of human-generated sketches. This limitation is particularly pronounced in the underexplored task of visual grading, where models should not only solve a...

💬 0 commentsarXiv:2601.06944v1PDF
0

Posted in cs.CV · 2026-01-11 · Chengwen Liu, Xiaomin Yu, Zhuoyue Chang, Zhe Huang, Shuo Zhang, Heng Lian, Jisheng Dang, Rui Xu, Sen Hu, Jianheng Hou, Chengwei Qin, Xiaobin Hu, Kunyi Wang, Zhi Yang, Hao Peng, Hong Peng, Ronghao Chen, Huacan Wang

Watching, Reasoning, and Searching: A Video Deep Research Benchmark on Open Web for Agentic Video Reasoning

In real-world video question answering scenarios, videos often provide only localized visual cues, while verifiable answers are distributed across the open web; models therefore need to jointly perform cross-frame clue extraction, iterative retrieval, and multi-hop reasoning-based verification. To bridge this gap, we construct the...

💬 0 commentsarXiv:2601.06943v2PDF
0

Posted in cs.LG · 2026-01-11 · James Tlhomole, Edoardo Borgomeo, Karthikeyan Matheswaran, Mariangel Garcia Andarcia

Towards Operational Streamflow Forecasting in the Limpopo River Basin using Long Short-Term Memory Networks

Robust hydrological simulation is key for sustainable development, water management strategies, and climate change adaptation. In recent years, deep learning methods have been demonstrated to outperform mechanistic models at the task of hydrological discharge simulation. Adoption of these methods has been catalysed by the...

💬 0 commentsarXiv:2601.06941v1PDF
0

Posted in cs.DB · 2026-01-11 · Hengyu Liu, Tianyi Li, Haoyu Wang, Kristian Torp, Tiancheng Zhang, Yushuai Li, Christian S. Jensen

VISTA: Knowledge-Driven Vessel Trajectory Imputation with Repair Provenance

Repairing incomplete trajectory data is essential for downstream spatio-temporal applications. Yet, existing repair methods focus solely on reconstruction without documenting the reasoning behind repair decisions, undermining trust in safety-critical applications where repaired trajectories affect operational decisions, such as in...

💬 0 commentsarXiv:2601.06940v2PDF
0

Posted in cs.LG · 2026-01-11 · Heng Xu, Tianqing Zhu, Dayong Ye, Lefeng Zhang, Le Wang, Wanlei Zhou

Forgetting Similar Samples: Can Machine Unlearning Do it Better?

Machine unlearning, a process enabling pre-trained models to remove the influence of specific training samples, has attracted significant attention in recent years. Although extensive research has focused on developing efficient machine unlearning strategies, we argue that these methods mainly aim at removing samples rather than...

💬 0 commentsarXiv:2601.06938v1PDF
0

Posted in cs.AI · 2026-01-11 · Fozle Rabbi Shafi, M. Anwar Hossain, Salimur Choudhury

mind_call: A Dataset for Mental Health Function Calling with Large Language Models

Large Language Model (LLM)-based systems increasingly rely on function calling to enable structured and controllable interaction with external data sources, yet existing datasets do not address mental health-oriented access to wearable sensor data. This paper presents a synthetic function-calling dataset designed for mental health...

💬 0 commentsarXiv:2601.06937v1PDF
0

Posted in cs.CL · 2026-01-11 · Stephen Gadd

Symphonym: Universal Phonetic Embeddings for Cross-Script Name Matching

Matching place names across writing systems is a persistent obstacle to the integration of multilingual geographic sources, whether modern gazetteers, medieval itineraries, or colonial-era surveys. Existing approaches depend on language-specific phonetic algorithms or romanisation steps that discard phonetic information, and none...

💬 0 commentsarXiv:2601.06932v4PDF
0

Posted in cs.CV · 2026-01-11 · Haodong Chen, Qiang Huang, Jiaqi Zhao, Qiuping Jiang, Xiaojun Chang, Jun Yu

Measuring Social Bias in Vision-Language Models with Face-Only Counterfactuals from Real Photos

Vision-Language Models (VLMs) are increasingly deployed in socially consequential settings, raising concerns about social bias driven by demographic cues. A central challenge in measuring such social bias is attribution under visual confounding: real-world images entangle race and gender with correlated factors such as background and...

💬 0 commentsarXiv:2601.06931v2PDF
0

Posted in cs.CV · 2026-01-11 · Shenghao Zhang, Runtao Liu, Christopher Schroers, Yang Zhang

RenderFlow: Single-Step Neural Rendering via Flow Matching

Conventional physically based rendering (PBR) pipelines generate photorealistic images through computationally intensive light transport simulations. Although recent deep learning approaches leverage diffusion model priors with geometry buffers (G-buffers) to produce visually compelling results without explicit scene geometry or light...

💬 0 commentsarXiv:2601.06928v2PDF
0

Posted in cs.IT · 2026-01-11 · Hui Zhao, Dirk Slock, Petros Elia

Caching Yields up to 5x Spectral Efficiency in Multi-Beam Satellite Communications

This paper examines the integration of vector coded caching (VCC) into multi-beam satellite communications (SATCOM) systems and demonstrates that even limited receiver-side caching can substantially enhance spectral efficiency. By leveraging cached content to suppress interference, VCC enables the concurrent transmission of multiple...

💬 0 commentsarXiv:2601.06925v1PDF
0

Posted in cs.CL · 2026-01-11 · Tianhua Zhang, Kun Li, Junan Li, Yunxiang Li, Hongyin Luo, Xixin Wu, James Glass, Helen Meng

TreePS-RAG: Tree-based Process Supervision for Reinforcement Learning in Agentic RAG

Agentic retrieval-augmented generation (RAG) formulates question answering as a multi-step interaction between reasoning and information retrieval, and has recently been advanced by reinforcement learning (RL) with outcome-based supervision. While effective, relying solely on sparse final rewards limits step-wise credit assignment and...

💬 0 commentsarXiv:2601.06922v1PDF
0

Posted in cs.IT · 2026-01-11 · Tadashi Wadayama, Takumi Takahashi

Score-Based VAMP with Fisher-Information-Based Onsager Correction

We propose score-based VAMP (SC-VAMP), a variant of vector approximate message passing (VAMP) in which the Onsager correction is expressed and computed via conditional Fisher information, thereby enabling a Jacobian-free implementation. Using learned score functions, SC-VAMP constructs nonlinear MMSE estimators through Tweedie's...

💬 0 commentsarXiv:2601.07095v1PDF
0

Posted in cs.CV · 2026-01-11 · Peiyuan Jing, Yue Yang, Chun-Wun Cheng, Zhenxuan Zhang, Liutao Yang, Thiago V. Lima, Klaus Strobel, Antoine Leimgruber, Angelica Aviles-Rivero, Guang Yang, Javier A. Montoya-Zegarra

3D Wavelet-Based Structural Priors for Controlled Diffusion in Whole-Body Low-Dose PET Denoising

Low-dose Positron Emission Tomography (PET) imaging reduces patient radiation exposure but suffers from increased noise that degrades image quality and diagnostic reliability. Although diffusion models have demonstrated strong denoising capability, their stochastic nature makes it challenging to enforce anatomically consistent...

💬 0 commentsarXiv:2601.07093v4PDF