Qwen Councils
arXiv could not process that search. Try a simpler keyword search or an arXiv field query such as all:quantum.
Showing downloaded papers while arXiv is unavailable.

Computer Science

arXiv preprints from January 1, 2026 through July 21, 2026 — 17:40:49 EST

0

Posted in cs.LG · 2026-01-15 · Tianqi Zhang, Flavio Ponzina, Tajana Rosing

FaTRQ: Tiered Residual Quantization for LLM Vector Search in Far-Memory-Aware ANNS Systems

Approximate Nearest-Neighbor Search (ANNS) is a key technique in retrieval-augmented generation (RAG), enabling rapid identification of the most relevant high-dimensional embeddings from massive vector databases. Modern ANNS engines accelerate this process using prebuilt indexes and store compressed vector-quantized representations in...

💬 0 commentsarXiv:2601.09985v1PDF
0

Posted in cs.CL · 2026-01-15 · David Samuel Setiawan, Raphaël Merx, Jey Han Lau

Context Volume Drives Performance: Tackling Domain Shift in Extremely Low-Resource Translation via RAG

Neural Machine Translation (NMT) models for low-resource languages suffer significant performance degradation under domain shift. We quantify this challenge using Dhao, an indigenous language of Eastern Indonesia with no digital footprint beyond the New Testament (NT). When applied to the unseen Old Testament (OT), a standard NMT...

💬 0 commentsarXiv:2601.09982v2PDF
0

Posted in cs.CV · 2026-01-15 · Yulin He, Wei Chen, Zhikang Jian, Tianhang Guo, Wenjuan Zhou, Minglong Li, Shaowu Yang, Wenjing Yang

DR$^2$Seg: Decomposed Two-Stage Rollouts for Efficient Reasoning Segmentation in Multimodal Large Language Models

Reasoning segmentation is an emerging vision-language task that requires reasoning over intricate text queries to precisely segment objects. However, existing methods typically suffer from overthinking, generating verbose reasoning chains that interfere with object localization in multimodal large language models (MLLMs). To address...

💬 0 commentsarXiv:2601.09981v2PDF
0

Posted in cs.LG · 2026-01-15 · Frank Cole, Dixi Wang, Yineng Chen, Yulong Lu, Rongjie Lai

In-Context Operator Learning on the Space of Probability Measures

We introduce \emph{in-context operator learning on probability measure spaces} for optimal transport (OT). The goal is to learn a single solution operator that maps a pair of distributions to the OT map, using only few-shot samples from each distribution as a prompt and \emph{without} gradient updates at inference. We parameterize the...

💬 0 commentsarXiv:2601.09979v1PDF
0

Posted in cs.NI · 2026-01-15 · Jie Zheng, Ruichen Zhang, Dusit Niyato, Haijun Zhang, Jiacheng Wang, Hongyang Du, Jiawen Kang, Zehui Xiong

Large Language Model (LLM)-enabled Reinforcement Learning for Wireless Network Optimization

Enhancing future wireless networks presents a significant challenge for networking systems due to diverse user demands and the emergence of 6G technology. While reinforcement learning (RL) is a powerful framework, it often encounters difficulties with high-dimensional state spaces and complex environments, leading to substantial...

💬 0 commentsarXiv:2602.13210v1PDF
0

Posted in cs.DC · 2026-01-15 · Jer Shyuan Ng, Wathsara Daluwatta, Shehan Edirimannage, Charitha Elvitigala, Asitha Kottahachchi Kankanamge Don, Ibrahim Khalil, Heng Zhang, Dusit Niyato

Federated Unlearning in Edge Networks: A Survey of Fundamentals, Challenges, Practical Applications and Future Directions

The proliferation of connected devices and privacy-sensitive applications has accelerated the adoption of Federated Learning (FL), a decentralized paradigm that enables collaborative model training without sharing raw data. While FL addresses data locality and privacy concerns, it does not inherently support data deletion requests...

💬 0 commentsarXiv:2601.09978v1PDF
0

Posted in cs.AI · 2026-01-15 · Seoyeon Kim, Jaehyung Kim

SPRInG: Continual LLM Personalization via Selective Parametric Adaptation and Retrieval-Interpolated Generation

Personalizing Large Language Models typically relies on static retrieval or one-time adaptation, assuming user preferences remain invariant over time. However, real-world interactions are dynamic, where user interests continuously evolve, posing a challenge for models to adapt to preference drift without catastrophic forgetting....

💬 0 commentsarXiv:2601.09974v1PDF
0

Posted in cs.CC · 2026-01-15 · Samuel Everett

Correspondences in computational and dynamical complexity II: forcing complex reductions

An algebraic telic problem is a decision problem in $\textsf{NP}_\mathbb{R}$ formalizing finite-time reachability questions for one-dimensional dynamical systems. We prove that the existence of "natural" mapping reductions between algebraic telic problems coming from distinct dynamical systems implies the two dynamical systems exhibit...

💬 0 commentsarXiv:2601.09973v1PDF
0

Posted in cs.AI · 2026-01-15 · Zixun Lan, Maochun Xu, Yifan Ren, Rui Wu, Jianghui Zhou, Xueyang Cheng, Jianan Ding Ding, Xinheng Wang, Mingmin Chi, Fei Ma

Chinese Labor Law Large Language Model Benchmark

Recent advances in large language models (LLMs) have led to substantial progress in domain-specific applications, particularly within the legal domain. However, general-purpose models such as GPT-4 often struggle with specialized subdomains that require precise legal knowledge, complex reasoning, and contextual sensitivity. To address...

💬 0 commentsarXiv:2601.09972v1PDF
0

Posted in cs.LG · 2026-01-15 · Hansen He, Shuheng Li

An Exploratory Study to Repurpose LLMs to a Unified Architecture for Time Series Classification

Time series classification (TSC) is a core machine learning problem with broad applications. Recently there has been growing interest in repurposing large language models (LLMs) for TSC, motivated by their strong reasoning and generalization ability. Prior work has primarily focused on alignment strategies that explicitly map time...

💬 0 commentsarXiv:2601.09971v1PDF
0

Posted in cs.LG · 2026-01-15 · Ruoxi Jia, Luis Oala, Wenjie Xiong, Suqin Ge, Jiachen T. Wang, Feiyang Kang, Dawn Song

A Sustainable AI Economy Needs Data Deals That Work for Generators

We argue that the machine learning value chain is structurally unsustainable due to an economic data processing inequality: each state in the data cycle from inputs to model weights to synthetic outputs refines technical signal but strips economic equity from data generators. We show, by analyzing seventy-three public data deals, that...

💬 0 commentsarXiv:2601.09966v1PDF
0

Posted in cs.GT · 2026-01-15 · Zehua Cheng, Wei Dai, Zhipeng Wang, Rui Sun, Nick Wen, Jiahao Sun

A Control Theoretic Approach to Decentralized AI Economy Stabilization via Dynamic Buyback-and-Burn Mechanisms

The democratization of artificial intelligence through decentralized networks represents a paradigm shift in computational provisioning, yet the long-term viability of these ecosystems is critically endangered by the extreme volatility of their native economic layers. Current tokenomic models, which predominantly rely on static or...

💬 0 commentsarXiv:2601.09961v1PDF
0

Posted in cs.IT · 2026-01-15 · Yingying Huangfu, Tian Bai

On the Leaky Private Information Retrieval with Side Information

This paper investigates the problem of Leaky Private Information Retrieval with Side Information (L-PIR-SI), providing a fundamental characterization of the trade-off among leaky privacy, side information, and download cost. We propose a unified probabilistic framework to design L-PIR-SI schemes under $\varepsilon$-differential...

💬 0 commentsarXiv:2601.09960v2PDF
0

Posted in cs.IT · 2026-01-15 · Vayur Shanbhag, Prasad Krishnan

Private Information Retrieval for Graph-based Replication with Minimal Subpacketization

We design new minimal-subpacketization schemes for information-theoretic private information retrieval on graph-based replicated databases. In graph-based replication, the system consists of $K$ files replicated across $N$ servers according to a graph with $N$ vertices and $K$ edges. The client wants to retrieve one desired file,...

💬 0 commentsarXiv:2601.09957v1PDF
0

Posted in cs.CV · 2026-01-15 · Nahid Alam, Leema Krishna Murali, Siddhant Bharadwaj, Patrick Liu, Timothy Chung, Drishti Sharma, Akshata A, Kranthi Kiran, Wesley Tam, Bala Krishna S Vegesna

The Spatial Blindspot of Vision-Language Models

Vision-language models (VLMs) have advanced rapidly, but their ability to capture spatial relationships remains a blindspot. Current VLMs are typically built with contrastive language-image pretraining (CLIP) style image encoders. The training recipe often flattens images into 1D patch sequences, discarding the 2D structure necessary...

💬 0 commentsarXiv:2601.09954v2PDF
0

Posted in cs.CL · 2026-01-15 · Christabel Acquaye, Yi Ting Huang, Marine Carpuat, Rachel Rudinger

Take Out Your Calculators: Estimating the Real Difficulty of Question Items with LLM Student Simulations

Standardized math assessments require expensive human pilot studies to establish the difficulty of test items. We investigate the predictive value of open-source large language models (LLMs) for evaluating the difficulty of multiple-choice math questions for real-world students. We show that, while LLMs are poor direct judges of...

💬 0 commentsarXiv:2601.09953v2PDF
0

Posted in cs.CV · 2026-01-15 · Zhihua Zhao, Guoqiang Li, Chen Min, Kangping Lu

OT-Drive: Out-of-Distribution Off-Road Traversable Area Segmentation via Optimal Transport

Reliable traversable area segmentation in unstructured environments is critical for planning and decision-making in autonomous driving. However, existing data-driven approaches often suffer from degraded segmentation performance in out-of-distribution (OOD) scenarios, consequently impairing downstream driving tasks. To address this...

💬 0 commentsarXiv:2601.09952v1PDF
0

Posted in cs.LG · 2026-01-15 · Griffin Kearney

Kinematic Tokenization: Optimization-Based Continuous-Time Tokens for Learnable Decision Policies in Noisy Time Series

Transformers are designed for discrete tokens, yet many real-world signals are continuous processes observed through noisy sampling. Discrete tokenizations (raw values, patches, finite differences) can be brittle in low signal-to-noise regimes, especially when downstream objectives impose asymmetric penalties that rationally encourage...

💬 0 commentsarXiv:2601.09949v2PDF
0

Posted in cs.IT · 2026-01-15 · Shubhransh Singhvi, Han Mao Kiah, Eitan Yaakobi

Reconstructing Reed-Solomon Codes from Multiple Noisy Channel Outputs

The sequence reconstruction problem, introduced by Levenshtein in 2001, considers a communication setting in which a sender transmits a codeword and the receiver observes K independent noisy versions of this codeword. In this work, we study the problem of efficient reconstruction when each of the $K$ outputs is corrupted by a $q$-ary...

💬 0 commentsarXiv:2601.09947v1PDF
0

Posted in cs.LG · 2026-01-15 · Chenxi Qiu

Interpolation-Based Optimization for Enforcing lp-Norm Metric Differential Privacy in Continuous and Fine-Grained Domains

Metric Differential Privacy (mDP) generalizes Local Differential Privacy (LDP) by adapting privacy guarantees based on pairwise distances, enabling context-aware protection and improved utility. While existing optimization-based methods reduce utility loss effectively in coarse-grained domains, optimizing mDP in fine-grained or...

💬 0 commentsarXiv:2601.09946v1PDF
0

Posted in cs.CY · 2026-01-15 · Richard Q. Blackwell, Eman Hammad, Congrui Jin, Jisoo Park, Albert E. Patterson

Modeling conflicting incentives in engineering senior capstone projects: A multi-player game theory approach

University engineering capstone projects involve sustained interaction among students, faculty, and industry sponsors whose objectives are only partially aligned. While capstones are widely used in engineering education, existing analyses typically treat stakeholder behavior informally or descriptively, leaving incentive conflicts,...

💬 0 commentsarXiv:2601.09944v1PDF
0

Posted in cs.CV · 2026-01-15 · Wenwen Liao, Hang Ruan, Jianbo Yu, Yuansong Wang, Qingchao Jiang, Xiaofeng Yang

InfoSculpt: Sculpting the Latent Space for Generalized Category Discovery

Generalized Category Discovery (GCD) aims to classify instances from both known and novel categories within a large-scale unlabeled dataset, a critical yet challenging task for real-world, open-world applications. However, existing methods often rely on pseudo-labeling, or two-stage clustering, which lack a principled mechanism to...

💬 0 commentsarXiv:2601.10098v1PDF
0

Posted in cs.LG · 2026-01-15 · Piyush Singh Pasi

Multilingual-To-Multimodal (M2M): Unlocking New Languages with Monolingual Text

Multimodal models excel in English, supported by abundant image-text and audio-text data, but performance drops sharply for other languages due to limited multilingual multimodal resources. Existing solutions rely on machine translation, while advances in multilingual text modeling remain underutilized. We introduce M2M, a lightweight...

💬 0 commentsarXiv:2601.10096v2PDF
0

Posted in cs.CV · 2026-01-15 · Han Wang, Yi Yang, Jingyuan Hu, Minfeng Zhu, Wei Chen

V-Zero: Self-Improving Multimodal Reasoning with Zero Annotation

Recent advances in multimodal learning have significantly enhanced the reasoning capabilities of vision-language models (VLMs). However, state-of-the-art approaches rely heavily on large-scale human-annotated datasets, which are costly and time-consuming to acquire. To overcome this limitation, we introduce V-Zero, a general...

💬 0 commentsarXiv:2601.10094v1PDF
0

Posted in cs.SE · 2026-01-15 · Yiding Qiu, Seyed Mahdi Azimi, Artem Lensky

Mark My Works Autograder for Programming Courses

Large programming courses struggle to provide timely, detailed feedback on student code. We developed Mark My Works, a local autograding system that combines traditional unit testing with LLM-generated explanations. The system uses role-based prompts to analyze submissions, critique code quality, and generate pedagogical feedback...

💬 0 commentsarXiv:2601.10093v1PDF