Qwen Councils
arXiv is taking too long to respond. Please try again or narrow your search.
Showing downloaded papers while arXiv is unavailable.

Computer Science

arXiv preprints from January 1, 2026 through July 20, 2026 — 16:41:00 EST

0

Posted in cs.AI · 2026-01-15 · Seoyeon Kim, Jaehyung Kim

SPRInG: Continual LLM Personalization via Selective Parametric Adaptation and Retrieval-Interpolated Generation

Personalizing Large Language Models typically relies on static retrieval or one-time adaptation, assuming user preferences remain invariant over time. However, real-world interactions are dynamic, where user interests continuously evolve, posing a challenge for models to adapt to preference drift without catastrophic forgetting....

💬 0 commentsarXiv:2601.09974v1PDF
0

Posted in cs.CC · 2026-01-15 · Samuel Everett

Correspondences in computational and dynamical complexity II: forcing complex reductions

An algebraic telic problem is a decision problem in $\textsf{NP}_\mathbb{R}$ formalizing finite-time reachability questions for one-dimensional dynamical systems. We prove that the existence of "natural" mapping reductions between algebraic telic problems coming from distinct dynamical systems implies the two dynamical systems exhibit...

💬 0 commentsarXiv:2601.09973v1PDF
0

Posted in cs.AI · 2026-01-15 · Zixun Lan, Maochun Xu, Yifan Ren, Rui Wu, Jianghui Zhou, Xueyang Cheng, Jianan Ding Ding, Xinheng Wang, Mingmin Chi, Fei Ma

Chinese Labor Law Large Language Model Benchmark

Recent advances in large language models (LLMs) have led to substantial progress in domain-specific applications, particularly within the legal domain. However, general-purpose models such as GPT-4 often struggle with specialized subdomains that require precise legal knowledge, complex reasoning, and contextual sensitivity. To address...

💬 0 commentsarXiv:2601.09972v1PDF
0

Posted in cs.LG · 2026-01-15 · Hansen He, Shuheng Li

An Exploratory Study to Repurpose LLMs to a Unified Architecture for Time Series Classification

Time series classification (TSC) is a core machine learning problem with broad applications. Recently there has been growing interest in repurposing large language models (LLMs) for TSC, motivated by their strong reasoning and generalization ability. Prior work has primarily focused on alignment strategies that explicitly map time...

💬 0 commentsarXiv:2601.09971v1PDF
0

Posted in cs.LG · 2026-01-15 · Ruoxi Jia, Luis Oala, Wenjie Xiong, Suqin Ge, Jiachen T. Wang, Feiyang Kang, Dawn Song

A Sustainable AI Economy Needs Data Deals That Work for Generators

We argue that the machine learning value chain is structurally unsustainable due to an economic data processing inequality: each state in the data cycle from inputs to model weights to synthetic outputs refines technical signal but strips economic equity from data generators. We show, by analyzing seventy-three public data deals, that...

💬 0 commentsarXiv:2601.09966v1PDF
0

Posted in cs.GT · 2026-01-15 · Zehua Cheng, Wei Dai, Zhipeng Wang, Rui Sun, Nick Wen, Jiahao Sun

A Control Theoretic Approach to Decentralized AI Economy Stabilization via Dynamic Buyback-and-Burn Mechanisms

The democratization of artificial intelligence through decentralized networks represents a paradigm shift in computational provisioning, yet the long-term viability of these ecosystems is critically endangered by the extreme volatility of their native economic layers. Current tokenomic models, which predominantly rely on static or...

💬 0 commentsarXiv:2601.09961v1PDF
0

Posted in cs.IT · 2026-01-15 · Yingying Huangfu, Tian Bai

On the Leaky Private Information Retrieval with Side Information

This paper investigates the problem of Leaky Private Information Retrieval with Side Information (L-PIR-SI), providing a fundamental characterization of the trade-off among leaky privacy, side information, and download cost. We propose a unified probabilistic framework to design L-PIR-SI schemes under $\varepsilon$-differential...

💬 0 commentsarXiv:2601.09960v2PDF
0

Posted in cs.IT · 2026-01-15 · Vayur Shanbhag, Prasad Krishnan

Private Information Retrieval for Graph-based Replication with Minimal Subpacketization

We design new minimal-subpacketization schemes for information-theoretic private information retrieval on graph-based replicated databases. In graph-based replication, the system consists of $K$ files replicated across $N$ servers according to a graph with $N$ vertices and $K$ edges. The client wants to retrieve one desired file,...

💬 0 commentsarXiv:2601.09957v1PDF
0

Posted in cs.CV · 2026-01-15 · Nahid Alam, Leema Krishna Murali, Siddhant Bharadwaj, Patrick Liu, Timothy Chung, Drishti Sharma, Akshata A, Kranthi Kiran, Wesley Tam, Bala Krishna S Vegesna

The Spatial Blindspot of Vision-Language Models

Vision-language models (VLMs) have advanced rapidly, but their ability to capture spatial relationships remains a blindspot. Current VLMs are typically built with contrastive language-image pretraining (CLIP) style image encoders. The training recipe often flattens images into 1D patch sequences, discarding the 2D structure necessary...

💬 0 commentsarXiv:2601.09954v2PDF
0

Posted in cs.CL · 2026-01-15 · Christabel Acquaye, Yi Ting Huang, Marine Carpuat, Rachel Rudinger

Take Out Your Calculators: Estimating the Real Difficulty of Question Items with LLM Student Simulations

Standardized math assessments require expensive human pilot studies to establish the difficulty of test items. We investigate the predictive value of open-source large language models (LLMs) for evaluating the difficulty of multiple-choice math questions for real-world students. We show that, while LLMs are poor direct judges of...

💬 0 commentsarXiv:2601.09953v2PDF
0

Posted in cs.CV · 2026-01-15 · Zhihua Zhao, Guoqiang Li, Chen Min, Kangping Lu

OT-Drive: Out-of-Distribution Off-Road Traversable Area Segmentation via Optimal Transport

Reliable traversable area segmentation in unstructured environments is critical for planning and decision-making in autonomous driving. However, existing data-driven approaches often suffer from degraded segmentation performance in out-of-distribution (OOD) scenarios, consequently impairing downstream driving tasks. To address this...

💬 0 commentsarXiv:2601.09952v1PDF
0

Posted in cs.LG · 2026-01-15 · Griffin Kearney

Kinematic Tokenization: Optimization-Based Continuous-Time Tokens for Learnable Decision Policies in Noisy Time Series

Transformers are designed for discrete tokens, yet many real-world signals are continuous processes observed through noisy sampling. Discrete tokenizations (raw values, patches, finite differences) can be brittle in low signal-to-noise regimes, especially when downstream objectives impose asymmetric penalties that rationally encourage...

💬 0 commentsarXiv:2601.09949v2PDF
0

Posted in cs.IT · 2026-01-15 · Shubhransh Singhvi, Han Mao Kiah, Eitan Yaakobi

Reconstructing Reed-Solomon Codes from Multiple Noisy Channel Outputs

The sequence reconstruction problem, introduced by Levenshtein in 2001, considers a communication setting in which a sender transmits a codeword and the receiver observes K independent noisy versions of this codeword. In this work, we study the problem of efficient reconstruction when each of the $K$ outputs is corrupted by a $q$-ary...

💬 0 commentsarXiv:2601.09947v1PDF
0

Posted in cs.LG · 2026-01-15 · Chenxi Qiu

Interpolation-Based Optimization for Enforcing lp-Norm Metric Differential Privacy in Continuous and Fine-Grained Domains

Metric Differential Privacy (mDP) generalizes Local Differential Privacy (LDP) by adapting privacy guarantees based on pairwise distances, enabling context-aware protection and improved utility. While existing optimization-based methods reduce utility loss effectively in coarse-grained domains, optimizing mDP in fine-grained or...

💬 0 commentsarXiv:2601.09946v1PDF
0

Posted in cs.CY · 2026-01-15 · Richard Q. Blackwell, Eman Hammad, Congrui Jin, Jisoo Park, Albert E. Patterson

Modeling conflicting incentives in engineering senior capstone projects: A multi-player game theory approach

University engineering capstone projects involve sustained interaction among students, faculty, and industry sponsors whose objectives are only partially aligned. While capstones are widely used in engineering education, existing analyses typically treat stakeholder behavior informally or descriptively, leaving incentive conflicts,...

💬 0 commentsarXiv:2601.09944v1PDF
0

Posted in cs.CV · 2026-01-15 · Wenwen Liao, Hang Ruan, Jianbo Yu, Yuansong Wang, Qingchao Jiang, Xiaofeng Yang

InfoSculpt: Sculpting the Latent Space for Generalized Category Discovery

Generalized Category Discovery (GCD) aims to classify instances from both known and novel categories within a large-scale unlabeled dataset, a critical yet challenging task for real-world, open-world applications. However, existing methods often rely on pseudo-labeling, or two-stage clustering, which lack a principled mechanism to...

💬 0 commentsarXiv:2601.10098v1PDF
0

Posted in cs.LG · 2026-01-15 · Piyush Singh Pasi

Multilingual-To-Multimodal (M2M): Unlocking New Languages with Monolingual Text

Multimodal models excel in English, supported by abundant image-text and audio-text data, but performance drops sharply for other languages due to limited multilingual multimodal resources. Existing solutions rely on machine translation, while advances in multilingual text modeling remain underutilized. We introduce M2M, a lightweight...

💬 0 commentsarXiv:2601.10096v2PDF
0

Posted in cs.CV · 2026-01-15 · Han Wang, Yi Yang, Jingyuan Hu, Minfeng Zhu, Wei Chen

V-Zero: Self-Improving Multimodal Reasoning with Zero Annotation

Recent advances in multimodal learning have significantly enhanced the reasoning capabilities of vision-language models (VLMs). However, state-of-the-art approaches rely heavily on large-scale human-annotated datasets, which are costly and time-consuming to acquire. To overcome this limitation, we introduce V-Zero, a general...

💬 0 commentsarXiv:2601.10094v1PDF
0

Posted in cs.SE · 2026-01-15 · Yiding Qiu, Seyed Mahdi Azimi, Artem Lensky

Mark My Works Autograder for Programming Courses

Large programming courses struggle to provide timely, detailed feedback on student code. We developed Mark My Works, a local autograding system that combines traditional unit testing with LLM-generated explanations. The system uses role-based prompts to analyze submissions, critique code quality, and generate pedagogical feedback...

💬 0 commentsarXiv:2601.10093v1PDF
0

Posted in cs.LG · 2026-01-15 · Jongseok Kim, Seongae Kang, Jonghwan Shin, Yuhan Lee, Ohyun Jo

LeMoF: Level-guided Multimodal Fusion for Heterogeneous Clinical Data

Multimodal clinical prediction is widely used to integrate heterogeneous data such as Electronic Health Records (EHR) and biosignals. However, existing methods tend to rely on static modality integration schemes and simple fusion strategies. As a result, they fail to fully exploit modality-specific representations. In this paper, we...

💬 0 commentsarXiv:2601.10092v1PDF
0

Posted in cs.CV · 2026-01-15 · Mingzhuo Li, Guang Li, Linfeng Ye, Jiafeng Mao, Takahiro Ogawa, Konstantinos N. Plataniotis, Miki Haseyama

Difficulty-guided Sampling: Bridging the Target Gap between Dataset Distillation and Downstream Tasks

In this paper, we propose difficulty-guided sampling (DGS) to bridge the target gap between the distillation objective and the downstream task, therefore improving the performance of dataset distillation. Deep neural networks achieve remarkable performance but have time and storage-consuming training processes. Dataset distillation is...

💬 0 commentsarXiv:2601.10090v1PDF
0

Posted in cs.LG · 2026-01-15 · Ashley Klein, Edward Raff, Marcia DesJardin

Bayesian Meta-Analyses Could Be More: A Case Study in Trial of Labor After a Cesarean-section Outcomes and Complications

The meta-analysis's utility is dependent on previous studies having accurately captured the variables of interest, but in medical studies, a key decision variable that impacts a physician's decisions was not captured. This results in an unknown effect size and unreliable conclusions. A Bayesian approach may allow analysis to determine...

💬 0 commentsarXiv:2601.10089v1PDF
0

Posted in cs.AI · 2026-01-15 · Malika Aubakirova, Alex Atallah, Chris Clark, Justin Summerville, Anjney Midha

State of AI: An Empirical 100 Trillion Token Study with OpenRouter

The past year has marked a turning point in the evolution and real-world use of large language models (LLMs). With the release of the first widely adopted reasoning model, o1, on December 5th, 2024, the field shifted from single-pass pattern generation to multi-step deliberation inference, accelerating deployment, experimentation, and...

💬 0 commentsarXiv:2601.10088v1PDF
0

Posted in cs.CL · 2026-01-15 · Viet Cuong Nguyen, Nhi Yen Nguyen, Kristin A. Candan, Mary Conlon, Vanessa Rumie, Kristen Risola, Michael L. Birnbaum, Munmun De Choudhury

CALM-IT: Generating Realistic Long-Form Motivational Interviewing Dialogues with Dual-Actor Conversational Dynamics Tracking

Therapeutic dialogue is not a sequence of isolated responses: client goals, motivation, resistance, and therapeutic alliance evolve over time. Yet current LLM-based mental health dialogue systems often lack explicit mechanisms for tracking these dynamics across extended interactions, which can lead to poorly timed interventions or...

💬 0 commentsarXiv:2601.10085v2PDF
0

Posted in cs.LG · 2026-01-15 · Zan Chaudhry, Noam H. Rotenberg, Brian Caffo, Craig K. Jones, Haris I. Sair

Adaptive Label Error Detection: A Bayesian Approach to Mislabeled Data Detection

Machine learning classification systems are susceptible to poor performance when trained with incorrect ground truth labels, even when data is well-curated by expert annotators. As machine learning becomes more widespread, it is increasingly imperative to identify and correct mislabeling to develop more powerful models. In this work,...

💬 0 commentsarXiv:2601.10084v1PDF