Qwen Councils
arXiv is taking too long to respond. Please try again or narrow your search.
Showing downloaded papers while arXiv is unavailable.

Computer Science

arXiv preprints from January 1, 2026 through July 20, 2026 — 23:07:00 EST

0

Posted in cs.CL · 2026-01-15 · Zhanming Shen, Jiaqi Hu, Zeyu Qin, Hao Chen, Wentao Ye, Zenan Huang, Yihong Zhuang, Guoshan Lu, Junlin Zhou, Junbo Zhao

Training-Trajectory-Aware Token Selection

Efficient distillation is a key pathway for converting expensive reasoning capability into deployable efficiency, yet in the frontier regime where the student already has strong reasoning ability, naive continual distillation often yields limited gains or even degradation. We observe a characteristic training phenomenon: even as loss...

💬 0 commentsarXiv:2601.10348v2PDF
0

Posted in cs.SD · 2026-01-15 · Yunyi Liu, Taketo Akama

Self-supervised restoration of singing voice degraded by pitch shifting using shallow diffusion

Pitch shifting has been an essential feature in singing voice production. However, conventional signal processing approaches exhibit well known trade offs such as formant shifts and robotic coloration that becomes more severe at larger transposition jumps. This paper targets high quality pitch shifting for singing by reframing it as a...

💬 0 commentsarXiv:2601.10345v1PDF
0

Posted in cs.CL · 2026-01-15 · Deming Ding, Shichun Liu, Enhui Yang, Jiahang Lin, Ziying Chen, Shihan Dou, Honglin Guo, Weiyu Cheng, Pengyu Zhao, Chengjun Xiao, Qunhong Zeng, Qi Zhang, Xuanjing Huang, Qidi Xu, Tao Gui

OctoBench: Benchmarking Scaffold-Aware Instruction Following in Repository-Grounded Agentic Coding

Modern coding scaffolds turn LLMs into capable software agents, but their ability to follow scaffold-specified instructions remains under-examined, especially when constraints are heterogeneous and persist across interactions. To fill this gap, we introduce OctoBench, which benchmarks scaffold-aware instruction following in...

💬 0 commentsarXiv:2601.10343v2PDF
0

Posted in cs.AI · 2026-01-15 · Cheng Lin Cheng, Ting Chuan Lin, Chai Kai Chang

C-GRASP: Clinically-Grounded Reasoning for Affective Signal Processing

Heart rate variability (HRV) is a pivotal noninvasive marker for autonomic monitoring; however, applying Large Language Models (LLMs) to HRV interpretation is hindered by physiological hallucinations. These include respiratory sinus arrhythmia (RSA) contamination, short-data instability in nonlinear metrics, and the neglect of...

💬 0 commentsarXiv:2601.10342v1PDF
0

Posted in cs.IT · 2026-01-15 · Anina Gruica, Benjamin Jany, Stanislav Kruglik

Convertible Codes for Data and Device Heterogeneity

Distributed storage systems must handle both data heterogeneity, arising from non-uniform access demands, and device heterogeneity, caused by time-varying node reliability. In this paper, we study convertible codes, which enable the transformation of one code into another with minimum cost in the merge regime, addressing the latter....

💬 0 commentsarXiv:2601.10341v1PDF
0

Posted in cs.RO · 2026-01-15 · David Morilla-Cabello, Eduardo Montijano

CHORAL: Traversal-Aware Planning for Safe and Efficient Heterogeneous Multi-Robot Routing

Monitoring large, unknown, and complex environments with autonomous robots poses significant navigation challenges, where deploying teams of heterogeneous robots with complementary capabilities can substantially improve both mission performance and feasibility. However, effectively modeling how different robotic platforms interact...

💬 0 commentsarXiv:2601.10340v1PDF
0

Posted in cs.CR · 2026-01-15 · Yi Liu, Weizhe Wang, Ruitao Feng, Yao Zhang, Guangquan Xu, Gelei Deng, Yuekang Li, Leo Zhang

Agent Skills in the Wild: An Empirical Study of Security Vulnerabilities at Scale

The rise of AI agent frameworks has introduced agent skills, modular packages containing instructions and executable code that dynamically extend agent capabilities. While this architecture enables powerful customization, skills execute with implicit trust and minimal vetting, creating a significant yet uncharacterized attack surface....

💬 0 commentsarXiv:2601.10338v1PDF
0

Posted in cs.CV · 2026-01-15 · Minh Hai Nguyen, Quoc Bao Do, Edouard Pauwels, Pierre Weiss

An analytic theory of convolutional neural network inverse problems solvers

Supervised convolutional neural networks (CNNs) are widely used to solve imaging inverse problems, achieving state-of-the-art performance in numerous applications. However, despite their empirical success, these methods are poorly understood from a theoretical perspective and often treated as black boxes. To bridge this gap, we...

💬 0 commentsarXiv:2601.10334v2PDF
0

Posted in cs.CV · 2026-01-15 · Siqi Kou, Jiachun Jin, Zetong Zhou, Ye Ma, Yugang Wang, Quan Chen, Peng Jiang, Xiao Yang, Jun Zhu, Kai Yu, Zhijie Deng

Think-Then-Generate: Reasoning-Aware Text-to-Image Diffusion with LLM Encoders

Recent progress in text-to-image (T2I) diffusion models (DMs) has enabled high-quality visual synthesis from diverse textual prompts. Yet, most existing T2I DMs, even those equipped with large language model (LLM)-based text encoders, remain text-pixel mappers -- they employ LLMs merely as text encoders, without leveraging their...

💬 0 commentsarXiv:2601.10332v1PDF
0

Posted in cs.IT · 2026-01-15 · Yuval Gerzon, Ilan Shomorony, Nir Weinberger

On the Capacity of Noisy Frequency-based Channels

We investigate the capacity of noisy frequency-based channels, motivated by DNA data storage in the short-molecule regime, where information is encoded in the frequency of items types rather than their order. The channel output is a histogram formed by random sampling of items, followed by noisy item identification. While the capacity...

💬 0 commentsarXiv:2601.10329v1PDF
0

Posted in cs.LG · 2026-01-15 · Yiqing Zou, Hanning Yuan, Qianyu Yang, Ziqiang Yuan, Shuliang Wang, Sijie Ruan

Meta Dynamic Graph for Traffic Flow Prediction

Traffic flow prediction is a typical spatio-temporal prediction problem and has a wide range of applications. The core challenge lies in modeling the underlying complex spatio-temporal dependencies. Various methods have been proposed, and recent studies show that the modeling of dynamics is useful to meet the core challenge. While...

💬 0 commentsarXiv:2601.10328v1PDF
0

Posted in cs.CV · 2026-01-15 · Yiming Zhang, Weibo Qin, Yuntian Liu, Feng Wang

SRAW-Attack: Space-Reweighted Adversarial Warping Attack for SAR Target Recognition

Synthetic aperture radar (SAR) imagery exhibits intrinsic information sparsity due to its unique electromagnetic scattering mechanism. Despite the widespread adoption of deep neural network (DNN)-based SAR automatic target recognition (SAR-ATR) systems, they remain vulnerable to adversarial examples and tend to over-rely on background...

💬 0 commentsarXiv:2601.10324v2PDF
0

Posted in cs.CV · 2026-01-15 · Xueyun Tian, Wei Li, Bingbing Xu, Heng Dong, Yuanzhuo Wang, Huawei Shen

ROMA: Real-time Omni-Multimodal Assistant with Interactive Streaming Understanding

Recent Omni-multimodal Large Language Models show promise in unified audio, vision, and text modeling. However, streaming audio-video understanding remains challenging, as existing approaches suffer from disjointed capabilities: they typically exhibit incomplete modality support or lack autonomous proactive monitoring. To address...

💬 0 commentsarXiv:2601.10323v1PDF
0

Posted in cs.CL · 2026-01-15 · Warren Jouanneau, Emma Jouffroy, Marc Palyart

An Efficient Long-Context Ranking Architecture With Calibrated LLM Distillation: Application to Person-Job Fit

Finding the most relevant person for a job proposal in real time is challenging, especially when resumes are long, structured, and multilingual. In this paper, we propose a re-ranking model based on a new generation of late cross-attention architecture, that decomposes both resumes and project briefs to efficiently handle long-context...

💬 0 commentsarXiv:2601.10321v2PDF
0

Posted in cs.CL · 2026-01-15 · Songsong Tian, Kongsheng Zhuo, Zhendong Wang, Rong Shen, Shengtao Zhang, Yong Wu

Boundary-Aware NL2SQL: Integrating Reliability through Hybrid Reward and Data Synthesis

In this paper, we present BAR-SQL (Boundary-Aware Reliable NL2SQL), a unified training framework that embeds reliability and boundary awareness directly into the generation process. We introduce a Seed Mutation data synthesis paradigm that constructs a representative enterprise corpus, explicitly encompassing multi-step analytical...

💬 0 commentsarXiv:2601.10318v1PDF
0

Posted in cs.CL · 2026-01-15 · Aniket Deroy

ADVOSYNTH: A Synthetic Multi-Advocate Dataset for Speaker Identification in Courtroom Scenarios

As large-scale speech-to-speech models achieve high fidelity, the distinction between synthetic voices in structured environments becomes a vital area of study. This paper introduces Advosynth-500, a specialized dataset comprising 100 synthetic speech files featuring 10 unique advocate identities. Using the Speech Llama Omni model, we...

💬 0 commentsarXiv:2601.10315v1PDF
0

Posted in cs.CV · 2026-01-15 · Peng-Fei Zhang, Zi Huang

Hierarchical Refinement of Universal Multimodal Attacks on Vision-Language Models

Existing adversarial attacks for VLP models are mostly sample-specific, resulting in substantial computational overhead when scaled to large datasets or new scenarios. To overcome this limitation, we propose Hierarchical Refinement Attack (HRA), a multimodal universal attack framework for VLP models. For the image modality, we refine...

💬 0 commentsarXiv:2601.10313v3PDF
0

Posted in cs.LG · 2026-01-15 · Zhipeng Liu, Peibo Duan, Xuan Tang, Haodong Jing, Mingyang Geng, Yongsheng Huang, Jialu Xu, Bin Zhang, Binwu Wang

We Need a More Robust Classifier: Dual Causal Learning Empowers Domain-Incremental Time Series Classification

The World Wide Web thrives on intelligent services that rely on accurate time series classification, which has recently witnessed significant progress driven by advances in deep learning. However, existing studies face challenges in domain incremental learning. In this paper, we propose a lightweight and robust dual-causal...

💬 0 commentsarXiv:2601.10312v1PDF
0

Posted in cs.CL · 2026-01-15 · Jan Christian Blaise Cruz, David Ifeoluwa Adelani, Alham Fikri Aji

Multilinguality as Sense Adaptation

We approach multilinguality as sense adaptation: aligning latent meaning representations across languages rather than relying solely on shared parameters and scale. In this paper, we introduce SENse-based Symmetric Interlingual Alignment (SENSIA), which adapts a Backpack language model from one language to another by explicitly...

💬 0 commentsarXiv:2601.10310v1PDF
0

Posted in cs.CL · 2026-01-15 · Luoming Hu, Jingjie Zeng, Liang Yang, Hongfei Lin

The Straight and Narrow: Do LLMs Possess an Internal Moral Path?

Enhancing the moral alignment of Large Language Models (LLMs) is a critical challenge in AI safety. Current alignment techniques often act as superficial guardrails, leaving the intrinsic moral representations of LLMs largely untouched. In this paper, we bridge this gap by leveraging Moral Foundations Theory (MFT) to map and...

💬 0 commentsarXiv:2601.10307v1PDF
0

Posted in cs.AI · 2026-01-15 · Xin Guan, Zijian Li, Shen Huang, Pengjun Xie, Jingren Zhou, Jiuxin Cao

Evidence-Augmented Policy Optimization with Reward Co-Evolution for Long-Context Reasoning

While Reinforcement Learning (RL) has advanced LLM reasoning, applying it to long-context scenarios is hindered by sparsity of outcome rewards. This limitation fails to penalize ungrounded "lucky guesses," leaving the critical process of needle-in-a-haystack evidence retrieval largely unsupervised. To address this, we propose EAPO...

💬 0 commentsarXiv:2601.10306v2PDF
0

Posted in cs.CV · 2026-01-15 · Hengyu Shen, Tiancheng Gu, Bin Qin, Lan Wu, Yuling Wu, Shuo Tan, Zelong Sun, Jun Wang, Nan Wu, Xiang An, Weidong Cai, Ziyong Feng, Kaicheng Yang

DanQing: An Up-to-Date Large-Scale Chinese Vision-Language Pre-training Dataset

Vision-Language Pre-training (VLP) models have achieved remarkable success by leveraging large-scale image-text pairs. While English-centric models like CLIP and SigLIP benefit from massive datasets (e.g., LAION-400M), the development of Chinese VLP remains bottlenecked by the lack of high-quality, large-scale open-source data. In...

💬 0 commentsarXiv:2601.10305v3PDF
0

Posted in cs.MA · 2026-01-15 · Zhenyu Zhao, Tiankui Zhang, Xiaoxia Xu, Junjie Li, Yuanwei Liu, Wenjuan Xing

Multipath Routing for Multi-Hop UAV Networks

Multi-hop uncrewed aerial vehicle (UAV) networks are promising to extend the terrestrial network coverage. Existing multi-hop UAV networks employ a single routing path by selecting the next-hop forwarding node in a hop-by-hop manner, which leads to local congestion and increases traffic delays. In this paper, a novel traffic-adaptive...

💬 0 commentsarXiv:2601.10299v1PDF
0

Posted in cs.AI · 2026-01-15 · Nina Bočková, Barbora Volná, Mirko Dohnal

Optimisation of complex product innovation processes based on trend models with three-valued logic

This paper investigates complex product-innovation processes using models grounded in a set of heuristics. Each heuristic is expressed through simple trends -- increasing, decreasing, or constant -- which serve as minimally information-intensive quantifiers, avoiding reliance on numerical values or rough sets. A solution to a trend...

💬 0 commentsarXiv:2601.10768v1PDF
0

Posted in cs.CR · 2026-01-15 · Yuansen Liu, Yixuan Tang, Anthony Kum Hoe Tun

Reasoning Hijacking: The Fragility of Reasoning Alignment in Large Language Models

Current LLM safety research predominantly focuses on mitigating Goal Hijacking, preventing attackers from redirecting a model's high-level objective (e.g., from "summarizing emails" to "phishing users"). In this paper, we argue that this perspective is incomplete and highlight a critical vulnerability in Reasoning Alignment. We expose...

💬 0 commentsarXiv:2601.10294v6PDF