Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through July 20, 2026 — 10:01:12 EST

0

Posted in cs.CV · 2026-01-14 · Haoyan Gong, Hongbin Liu

LP-LLM: End-to-End Real-World Degraded License Plate Text Recognition via Large Multimodal Models

Real-world License Plate Recognition (LPR) faces significant challenges from severe degradations such as motion blur, low resolution, and complex illumination. The prevailing "restoration-then-recognition" two-stage paradigm suffers from a fundamental flaw: the pixel-level optimization objectives of image restoration models are...

💬 0 commentsarXiv:2601.09116v1PDF
0

Posted in cs.CR · 2026-01-14 · Fengchao Chen, Tingmin Wu, Van Nguyen, Surya. Nepal, Carsten Rudolph

Agents at Risk: How Users Unwittingly Undermine LLM Safety

Large language model (LLM)-based agents are increasingly deployed in applications, such as trip-planning agents and web-use agents, to perform complex planning and execution tasks. Prior work has shown that LLM-based agents are vulnerable to context confusion, where external adversarial content incorporated into the agent's reasoning...

💬 0 commentsarXiv:2601.10758v3PDF
0

Posted in cs.DC · 2026-01-14 · Yufan Xia, Marco De La Pierre, Amanda S. Barnard, Giuseppe Maria Junior Barca

A Machine Learning Approach Towards Runtime Optimisation of Matrix Multiplication

The GEneral Matrix Multiplication (GEMM) is one of the essential algorithms in scientific computing. Single-thread GEMM implementations are well-optimised with techniques like blocking and autotuning. However, due to the complexity of modern multi-core shared memory systems, it is challenging to determine the number of threads that...

💬 0 commentsarXiv:2601.09114v1PDF
0

Posted in cs.AI · 2026-01-14 · Zixia Jia, Jiaqi Li, Yipeng Kang, Yuxuan Wang, Tong Wu, Quansen Wang, Xiaobo Wang, Shuyi Zhang, Junzhe Shen, Qing Li, Siyuan Qi, Yitao Liang, Di He, Zilong Zheng, Song-Chun Zhu

The AI Hippocampus: How Far are We From Human Memory?

Memory plays a foundational role in augmenting the reasoning, adaptability, and contextual fidelity of modern Large Language Models and Multi-Modal LLMs. As these models transition from static predictors to interactive systems capable of continual learning and personalized inference, the incorporation of memory mechanisms has emerged...

💬 0 commentsarXiv:2601.09113v1PDF
0

Posted in cs.CY · 2026-01-14 · Ying He, Baiyang Li, Yule Cao, Huirun Xu, Qiuxian Chen, Shu Chen, Shangsheng Ren

Seeking Human Security Consensus: A Unified Value Scale for Generative AI Value Safety

The rapid development of generative AI has brought value- and ethics-related risks to the forefront, making value safety a critical concern while a unified consensus remains lacking. In this work, we propose an internationally inclusive and resilient unified value framework, the GenAI Value Safety Scale (GVS-Scale): Grounded in a...

💬 0 commentsarXiv:2601.09112v1PDF
0

Posted in cs.CV · 2026-01-14 · Yang Li, Aming Wu, Zihao Zhang, Yahong Han

Towards Open Environments and Instructions: General Vision-Language Navigation via Fast-Slow Interactive Reasoning

Vision-Language Navigation (VLN) aims to enable agents to navigate to a target location based on language instructions. Traditional VLN often follows a close-set assumption, i.e., training and test data share the same style of the input images and instructions. However, the real world is open and filled with various unseen...

💬 0 commentsarXiv:2601.09111v2PDF
0

Posted in cs.CL · 2026-01-14 · Ziyang Zhou, Ziqi Liu, Yan Wang, Yiming Lin, Yangbin Chen

RAM-SD: Retrieval-Augmented Multi-agent framework for Sarcasm Detection

Sarcasm detection remains a significant challenge due to its reliance on nuanced contextual understanding, world knowledge, and multi-faceted linguistic cues that vary substantially across different sarcastic expressions. Existing approaches, from fine-tuned transformers to large language models, apply a uniform reasoning strategy to...

💬 0 commentsarXiv:2601.17002v1PDF
0

Posted in cs.CV · 2026-01-14 · Kai Hu, Yaozu Feng, Vladimir Lysenko, Ya Guo, Huayi Wu

SAM-Aug: Leveraging SAM Priors for Few-Shot Parcel Segmentation in Satellite Time Series

Few-shot semantic segmentation of time-series remote sensing images remains a critical challenge, particularly in regions where labeled data is scarce or costly to obtain. While state-of-the-art models perform well under full supervision, their performance degrades significantly under limited labeling, limiting their real-world...

💬 0 commentsarXiv:2601.09110v2PDF
0

Posted in cs.CV · 2026-01-14 · Yanguang Sun, Chao Wang, Jian Yang, Lei Luo

Small but Mighty: Dynamic Wavelet Expert-Guided Fine-Tuning of Large-Scale Models for Optical Remote Sensing Object Segmentation

Accurately localizing and segmenting relevant objects from optical remote sensing images (ORSIs) is critical for advancing remote sensing applications. Existing methods are typically built upon moderate-scale pre-trained models and employ diverse optimization strategies to achieve promising performance under full-parameter...

💬 0 commentsarXiv:2601.09108v1PDF
0

Posted in cs.CV · 2026-01-14 · Lachlan Holden, Feras Dayoub, Alberto Candela, David Harvey, Tat-Jun Chin

Vision Foundation Models for Domain Generalisable Cross-View Localisation in Planetary Ground-Aerial Robotic Teams

Accurate localisation in planetary robotics enables the advanced autonomy required to support the increased scale and scope of future missions. The successes of the Ingenuity helicopter and multiple planetary orbiters lay the groundwork for future missions that use ground-aerial robotic teams. In this paper, we consider rovers using...

💬 0 commentsarXiv:2601.09107v1PDF
0

Posted in cs.AI · 2026-01-14 · Wenbin Li, Jingling Wu, Xiaoyong Lin. Jing Chen, Cong Chen

AviationLMM: A Large Multimodal Foundation Model for Civil Aviation

Civil aviation is a cornerstone of global transportation and commerce, and ensuring its safety, efficiency and customer satisfaction is paramount. Yet conventional Artificial Intelligence (AI) solutions in aviation remain siloed and narrow, focusing on isolated tasks or single modalities. They struggle to integrate heterogeneous data...

💬 0 commentsarXiv:2601.09105v2PDF
0

Posted in cs.RO · 2026-01-14 · Ko Yamamoto, Kyosuke Ishibashi, Hiroki Ishikawa, Osamu Azami

Design Methodology of Hydraulically-driven Soft Robotic Gripper for a Large and Heavy Object

This paper presents a design methodology of a hydraulically-driven soft robotic gripper for grasping a large and heavy object -- approximately 10 - 20 kg with 20 - 30 cm diameter. Most existing soft grippers are pneumatically actuated with several hundred kPa pressure, and cannot generate output force sufficient for such a large and...

💬 0 commentsarXiv:2601.09104v1PDF
0

Posted in cs.LG · 2026-01-14 · Haijian Shao, Wei Liu, Xing Deng, Daze Lu

Enhancing Imbalanced Electrocardiogram Classification: A Novel Approach Integrating Data Augmentation through Wavelet Transform and Interclass Fusion

Imbalanced electrocardiogram (ECG) data hampers the efficacy and resilience of algorithms in the automated processing and interpretation of cardiovascular diagnostic information, which in turn impedes deep learning-based ECG classification. Notably, certain cardiac conditions that are infrequently encountered are disproportionately...

💬 0 commentsarXiv:2601.09103v1PDF
0

Posted in cs.CV · 2026-01-14 · Haonan Wei, Linyuan Wang, Nuolin Sun, Zhizhong Zheng, Lei Li, Bin Yan

A one-step generation model with a Single-Layer Transformer: Layer number re-distillation of FreeFlow

Currently, Flow matching methods aim to compress the iterative generation process of diffusion models into a few or even a single step, with MeanFlow and FreeFlow being representative achievements of one-step generation based on Ordinary Differential Equations (ODEs). We observe that the 28-layer Transformer architecture of FreeFlow...

💬 0 commentsarXiv:2601.11630v1PDF
0

Posted in cs.AI · 2026-01-14 · Lixiang Zhang, Chenggong Zhao, Qing Gao, Xiaoke Zhao, Gengyi Bai, Jinhu Lv

DScheLLM: Enabling Dynamic Scheduling through a Fine-Tuned Dual-System Large language Model

Production scheduling is highly susceptible to dynamic disruptions, such as variations in processing times, machine availability, and unexpected task insertions. Conventional approaches typically rely on event-specific models and explicit analytical formulations, which limits their adaptability and generalization across previously...

💬 0 commentsarXiv:2601.09100v2PDF
0

Posted in cs.CR · 2026-01-14 · Nghia T. Le, Alan Ritter, Kartik Goyal

Semantic Differentiation for Tackling Challenges in Watermarking Low-Entropy Constrained Generation Outputs

We demonstrate that while the current approaches for language model watermarking are effective for open-ended generation, they are inadequate at watermarking LM outputs for constrained generation tasks with low-entropy output spaces. Therefore, we devise SeqMark, a sequence-level watermarking algorithm with semantic differentiation...

💬 0 commentsarXiv:2601.11629v1PDF
0

Posted in cs.IT · 2026-01-14 · Yifeng Qin, Jing Chen, Zhi Hao Jiang, Zhi Ning Chen, Yongming Huang, Lingyang Song

Airy Beamforming for Radiative Near-Field MU-XL-MIMO: Overcoming Half-Space Blockage

The move to next-generation wireless communications with extremely large-scale antenna arrays (ELAAs) brings the communications into the radiative near-field (RNF) region, where distance-aware focusing is feasible. However, high-frequency RNF links are highly vulnerable to blockage in indoor environments dominated by half-space...

💬 0 commentsarXiv:2601.09098v3PDF
0

Posted in cs.AI · 2026-01-14 · Derrick Goh Xin Deik, Quanyu Long, Zhengyuan Liu, Nancy F. Chen, Wenya Wang

Programming over Thinking: Efficient and Robust Multi-Constraint Planning

Multi-constraint planning involves identifying, evaluating, and refining candidate plans while satisfying multiple, potentially conflicting constraints. Existing large language model (LLM) approaches face fundamental limitations in this domain. Pure reasoning paradigms, which rely on long natural language chains, are prone to...

💬 0 commentsarXiv:2601.09097v4PDF
0

Posted in cs.LG · 2026-01-14 · Md Asiful Islam, Md Ahmed Al Muzaddid, Afia Jahin Prema, Sreenath Reddy Vuske

Comparative Assessment of Concrete Compressive Strength Prediction at Industry Scale Using Embedding-based Neural Networks, Transformers, and Traditional Machine Learning Approaches

Concrete is the most widely used construction material worldwide; however, reliable prediction of compressive strength remains challenging due to material heterogeneity, variable mix proportions, and sensitivity to field and environmental conditions. Recent advances in artificial intelligence enable data-driven modeling frameworks...

💬 0 commentsarXiv:2601.09096v1PDF
0

Posted in cs.LG · 2026-01-14 · Zhixiang Liang, Beichen Huang, Zheng Wang, Minjia Zhang

Hidden States as Early Signals: Step-level Trace Evaluation and Pruning for Efficient Test-Time Scaling

Large Language Models (LLMs) can enhance reasoning capabilities through test-time scaling by generating multiple traces. However, the combination of lengthy reasoning traces with multiple sampling introduces substantial computation and high end-to-end latency. Prior work on accelerating this process has relied on similarity-based or...

💬 0 commentsarXiv:2601.09093v2PDF
0

Posted in cs.CR · 2026-01-14 · Christopher Blake, Chen Feng, Xuachao Wang, Qianyu Yu

Merged Bitcoin: Proof of Work Blockchains with Multiple Hash Types

Proof of work blockchain protocols using multiple hash types are considered. It is proven that the security region of such a protocol cannot be the AND of a 51\% attack on all the hash types. Nevertheless, a protocol called Merged Bitcoin is introduced, which is the Bitcoin protocol where links between blocks can be formed using...

💬 0 commentsarXiv:2601.09090v1PDF
0

Posted in cs.CL · 2026-01-14 · Shuyang Hou, Yi Hu, Muhan Zhang

SubTokenTest: A Practical Benchmark for Real-World Sub-token Understanding

Recent advancements in large language models (LLMs) have significantly enhanced their reasoning capabilities. However, they continue to struggle with basic character-level tasks, such as counting letters in words, a problem rooted in their tokenization process. While existing benchmarks have highlighted this weakness through basic...

💬 0 commentsarXiv:2601.09089v1PDF
0

Posted in cs.LG · 2026-01-14 · Shaotian Yan, Kaiyuan Liu, Chen Shen, Bing Wang, Sinan Fan, Jun Zhang, Yue Wu, Zheng Wang, Jieping Ye

Distribution-Aligned Sequence Distillation for Superior Long-CoT Reasoning

In this report, we introduce DASD-4B-Thinking, a lightweight yet highly capable, fully open-source reasoning model. It achieves SOTA performance among open-source models of comparable scale across challenging benchmarks in mathematics, scientific reasoning, and code generation -- even outperforming several larger models. We begin by...

💬 0 commentsarXiv:2601.09088v1PDF
0

Posted in cs.LG · 2026-01-14 · Kangda Wei, Ruihong Huang

MMR-GRPO: Accelerating GRPO-Style Training through Diversity-Aware Reward Reweighting

Group Relative Policy Optimization (GRPO) has become a standard approach for training mathematical reasoning models; however, its reliance on multiple completions per prompt makes training computationally expensive. Although recent work has reduced the number of training steps required to reach peak performance, the overall wall-clock...

💬 0 commentsarXiv:2601.09085v2PDF
0

Posted in cs.CL · 2026-01-14 · Wilson Y. Lee

How Many Human Judgments Are Enough? Feasibility Limits of Human Preference Evaluation

Human preference evaluations are widely used to compare generative models, yet it remains unclear how many judgments are required to reliably detect small improvements. We show that when preference signal is diffuse across prompts (i.e., all prompt types are similarly informative), proportional allocation is minimax-optimal: no...

💬 0 commentsarXiv:2601.09084v2PDF