Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through July 28, 2026 — 01:03:50 EST

0

Posted in cs.CV · 2026-01-05 · Yujie Hu, Zecheng Tang, Xu Jiang, Weiqi Li, Jian Zhang

TalkPhoto: A Versatile Training-Free Conversational Assistant for Intelligent Image Editing

Thanks to the powerful language comprehension capabilities of Large Language Models (LLMs), existing instruction-based image editing methods have introduced Multimodal Large Language Models (MLLMs) to promote information exchange between instructions and images, ensuring the controllability and flexibility of image editing. However,...

💬 0 commentsarXiv:2601.01915v1PDF
0

Posted in cs.CV · 2026-01-05 · Zhibo Wang, Zuoyuan Zhang, Xiaoyi Pang, Qile Zhang, Xuanyi Hao, Shuguo Zhuo, Peng Sun

TAP-ViTs: Task-Adaptive Pruning for On-Device Deployment of Vision Transformers

Vision Transformers (ViTs) have demonstrated strong performance across a wide range of vision tasks, yet their substantial computational and memory demands hinder efficient deployment on resource-constrained mobile and edge devices. Pruning has emerged as a promising direction for reducing ViT complexity. However, existing approaches...

💬 0 commentsarXiv:2601.02437v1PDF
0

Posted in cs.CV · 2026-01-05 · Arjun Ramesh Kaushik, Nalini K. Ratha, Venu Govindaraju

Learning Action Hierarchies via Hybrid Geometric Diffusion

Temporal action segmentation is a critical task in video understanding, where the goal is to assign action labels to each frame in a video. While recent advances leverage iterative refinement-based strategies, they fail to explicitly utilize the hierarchical nature of human actions. In this work, we propose HybridTAS - a novel...

💬 0 commentsarXiv:2601.01914v1PDF
0

Posted in cs.AI · 2026-01-05 · Minh Hieu Ha, Khanh Ly Ta, Hung Phan, Tung Doan, Tung Dao, Dao Tran, Huynh Thi Thanh Binh

MMP-A*: Multimodal Perception Enhanced Incremental Heuristic Search on Path Planning

Autonomous path planning requires a synergy between global reasoning and geometric precision, especially in complex or cluttered environments. While classical A* is valued for its optimality, it incurs prohibitive computational and memory costs in large-scale scenarios. Recent attempts to mitigate these limitations by using Large...

💬 0 commentsarXiv:2601.01910v2PDF
0

Posted in cs.CV · 2026-01-05 · Jingjing Wang, Qianglin Liu, Zhuo Xiao, Xinning Yao, Bo Liu, Lu Li, Lijuan Niu, Fugen Zhou

Nodule-DETR: A Novel DETR Architecture with Frequency-Channel Attention for Ultrasound Thyroid Nodule Detection

Thyroid cancer is the most common endocrine malignancy, and its incidence is rising globally. While ultrasound is the preferred imaging modality for detecting thyroid nodules, its diagnostic accuracy is often limited by challenges such as low image contrast and blurred nodule boundaries. To address these issues, we propose...

💬 0 commentsarXiv:2601.01908v1PDF
0

Posted in cs.LG · 2026-01-05 · Yuxuan Li, Harshith Reddy Kethireddy, Srijita Das

Evaluating Feature Dependent Noise in Preference-based Reinforcement Learning

Learning from Preferences in Reinforcement Learning (PbRL) has gained attention recently, as it serves as a natural fit for complicated tasks where the reward function is not easily available. However, preferences often come with uncertainty and noise if they are not from perfect teachers. Much prior literature aimed to detect noise,...

💬 0 commentsarXiv:2601.01904v2PDF
0

Posted in cs.LG · 2026-01-05 · Ungsik Kim, Suwon Lee

TT-FSI: Scalable Faithful Shapley Interactions via Tensor-Train

The Faithful Shapley Interaction (FSI) index uniquely satisfies the faithfulness axiom among Shapley interaction indices, but computing FSI requires $O(d^\ell \cdot 2^d)$ time and existing implementations use $O(4^d)$ memory. We present TT-FSI, which exploits FSI's algebraic structure via Matrix Product Operators (MPO). Our main...

💬 0 commentsarXiv:2601.01903v1PDF
0

Posted in cs.LG · 2026-01-05 · Yuexuan Xia, Yinghao Zhang, Yalin Liu, Hong-Ning Dai, Yong Xia

FedBiCross: Personalized One-Shot Federated Learning on Medical Images

Data-free knowledge distillation-based one-shot federated learning (OSFL) trains a model in a single communication round without sharing raw data, making OSFL attractive for privacy-sensitive medical applications. However, existing methods aggregate predictions from all clients to form a global teacher. Under non-IID data, conflicting...

💬 0 commentsarXiv:2601.01901v4PDF
0

Posted in cs.NE · 2026-01-05 · Yiran Tian, Yuanjia Liu

Multi-strategy Improved Northern Goshawk Optimization for WSN Coverage Enhancement

To enhance the coverage rate of Wireless Sensor Networks (WSNs), this paper proposes an advanced optimization strategy based on a multi-strategy integrated Northern Goshawk Optimization (NGO) algorithm. Specifically, multivariate chaotic mapping is first employed to improve the randomness and uniformity of the initial population. To...

💬 0 commentsarXiv:2601.01898v1PDF
0

Posted in cs.IR · 2026-01-05 · Lilu Cheng, Jingjun Lu, Yi Xuan Chan, Quoc Khai Nguyen, John Bi, Sean Ho

A Hybrid Architecture for Multi-Stage Claim Document Understanding: Combining Vision-Language Models and Machine Learning for Real-Time Processing

Claims documents are fundamental to healthcare and insurance operations, serving as the basis for reimbursement, auditing, and compliance. However, these documents are typically not born digital; they often exist as scanned PDFs or photographs captured under uncontrolled conditions. Consequently, they exhibit significant content...

💬 0 commentsarXiv:2601.01897v1PDF
0

Posted in cs.CL · 2026-01-05 · Jingyu Liu, Jiaen Lin, Yong Liu

Tackling the Inherent Difficulty of Noise Filtering in RAG

Retrieval-Augmented Generation (RAG) has become a widely adopted approach to enhance Large Language Models (LLMs) by incorporating external knowledge and reducing hallucinations. However, noisy or irrelevant documents are often introduced during RAG, potentially degrading performance and even causing hallucinated outputs. While...

💬 0 commentsarXiv:2601.01896v2PDF
0

Posted in cs.MM · 2026-01-05 · William Han, Tony Chen, Chaojing Duan, Xiaoyu Song, Yihang Yao, Yuzhe Yang, Michael A. Rosenberg, Emerson Liu, Ding Zhao

ELF: A Family of Encoder-Free ECG-Language Models

ECG-Language Models (ELMs) extend recent advances in Multimodal Large Language Models (MLLMs) to automated ECG interpretation. However, most existing ELMs inherit Vision-Language Model (VLM) design choices and rely on pretrained ECG encoders, introducing substantial architectural and training complexity. Inspired by encoder-free VLMs,...

💬 0 commentsarXiv:2601.18798v3PDF
0

Posted in cs.CV · 2026-01-05 · Arjun Ramesh Kaushik, Naresh Kumar Devulapally, Vishnu Suresh Lokhande, Nalini K. Ratha, Venu Govindaraju

Forget Less by Learning from Parents Through Hierarchical Relationships

Custom Diffusion Models (CDMs) offer impressive capabilities for personalization in generative modeling, yet they remain vulnerable to catastrophic forgetting when learning new concepts sequentially. Existing approaches primarily focus on minimizing interference between concepts, often neglecting the potential for positive...

💬 0 commentsarXiv:2601.01892v1PDF
0

Posted in cs.CV · 2026-01-05 · Niloufar Alipour Talemi, Julia Boone, Fatemeh Afghah

Agentic AI in Remote Sensing: Foundations, Taxonomy, and Emerging Systems

The paradigm of Earth Observation analysis is shifting from static deep learning models to autonomous agentic AI. Although recent vision foundation models and multimodal large language models advance representation learning, they often lack the sequential planning and active tool orchestration required for complex geospatial...

💬 0 commentsarXiv:2601.01891v1PDF
0

Posted in cs.DB · 2026-01-05 · Yifan Wu, Yuhan Li, Zhenhua Wang, Zhongle Xie, Dingyu Yang, Ke Chen, Lidan Shou, Bo Tang, Liang Lin, Huan Li, Gang Chen

SafeLoad: Efficient Admission Control Framework for Identifying Memory-Overloading Queries in Cloud Data Warehouses

Memory overload is a common form of resource exhaustion in cloud data warehouses. When database queries fail due to memory overload, it not only wastes critical resources such as CPU time but also disrupts the execution of core business processes, as memory-overloading (MO) queries are typically part of complex workflows. If such...

💬 0 commentsarXiv:2601.01888v1PDF
0

Posted in cs.LG · 2026-01-05 · Jiawen Zhang, Lipeng He, Kejia Chen, Jian Lou, Jian Liu, Xiaohu Yang, Ruoxi Jia

Safety at One Shot: Patching Fine-Tuned LLMs with A Single Instance

Fine-tuning safety-aligned large language models (LLMs) can substantially compromise their safety. Previous approaches require many safety samples or calibration sets, which not only incur significant computational overhead during realignment but also lead to noticeable degradation in model utility. Contrary to this belief, we show...

💬 0 commentsarXiv:2601.01887v2PDF
0

Posted in cs.CL · 2026-01-05 · Yi Yu, Liuyi Yao, Yuexiang Xie, Qingquan Tan, Jiaqi Feng, Yaliang Li, Libing Wu

Agentic Memory: Learning Unified Long-Term and Short-Term Memory Management for Large Language Model Agents

Large language model (LLM) agents face fundamental limitations in long-horizon reasoning due to finite context windows, making effective memory management critical. Existing methods typically handle long-term memory (LTM) and short-term memory (STM) as separate components, relying on heuristics or auxiliary controllers, which limits...

💬 0 commentsarXiv:2601.01885v2PDF
0

Posted in cs.SI · 2026-01-05 · Yichao Yao, Minyu Feng, Matjaž Perc, Jürgen Kurths

Fixed-Size Dynamic Scale-Free Networks: Modeling, Stationarity, and Resilience

Many real-world scale-free networks, such as neural networks and online communication networks, consist of a fixed number of nodes but exhibit dynamic edge fluctuations. However, traditional models frequently overlook scenarios where the node count remains constant, instead prioritizing node growth. In this work, we depart from the...

💬 0 commentsarXiv:2601.01882v1PDF
0

Posted in cs.LG · 2026-01-05 · Yifang Zhang, Shengwu Xiong, Henan Wang, Wenjie Yin, Jiawang Peng, Duan Zhou, Yuqiang Zhang, Chen Zhou, Hua Chen, Qile Zhao, Pengfei Duan

RainBalance: Alleviating Dual Imbalance in GNSS-based Precipitation Nowcasting via Continuous Probability Modeling

Global navigation satellite systems (GNSS) station-based Precipitation Nowcasting aims to predict rainfall within the next 0-6 hours by leveraging a GNSS station's historical observations of precipitation, GNSS-PWV, and related meteorological variables, which is crucial for disaster mitigation and real-time decision-making. In recent...

💬 0 commentsarXiv:2601.06137v1PDF
0

Posted in cs.AI · 2026-01-05 · Farzan Karimi-Malekabadi, Suhaib Abdurahman, Zhivar Sourati, Jackson Trager, Morteza Dehghani

Theory Trace Card: Theory-Driven Socio-Cognitive Evaluation of LLMs

Socio-cognitive benchmarks for large language models (LLMs) often fail to predict real-world behavior, even when models achieve high benchmark scores. Prior work has attributed this evaluation-deployment gap to problems of measurement and validity. While these critiques are insightful, we argue that they overlook a more fundamental...

💬 0 commentsarXiv:2601.01878v1PDF
0

Posted in cs.AI · 2026-01-05 · Kewen Cao, Jianxu Chen, Yongbing Zhang, Ye Zhang, Hongxiao Wang

Toward Auditable Neuro-Symbolic Reasoning in Pathology: SQL as an Explicit Trace of Evidence

Automated pathology image analysis is central to clinical diagnosis, but clinicians still ask which slide features drive a model's decision and why. Vision-language models can produce natural language explanations, but these are often correlational and lack verifiable evidence. In this paper, we introduce an SQL-centered agentic...

💬 0 commentsarXiv:2601.01875v1PDF
0

Posted in cs.CV · 2026-01-05 · Shuhang Chen, Yunqiu Xu, Junjie Xie, Aojun Lu, Tao Feng, Zeying Huang, Ning Zhang, Yi Sun, Yi Yang, Hangjie Yuan

CogFlow: Bridging Perception and Reasoning through Knowledge Internalization for Visual Mathematical Problem Solving

Despite significant progress, multimodal large language models continue to struggle with visual mathematical problem solving. Some recent works recognize that visual perception is a bottleneck in visual mathematical reasoning, but their solutions are limited to improving the extraction and interpretation of visual inputs. Notably,...

💬 0 commentsarXiv:2601.01874v3PDF
0

Posted in cs.RO · 2026-01-05 · Hongbo Duan, Shangyi Luo, Zhiyuan Deng, Yanbo Chen, Yuanhao Chiang, Yi Liu, Fangming Liu, Xueqian Wang

CausalNav: A Long-term Embodied Navigation System for Autonomous Mobile Robots in Dynamic Outdoor Scenarios

Autonomous language-guided navigation in large-scale outdoor environments remains a key challenge in mobile robotics, due to difficulties in semantic reasoning, dynamic conditions, and long-term stability. We propose CausalNav, the first scene graph-based semantic navigation framework tailored for dynamic outdoor environments. We...

💬 0 commentsarXiv:2601.01872v1PDF
0

Posted in cs.CV · 2026-01-05 · Wenyu Shao, Hongbo Liu, Yunchuan Ma, Ruili Wang

Entity-Guided Multi-Task Learning for Infrared and Visible Image Fusion

Existing text-driven infrared and visible image fusion approaches often rely on textual information at the sentence level, which can lead to semantic noise from redundant text and fail to fully exploit the deeper semantic value of textual information. To address these issues, we propose a novel fusion approach named Entity-Guided...

💬 0 commentsarXiv:2601.01870v1PDF
0

Posted in cs.DS · 2026-01-05 · Yi Zhou, Haoyu Jiang, Chenghao Zhu, André Rossi

Exact Clique Number Manipulation via Edge Interdiction

The Edge Interdiction Clique Problem (EICP) aims to remove at most $k$ edges from a graph so as to minimize the size of the largest clique in the remaining graph. This problem captures a fundamental question in graph manipulation: which edges are structurally critical for preserving large cliques? Such a problem is also motivated by...

💬 0 commentsarXiv:2601.01869v1PDF