Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through July 20, 2026 — 05:42:08 EST

0

Posted in cs.CV · 2026-01-21 · Qingling Shu, Sibao Chen, Wei Lu, Zhihui You, Chengzhuang Liu

UniRoute: Unified Routing Mixture-of-Experts for Modality-Adaptive Remote Sensing Change Detection

Current remote sensing change detection (CD) methods mainly rely on specialized models, which limits the scalability toward modality-adaptive Earth observation. For homogeneous CD, precise boundary delineation relies on fine-grained spatial cues and local pixel interactions, whereas heterogeneous CD instead requires broader contextual...

💬 0 commentsarXiv:2601.14797v1PDF
0

Posted in cs.SI · 2026-01-21 · Naomi Sasaya, Shigefumi Kishida, Ryo Kikuchi, Akira Tajima

Validating Behavioral Proxies for Disease Risk Monitoring via Large-Scale E-commerce Data

Digital traces of daily activities, such as e-commerce (EC) purchase histories, provide scalable signals for public health surveillance, yet their epidemiological validity remains unclear. This study validates a behavioral proxy for disease onset, defined as transitions from regular to therapeutic diets, by comparing large-scale EC...

💬 0 commentsarXiv:2601.14795v2PDF
0

Posted in cs.LG · 2026-01-21 · Dong Sun, Rahul Nittala, Rebekka Burkholz

Robustness of Mixtures of Experts to Feature Noise

Despite their practical success, it remains unclear why Mixture of Experts (MoE) models can outperform dense networks beyond sheer parameter scaling. We study an iso-parameter regime where inputs exhibit latent modular structure but are corrupted by feature noise, a proxy for noisy internal activations. We show that sparse expert...

💬 0 commentsarXiv:2601.14792v2PDF
0

Posted in cs.CV · 2026-01-21 · Ziyao Ling, Silvia Mirri, Paola Salomoni, Giovanni Delnevo

Synthetic Data Augmentation for Multi-Task Chinese Porcelain Classification: A Stable Diffusion Approach

The scarcity of training data presents a fundamental challenge in applying deep learning to archaeological artifact classification, particularly for the rare types of Chinese porcelain. This study investigates whether synthetic images generated through Stable Diffusion with Low-Rank Adaptation (LoRA) can effectively augment limited...

💬 0 commentsarXiv:2601.14791v1PDF
0

Posted in cs.AI · 2026-01-21 · Zhi Qiu, Jiazheng Sun, Chenxiao Xia, Jun Zheng, Xin Peng

CI4A: Semantic Component Interfaces for Agents Empowering Web Automation

While Large Language Models demonstrate remarkable proficiency in high-level semantic planning, they remain limited in handling fine-grained, low-level web component manipulations. To address this limitation, extensive research has focused on enhancing model grounding capabilities through techniques such as Reinforcement Learning....

💬 0 commentsarXiv:2601.14790v1PDF
0

Posted in cs.CV · 2026-01-21 · Yifei Liu, Changxing Ding, Ling Guo, Huaiguang Jiang, Qiong Cao

Reconstruction-Anchored Diffusion Model for Text-to-Motion Generation

Diffusion models have seen widespread adoption for text-driven human motion generation and related tasks due to their impressive generative capabilities and flexibility. However, current motion diffusion models face two major limitations: a representational gap caused by pre-trained text encoders that lack motion-specific information,...

💬 0 commentsarXiv:2601.14788v2PDF
0

Posted in cs.CV · 2026-01-21 · Hongjun An, Yiliang Song, Jiawei Shao, Zhe Sun, Xuelong Li

Single-Pixel Vision-Language Model for Intrinsic Privacy-Preserving Behavioral Intelligence

Adverse social interactions, such as bullying, harassment, and other illicit activities, pose significant threats to individual well-being and public safety, leaving profound impacts on physical and mental health. However, these critical events frequently occur in privacy-sensitive environments like restrooms, and changing rooms,...

💬 0 commentsarXiv:2601.17050v1PDF
0

Posted in cs.SD · 2026-01-21 · Wei-Jaw Lee, Fang-Chih Hsieh, Xuanjun Chen, Fang-Duo Tsai, Yi-Hsuan Yang

Training-Efficient Text-to-Music Generation with State-Space Modeling

Recent advances in text-to-music generation (TTM) have yielded high-quality results, but often at the cost of extensive compute and the use of large proprietary internal data. To improve the affordability and openness of TTM training, an open-source generative model backbone that is more training- and data-efficient is needed. In this...

💬 0 commentsarXiv:2601.14786v1PDF
0

Posted in cs.AI · 2026-01-21 · Amaury Guichard, Laurent Michel, Hélène Verhaeghe, Pierre Schaus

Towards Bound Consistency for the No-Overlap Constraint Using MDDs

Achieving bound consistency for the no-overlap constraint is known to be NP-complete. Therefore, several polynomial-time tightening techniques, such as edge finding, not-first-not-last reasoning, and energetic reasoning, have been introduced for this constraint. In this work, we derive the first bound-consistent algorithm for the...

💬 0 commentsarXiv:2601.14784v1PDF
0

Posted in cs.CL · 2026-01-21 · Anqi Li, Yuqian Chen, Yu Lu, Zhaoming Chen, Yuan Xie, Zhenzhong Lan

RECAP: Resistance Capture in Text-based Mental Health Counseling with Large Language Models

Recognizing and navigating client resistance is critical for effective mental health counseling, yet detecting such behaviors is particularly challenging in text-based interactions. Existing NLP approaches oversimplify resistance categories, ignore the sequential dynamics of therapeutic interventions, and offer limited...

💬 0 commentsarXiv:2601.14780v1PDF
0

Posted in cs.CR · 2026-01-21 · Yuang Qi, Na Zhao, Qiyi Yao, Benlong Wu, Weiming Zhang, Nenghai Yu, Kejiang Chen

STEAD: Robust Provably Secure Linguistic Steganography with Diffusion Language Model

Recent provably secure linguistic steganography (PSLS) methods rely on mainstream autoregressive language models (ARMs) to address historically challenging tasks, that is, to disguise covert communication as ``innocuous'' natural language communication. However, due to the characteristic of sequential generation of ARMs, the stegotext...

💬 0 commentsarXiv:2601.14778v1PDF
0

Posted in cs.CV · 2026-01-21 · Jiaxuan Liu, Yang Xiang, Han Zhao, Xiangang Li, Zhenhua Ling

FunCineForge: A Unified Dataset Toolkit and Model for Zero-Shot Movie Dubbing in Diverse Cinematic Scenes

Movie dubbing is the task of synthesizing speech from scripts conditioned on video scenes, requiring accurate lip sync, faithful timbre transfer, and proper modeling of character identity and emotion. However, existing methods face two major limitations: (1) high-quality multimodal dubbing datasets are limited in scale, suffer from...

💬 0 commentsarXiv:2601.14777v1PDF
0

Posted in cs.CV · 2026-01-21 · Xiaofan Yang, Yubin Liu, Wei Pan, Guoqing Chu, Junming Zhang, Jie Zhao, Zhuoqi Man, Xuanming Cao

M2I2HA: Multi-modal Object Detection Based on Intra- and Inter-Modal Hypergraph Attention

Recent advances in multi-modal detection have significantly improved detection accuracy in challenging environments (e.g., low light, overexposure). By integrating RGB with modalities such as thermal and depth, multi-modal fusion increases data redundancy and system robustness. However, significant challenges remain in effectively...

💬 0 commentsarXiv:2601.14776v3PDF
0

Posted in cs.CV · 2026-01-21 · Keita Takeda, Tomoya Sakai

Does medical specialization of VLMs enhance discriminative power?: A comprehensive investigation through feature distribution analysis

This study investigates the feature representations produced by publicly available open source medical vision-language models (VLMs). While medical VLMs are expected to capture diagnostically relevant features, their learned representations remain underexplored, and standard evaluations like classification accuracy do not fully reveal...

💬 0 commentsarXiv:2601.14774v1PDF
0

Posted in cs.AI · 2026-01-21 · Haizhou Liu, Haodong Jin, Yiming Wang, Hui Yu

Semantic-Guided Unsupervised Video Summarization

Video summarization is a crucial technique for social understanding, enabling efficient browsing of massive multimedia content and extraction of key information from social platforms. Most existing unsupervised summarization methods rely on Generative Adversarial Networks (GANs) to enhance keyframe selection and generate coherent,...

💬 0 commentsarXiv:2601.14773v1PDF
0

Posted in cs.CV · 2026-01-21 · Puneet Sharma, Kristian Dalsbø Hindberg, Eibe Frank, Benedicte Schelde-Olesen, Ulrik Deding

Using Multi-Instance Learning to Identify Unique Polyps in Colon Capsule Endoscopy Images

Identifying unique polyps in colon capsule endoscopy (CCE) images is a critical yet challenging task for medical personnel due to the large volume of images, the cognitive load it creates for clinicians, and the ambiguity in labeling specific frames. This paper formulates this problem as a multi-instance learning (MIL) task, where a...

💬 0 commentsarXiv:2601.14771v1PDF
0

Posted in cs.GR · 2026-01-21 · Chun Chen, Minseok Chae, Seung-Woo Nam, Myeong-Ho Choi, Minseong Kim, Eunbi Lee, Yoonchan Jeong, Jae-Hyeung Park

PAColorHolo: A Perceptually-Aware Color Management Framework for Holographic Displays

Holographic displays offer significant potential for augmented and virtual reality applications by reconstructing wavefronts that enable continuous depth cues and natural parallax without vergence-accommodation conflict. However, despite advances in pixel-level image quality, current systems struggle to achieve perceptually accurate...

💬 0 commentsarXiv:2601.14766v1PDF
0

Posted in cs.LG · 2026-01-21 · Harold Kiossou, Pierre Schaus, Siegfried Nijssen

Anytime Optimal Decision Tree Learning with Continuous Features

In recent years, significant progress has been made on algorithms for learning optimal decision trees, primarily in the context of binary features. Extending these methods to continuous features remains substantially more challenging due to the large number of potential splits for each feature. Recently, an elegant exact algorithm was...

💬 0 commentsarXiv:2601.14765v1PDF
0

Posted in cs.AI · 2026-01-21 · Thomas Eiter, Tobias Geibinger, Zeynep G. Saribatur

An XAI View on Explainable ASP: Methods, Systems, and Perspectives

Answer Set Programming (ASP) is a popular declarative reasoning and problem solving approach in symbolic AI. Its rule-based formalism makes it inherently attractive for explainable and interpretive reasoning, which is gaining importance with the surge of Explainable AI (XAI). A number of explanation approaches and tools for ASP have...

💬 0 commentsarXiv:2601.14764v2PDF
0

Posted in cs.LG · 2026-01-21 · Injin Kong, Hyoungjoon Lee, Yohan Jo

Mechanism Shift During Post-training from Autoregressive to Masked Diffusion Language Models

Post-training pretrained autoregressive models (ARMs) into masked diffusion models (MDMs) has emerged as a cost-effective way to overcome the limitations of sequential generation. Yet it remains unclear whether post-trained MDMs acquire genuinely new computational mechanisms or merely re-express autoregressive computation in a...

💬 0 commentsarXiv:2601.14758v4PDF
0

Posted in cs.CV · 2026-01-21 · Kangcheng Zhou, Jun Jiang, Qing Zhang, Shuang Zheng, Qingli Li, Shugong Xu

ReinPath: A Multimodal Reinforcement Learning Approach for Pathology

Interpretability is significant in computational pathology, leading to the development of multimodal information integration from histopathological image and corresponding text data.However, existing multimodal methods have limited interpretability due to the lack of high-quality dataset that support explicit reasoning and inference...

💬 0 commentsarXiv:2601.14757v1PDF
0

Posted in cs.IT · 2026-01-21 · Lei Xie, Peilan Wang, Guanxiong Shen, Guyue Li, Weidong Mei, Liquan Chen

Secure Communication in MIMOME Movable-Antenna Systems with Statistical Eavesdropper CSI

This paper investigates the potential of movable antennas (MAs) to enhance physical layer security within a multiple-input multiple-output multiple-antenna eavesdropper (MIMOME) system. We consider a practical scenario where the transmitter operates with imperfect eavesdropper channel state information (ECSI), knowing only the...

💬 0 commentsarXiv:2601.14755v1PDF
0

Posted in cs.DL · 2026-01-21 · Marilena Daquino, Francesca Mambelli, Artem Kozlov

Many-to-many. Usability challenges of entity reconciliation in art history and photographic studies

This article investigates challenges in reconciling heterogeneous records across cultural institutions, focusing on art historical photo archives within the PHAROS consortium. Through case studies, the study analyses reconciliation workflows and cataloguing traditions, with attention to institutional contexts, data granularities, and...

💬 0 commentsarXiv:2601.14753v1PDF
0

Posted in cs.CL · 2026-01-21 · Yifan Wang, Shiyu Li, Peiming Li, Xiaochen Yang, Yang Tang, Zheng Wei

Render-of-Thought: Rendering Textual Chain-of-Thought as Images for Visual Latent Reasoning

Chain-of-Thought (CoT) prompting has achieved remarkable success in unlocking the reasoning capabilities of Large Language Models (LLMs). Although CoT prompting enhances reasoning, its verbosity imposes substantial computational overhead. Recent works often focus exclusively on outcome alignment and lack supervision on the...

💬 0 commentsarXiv:2601.14750v4PDF
0

Posted in cs.LG · 2026-01-21 · Hongyue Wu, Hangyu Li, Guodong Fan, Haoran Zhu, Shizhan Chen, Zhiyong Feng

RefProtoFL: Communication-Efficient Federated Learning via External-Referenced Prototype Alignment

Federated learning (FL) enables collaborative model training without sharing raw data in edge environments, but is constrained by limited communication bandwidth and heterogeneous client data distributions. Prototype-based FL mitigates this issue by exchanging class-wise feature prototypes instead of full model parameters; however,...

💬 0 commentsarXiv:2601.14746v2PDF