Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through July 21, 2026 — 19:41:07 EST

0

Posted in cs.CL · 2026-01-21 · Tamunotonye Harry, Ivoline Ngong, Chima Nweke, Yuanyuan Feng, Joseph Near

Beyond Fixed Psychological Personas: State Beats Trait, but Language Models are State-Blind

User interactions with language models vary due to static properties of the user (trait) and the specific context of the interaction (state). However, existing persona datasets (like PersonaChat, PANDORA etc.) capture only trait, and ignore the impact of state. We introduce Chameleon, a dataset of 5,001 contextual psychological...

💬 0 commentsarXiv:2601.15395v2PDF
0

Posted in cs.CL · 2026-01-21 · Jaydeep Borkar, Karan Chadha, Niloofar Mireshghallah, Yuchen Zhang, Irina-Elena Veliche, Archi Mitra, David A. Smith, Zheng Xu, Diego Garcia-Olano

Memorization Dynamics in Knowledge Distillation for Language Models

Knowledge Distillation (KD) is increasingly adopted to transfer capabilities from large language models to smaller ones, offering significant improvements in efficiency and utility while often surpassing standard fine-tuning. Beyond performance, KD is also explored as a privacy-preserving mechanism to mitigate the risk of training...

💬 0 commentsarXiv:2601.15394v1PDF
0

Posted in cs.AI · 2026-01-21 · Francesca Pia Panaccione, Carlo Sgaravatti, Pietro Pinoli

GeMM-GAN: A Multimodal Generative Model Conditioned on Histopathology Images and Clinical Descriptions for Gene Expression Profile Generation

Biomedical research increasingly relies on integrating diverse data modalities, including gene expression profiles, medical images, and clinical metadata. While medical images and clinical metadata are routinely collected in clinical practice, gene expression data presents unique challenges for widespread research use, mainly due to...

💬 0 commentsarXiv:2601.15392v1PDF
0

Posted in cs.LG · 2026-01-21 · Zhaolong Su, Leheng Zhao, Xiaoying Wu, Ziyue Xu, Jindong Wang

FedUMM: A General Framework for Federated Learning with Unified Multimodal Models

Unified multimodal models (UMMs) are emerging as strong foundation models that can do both generation and understanding tasks in a single architecture. However, they are typically trained in centralized settings where all training and downstream datasets are gathered in a central server, limiting the deployment in privacy-sensitive...

💬 0 commentsarXiv:2601.15390v1PDF
0

Posted in cs.HC · 2026-01-21 · Marko Hostnik, Rauf Kurbanov, Yaroslav Sokolov, Artem Trofimov

VegaChat: A Robust Framework for LLM-Based Chart Generation and Assessment

Natural-language-to-visualization (NL2VIS) systems based on large language models (LLMs) have substantially improved the accessibility of data visualization. However, their further adoption is hindered by two coupled challenges: (i) the absence of standardized evaluation metrics makes it difficult to assess progress in the field and...

💬 0 commentsarXiv:2601.15385v1PDF
0

Posted in cs.LG · 2026-01-21 · Elon Litman, Gabe Guo

You Need Better Attention Priors

We generalize the attention mechanism by viewing it through the lens of Entropic Optimal Transport, revealing that standard attention corresponds to a transport problem regularized by an implicit uniform prior. We introduce Generalized Optimal transport Attention with Trainable priors (GOAT), a new attention mechanism that replaces...

💬 0 commentsarXiv:2601.15380v1PDF
0

Posted in cs.CV · 2026-01-21 · Jiwon Kang, Yeji Choi, JoungBin Lee, Wooseok Jang, Jinhyeok Choi, Taekeun Kang, Yongjae Park, Myungin Kim, Seungryong Kim

APPLE: Attribute-Preserving Pseudo-Labeling for Diffusion-Based Face Swapping

Face swapping aims to transfer the identity of a source face onto a target face while preserving target-specific attributes such as pose, expression, lighting, skin tone, and makeup. However, since real ground truth for face swapping is unavailable, achieving both accurate identity transfer and high-quality attribute preservation...

💬 0 commentsarXiv:2601.15288v2PDF
0

Posted in cs.CV · 2026-01-21 · Gautom Das, Vincent La, Ethan Lau, Abhinav Shrivastava, Matthew Gwilliam

Towards Understanding Best Practices for Quantization of Vision-Language Models

Large language models (LLMs) deliver impressive results for a variety of tasks, but state-of-the-art systems require fast GPUs with large amounts of memory. To reduce both the memory and latency of these systems, practitioners quantize their learned parameters, typically at half precision. A growing body of research focuses on...

💬 0 commentsarXiv:2601.15287v1PDF
0

Posted in cs.CV · 2026-01-21 · Shantanu Jaiswal, Mihir Prabhudesai, Nikash Bhardwaj, Zheyang Qin, Amir Zadeh, Chuan Li, Katerina Fragkiadaki, Deepak Pathak

Iterative Refinement Improves Compositional Image Generation

Text-to-image (T2I) models have achieved remarkable progress, yet they continue to struggle with complex prompts that require simultaneously handling multiple objects, relations, and attributes. Existing inference-time strategies, such as parallel sampling with verifiers or simply increasing denoising steps, can improve prompt...

💬 0 commentsarXiv:2601.15286v1PDF
0

Posted in cs.CV · 2026-01-21 · Anurag Bagchi, Zhipeng Bao, Homanga Bharadhwaj, Yu-Xiong Wang, Pavel Tokmakov, Martial Hebert

Walk through Paintings: Egocentric World Models from Internet Priors

What if a video generation model could not only imagine a plausible future, but the correct one, accurately reflecting how the world changes with each action? We address this question by presenting the Egocentric World Model (EgoWM), a simple, architecture-agnostic method that transforms any pretrained video diffusion model into an...

💬 0 commentsarXiv:2601.15284v1PDF
0

Posted in cs.CV · 2026-01-21 · Ruofan Liang, Norman Müller, Ethan Weber, Duncan Zauss, Nandita Vijaykumar, Peter Kontschieder, Christian Richardt

LuxRemix: Lighting Decomposition and Remixing for Indoor Scenes

We present a novel approach for interactive light editing in indoor scenes from a single multi-view scene capture. Our method leverages a generative image-based light decomposition model that factorizes complex indoor scene illumination into its constituent light sources. This factorization enables independent manipulation of...

💬 0 commentsarXiv:2601.15283v2PDF
0

Posted in cs.CV · 2026-01-21 · Yufan Deng, Zilin Pan, Hongyu Zhang, Xiaojie Li, Ruoqing Hu, Yufei Ding, Yiming Zou, Yan Zeng, Daquan Zhou

Rethinking Video Generation Model for the Embodied World

Video generation models have significantly advanced embodied intelligence, unlocking new possibilities for generating diverse robot data that capture perception, reasoning, and action in the physical world. However, synthesizing high-quality videos that accurately reflect real-world robotic interactions remains challenging, and the...

💬 0 commentsarXiv:2601.15282v1PDF
0

Posted in cs.CV · 2026-01-21 · Ying Yang, Zhengyao Lv, Tianlin Pan, Haofan Wang, Binxin Yang, Hubery Yin, Chen Li, Ziwei Liu, Chenyang Si

StableWorld: Towards Stable and Consistent Long Interactive Video Generation

In this paper, we explore the overlooked challenge of stability and temporal consistency in interactive video generation, which synthesizes dynamic and controllable video worlds through interactive behaviors such as camera movements and text prompts. Despite remarkable progress in world modeling, current methods still suffer from...

💬 0 commentsarXiv:2601.15281v1PDF
0

Posted in cs.HC · 2026-01-21 · Chloe Qianhui Zhao, Jie Cao, Jionghao Lin, Kenneth R. Koedinger

LLM-based Multimodal Feedback Produces Equivalent Learning and Better Student Perceptions than Educator Feedback

Providing timely, targeted, and multimodal feedback helps students quickly correct errors, build deep understanding and stay motivated, yet making it at scale remains a challenge. This study introduces a real-time AI-facilitated multimodal feedback system that integrates structured textual explanations with dynamic multimedia...

💬 0 commentsarXiv:2601.15280v1PDF
0

Posted in cs.LG · 2026-01-21 · Christoph Bartmann, Johannes Schimunek, Mykyta Ielanskyi, Philipp Seidl, Günter Klambauer, Sohvi Luukkonen

MolecularIQ: Characterizing Chemical Reasoning Capabilities Through Symbolic Verification on Molecular Graphs

A molecule's properties are fundamentally determined by its composition and structure encoded in its molecular graph. Thus, reasoning about molecular properties requires the ability to parse and understand the molecular graph. Large Language Models (LLMs) are increasingly applied to chemistry, tackling tasks such as molecular name...

💬 0 commentsarXiv:2601.15279v1PDF
0

Posted in cs.MM · 2026-01-21 · Mingyue Zha, Ho-Chun Herbert Chang

Interpreting Multimodal Communication at Scale in Short-Form Video: Visual, Audio, and Textual Mental Health Discourse on TikTok

Short-form video platforms integrate text, visuals, and audio into complex communicative acts, yet existing research analyzes these modalities in isolation, lacking scalable frameworks to interpret their joint contributions. This study introduces a pipeline combining automated multimodal feature extraction with Shapley value-based...

💬 0 commentsarXiv:2601.15278v1PDF
0

Posted in cs.CL · 2026-01-21 · Sahar Tahmasebi, Eric Müller-Budack, Ralph Ewerth

Robust Fake News Detection using Large Language Models under Adversarial Sentiment Attacks

Misinformation and fake news have become a pressing societal challenge, driving the need for reliable automated detection methods. Prior research has highlighted sentiment as an important signal in fake news detection, either by analyzing which sentiments are associated with fake news or by using sentiment and emotion features for...

💬 0 commentsarXiv:2601.15277v1PDF
0

Posted in cs.CV · 2026-01-21 · Yu Wu, Minsik Jeon, Jen-Hao Rick Chang, Oncel Tuzel, Shubham Tulsiani

RayRoPE: Projective Ray Positional Encoding for Multi-view Attention

We study positional encodings for multi-view transformers that process tokens from a set of posed input images, and seek a mechanism that encodes patches uniquely, allows SE(3)-invariant attention with multi-frequency similarity, and can adapt to the geometry of the underlying 3D scene. We find that prior (absolute or relative)...

💬 0 commentsarXiv:2601.15275v3PDF
0

Posted in cs.LG · 2026-01-21 · Maciej Kilian, Oleg Mkrtchyan, Luke Zettlemoyer, Akshat Shrivastava, Armen Aghajanyan

Improving MoE Compute Efficiency by Composing Weight and Data Sparsity

Mixture-of-Experts layers achieve compute efficiency through weight sparsity: each token activates only a subset of experts. Data sparsity, where each expert processes only a subset of tokens, offers a complementary axis. Expert-choice routing implements data sparsity directly but violates causality in autoregressive models, creating...

💬 0 commentsarXiv:2601.15370v1PDF
0

Posted in cs.RO · 2026-01-21 · Heng Zhang, Wei-Hsing Huang, Qiyi Tong, Gokhan Solak, Puze Liu, Kaidi Zhang, Sheng Liu, Jan Peters, Yu She, Arash Ajoudani

CompliantVLA-adaptor: VLM-Guided Variable Impedance Action for Safe Contact-Rich Manipulation

We propose a CompliantVLA-adaptor that augments the state-of-the-art Vision-Language-Action (VLA) models with vision-language model (VLM)-informed context-aware variable impedance control (VIC) to improve the safety and effectiveness of contact-rich robotic manipulation tasks. Existing VLA systems (e.g., RDT, Pi0.5, OpenVLA-oft)...

💬 0 commentsarXiv:2601.15541v2PDF
0

Posted in cs.LG · 2026-01-21 · Dongchen Huang

PRISM: Deriving a White-Box Transformer as a Signal-Noise Decomposition Operator via Maximum Coding Rate Reduction

Deep learning models, particularly Transformers, are often criticized as "black boxes" and lack interpretability. We propose Prism, a white-box attention-based architecture derived from the principles of Maximizing Coding Rate Reduction ($\text{MCR}^2$). By modeling the attention mechanism as a gradient ascent process on a distinct...

💬 0 commentsarXiv:2601.15540v2PDF
0

Posted in cs.LG · 2026-01-21 · Himanshu Mishra, Kanwal Mehreen

QUAIL: Quantization Aware Unlearning for Mitigating Misinformation in LLMs

Machine unlearning aims to remove specific knowledge (e.g., copyrighted or private data) from a trained model without full retraining. In practice, models are often quantized (e.g., 4-bit) for deployment, but we find that quantization can catastrophically restore forgotten information [1]. In this paper, we (1) analyze why low-bit...

💬 0 commentsarXiv:2601.15538v1PDF
0

Posted in cs.AI · 2026-01-21 · Zhikang Chen, Tingting Zhu

From Generative Engines to Actionable Simulators: The Imperative of Physical Grounding in World Models

A world model is an AI system that simulates how an environment evolves under actions, enabling planning through imagined futures rather than reactive perception. Current world models, however, suffer from visual conflation: the mistaken assumption that high-fidelity video generation implies an understanding of physical and causal...

💬 0 commentsarXiv:2601.15533v1PDF
0

Posted in cs.NI · 2026-01-21 · Abd Ullah Khan, Wali Ullah Khan, Haejoon Jung, Hyundong Shin

Resource Allocation and Sharing for UAV-Assisted Integrated TN-NTN with Multi-Connectivity

Unmanned aerial vehicles (UAVs) with multi-connectivity (MC) capabilities efficiently and reliably transfer data between terrestrial networks (TNs) and non-terrestrial networks (NTNs). However, optimally sharing and allocating spectrum and power resources to maintain MC while ensuring reliable connectivity and optimal performance...

💬 0 commentsarXiv:2601.15532v2PDF
0

Posted in cs.LG · 2026-01-21 · Megan A. Witherow, Michael L. Evans, Ahmed Temtam, Hamid R. Okhravi, Khan M. Iftekharuddin

Machine learning-enhanced non-amnestic Alzheimer's disease diagnosis from MRI and clinical features

Alzheimer's disease (AD), defined as an abnormal buildup of amyloid plaques and tau tangles in the brain can be diagnosed with high accuracy based on protein biomarkers via PET or CSF analysis. However, due to the invasive nature of biomarker collection, most AD diagnoses are made in memory clinics using cognitive tests and evaluation...

💬 0 commentsarXiv:2601.15530v2PDF