Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through September 22, 2026 — 02:40:15 EST

0

Posted in cs.CV · 2026-01-21 · Cuong Tran Van, Trong-Thang Pham, Ngoc-Son Nguyen, Duy Minh Ho Nguyen, Ngan Le

DuFal: Dual-Frequency-Aware Learning for High-Fidelity Extremely Sparse-view CBCT Reconstruction

Sparse-view Cone-Beam Computed Tomography reconstruction from limited X-ray projections remains a challenging problem in medical imaging due to the inherent undersampling of fine-grained anatomical details, which correspond to high-frequency components. Conventional CNN-based methods often struggle to recover these fine structures, as...

💬 0 commentsarXiv:2601.15416v2PDF
0

Posted in cs.HC · 2026-01-21 · Shreya Haran, Samiha Thatikonda, Dong Whi Yoo, Koustuv Saha

A Checklist for Trustworthy, Safe, and User-Friendly Mental Health Chatbots

Mental health concerns are rising globally, prompting increased reliance on technology to address the demand-supply gap in mental health services. In particular, mental health chatbots are emerging as a promising solution, but these remain largely untested, raising concerns about safety and potential harms. In this paper, we dive into...

💬 0 commentsarXiv:2601.15412v1PDF
0

Posted in cs.CV · 2026-01-21 · Pablo Messina, Andrés Villa, Juan León Alcázar, Karen Sánchez, Carlos Hinojosa, Denis Parra, Álvaro Soto, Bernard Ghanem

CURE: Curriculum-guided Multi-task Training for Reliable Anatomy Grounded Report Generation

Medical vision-language models can automate the generation of radiology reports but struggle with accurate visual grounding and factual consistency. Existing models often misalign textual findings with visual evidence, leading to unreliable or weakly grounded predictions. We present CURE, an error-aware curriculum learning framework...

💬 0 commentsarXiv:2601.15408v2PDF
0

Posted in cs.CV · 2026-01-21 · Hatef Otroshi Shahreza, Anjith George, Sébastien Marcel

Evaluating Multimodal Large Language Models for Heterogeneous Face Recognition

Multimodal Large Language Models (MLLMs) have recently demonstrated strong performance on a wide range of vision-language tasks, raising interest in their potential use for biometric applications. In this paper, we conduct a systematic evaluation of state-of-the-art MLLMs for heterogeneous face recognition (HFR), where enrollment and...

💬 0 commentsarXiv:2601.15406v1PDF
0

Posted in cs.IT · 2026-01-21 · Arman Fazeli, Mohammad M. Mansour, Ziyuan Zhu, Louay Jalloul

Partially Polarized Polar Codes: A New Design for 6G Control Channels

We introduce a new family of polar-like codes, called Partially Polarized Polar (PPP) codes. PPP codes are constructed from conventional polar codes by selectively pruning polarization kernels, thereby modifying the synthesized bit-channel capacities to ensure a guaranteed number of non-frozen bits available early in decoding. These...

💬 0 commentsarXiv:2601.15404v1PDF
0

Posted in cs.CR · 2026-01-21 · Sajjad Akherati, Xinmiao Zhang

Multi-Input Ciphertext Multiplication for Homomorphic Encryption

Homomorphic encryption (HE) enables arithmetic operations to be performed directly on encrypted data. It is essential for privacy-preserving applications such as machine learning, medical diagnosis, and financial data analysis. In popular HE schemes, ciphertext multiplication is only defined for two inputs. However, the multiplication...

💬 0 commentsarXiv:2601.15401v1PDF
0

Posted in cs.LG · 2026-01-21 · Ashna Nawar Ahmed, Banooqa Banday, Terry Jones, Tanzima Z. Islam

Attention-Informed Surrogates for Navigating Power-Performance Trade-offs in HPC

High-Performance Computing (HPC) schedulers must balance user performance with facility-wide resource constraints. The task boils down to selecting the optimal number of nodes for a given job. We present a surrogate-assisted multi-objective Bayesian optimization (MOBO) framework to automate this complex decision. Our core hypothesis...

💬 0 commentsarXiv:2601.15399v1PDF
0

Posted in cs.AI · 2026-01-21 · Peidong Wang

Beyond Prompting: Efficient and Robust Contextual Biasing for Speech LLMs via Logit-Space Integration (LOGIC)

The rapid emergence of new entities -- driven by cultural shifts, evolving trends, and personalized user data -- poses a significant challenge for existing Speech Large Language Models (Speech LLMs). While these models excel at general conversational tasks, their static training knowledge limits their ability to recognize...

💬 0 commentsarXiv:2601.15397v2PDF
0

Posted in cs.CL · 2026-01-21 · Tamunotonye Harry, Ivoline Ngong, Chima Nweke, Yuanyuan Feng, Joseph Near

Beyond Fixed Psychological Personas: State Beats Trait, but Language Models are State-Blind

User interactions with language models vary due to static properties of the user (trait) and the specific context of the interaction (state). However, existing persona datasets (like PersonaChat, PANDORA etc.) capture only trait, and ignore the impact of state. We introduce Chameleon, a dataset of 5,001 contextual psychological...

💬 0 commentsarXiv:2601.15395v2PDF
0

Posted in cs.CL · 2026-01-21 · Jaydeep Borkar, Karan Chadha, Niloofar Mireshghallah, Yuchen Zhang, Irina-Elena Veliche, Archi Mitra, David A. Smith, Zheng Xu, Diego Garcia-Olano

Memorization Dynamics in Knowledge Distillation for Language Models

Knowledge Distillation (KD) is increasingly adopted to transfer capabilities from large language models to smaller ones, offering significant improvements in efficiency and utility while often surpassing standard fine-tuning. Beyond performance, KD is also explored as a privacy-preserving mechanism to mitigate the risk of training...

💬 0 commentsarXiv:2601.15394v1PDF
0

Posted in cs.AI · 2026-01-21 · Francesca Pia Panaccione, Carlo Sgaravatti, Pietro Pinoli

GeMM-GAN: A Multimodal Generative Model Conditioned on Histopathology Images and Clinical Descriptions for Gene Expression Profile Generation

Biomedical research increasingly relies on integrating diverse data modalities, including gene expression profiles, medical images, and clinical metadata. While medical images and clinical metadata are routinely collected in clinical practice, gene expression data presents unique challenges for widespread research use, mainly due to...

💬 0 commentsarXiv:2601.15392v1PDF
0

Posted in cs.LG · 2026-01-21 · Zhaolong Su, Leheng Zhao, Xiaoying Wu, Ziyue Xu, Jindong Wang

FedUMM: A General Framework for Federated Learning with Unified Multimodal Models

Unified multimodal models (UMMs) are emerging as strong foundation models that can do both generation and understanding tasks in a single architecture. However, they are typically trained in centralized settings where all training and downstream datasets are gathered in a central server, limiting the deployment in privacy-sensitive...

💬 0 commentsarXiv:2601.15390v1PDF
0

Posted in cs.HC · 2026-01-21 · Marko Hostnik, Rauf Kurbanov, Yaroslav Sokolov, Artem Trofimov

VegaChat: A Robust Framework for LLM-Based Chart Generation and Assessment

Natural-language-to-visualization (NL2VIS) systems based on large language models (LLMs) have substantially improved the accessibility of data visualization. However, their further adoption is hindered by two coupled challenges: (i) the absence of standardized evaluation metrics makes it difficult to assess progress in the field and...

💬 0 commentsarXiv:2601.15385v1PDF
0

Posted in cs.LG · 2026-01-21 · Elon Litman, Gabe Guo

You Need Better Attention Priors

We generalize the attention mechanism by viewing it through the lens of Entropic Optimal Transport, revealing that standard attention corresponds to a transport problem regularized by an implicit uniform prior. We introduce Generalized Optimal transport Attention with Trainable priors (GOAT), a new attention mechanism that replaces...

💬 0 commentsarXiv:2601.15380v1PDF
0

Posted in cs.CV · 2026-01-21 · Jiwon Kang, Yeji Choi, JoungBin Lee, Wooseok Jang, Jinhyeok Choi, Taekeun Kang, Yongjae Park, Myungin Kim, Seungryong Kim

APPLE: Attribute-Preserving Pseudo-Labeling for Diffusion-Based Face Swapping

Face swapping aims to transfer the identity of a source face onto a target face while preserving target-specific attributes such as pose, expression, lighting, skin tone, and makeup. However, since real ground truth for face swapping is unavailable, achieving both accurate identity transfer and high-quality attribute preservation...

💬 0 commentsarXiv:2601.15288v2PDF
0

Posted in cs.CV · 2026-01-21 · Gautom Das, Vincent La, Ethan Lau, Abhinav Shrivastava, Matthew Gwilliam

Towards Understanding Best Practices for Quantization of Vision-Language Models

Large language models (LLMs) deliver impressive results for a variety of tasks, but state-of-the-art systems require fast GPUs with large amounts of memory. To reduce both the memory and latency of these systems, practitioners quantize their learned parameters, typically at half precision. A growing body of research focuses on...

💬 0 commentsarXiv:2601.15287v1PDF
0

Posted in cs.CV · 2026-01-21 · Shantanu Jaiswal, Mihir Prabhudesai, Nikash Bhardwaj, Zheyang Qin, Amir Zadeh, Chuan Li, Katerina Fragkiadaki, Deepak Pathak

Iterative Refinement Improves Compositional Image Generation

Text-to-image (T2I) models have achieved remarkable progress, yet they continue to struggle with complex prompts that require simultaneously handling multiple objects, relations, and attributes. Existing inference-time strategies, such as parallel sampling with verifiers or simply increasing denoising steps, can improve prompt...

💬 0 commentsarXiv:2601.15286v1PDF
0

Posted in cs.CV · 2026-01-21 · Anurag Bagchi, Zhipeng Bao, Homanga Bharadhwaj, Yu-Xiong Wang, Pavel Tokmakov, Martial Hebert

Walk through Paintings: Egocentric World Models from Internet Priors

What if a video generation model could not only imagine a plausible future, but the correct one, accurately reflecting how the world changes with each action? We address this question by presenting the Egocentric World Model (EgoWM), a simple, architecture-agnostic method that transforms any pretrained video diffusion model into an...

💬 0 commentsarXiv:2601.15284v1PDF
0

Posted in cs.CV · 2026-01-21 · Ruofan Liang, Norman Müller, Ethan Weber, Duncan Zauss, Nandita Vijaykumar, Peter Kontschieder, Christian Richardt

LuxRemix: Lighting Decomposition and Remixing for Indoor Scenes

We present a novel approach for interactive light editing in indoor scenes from a single multi-view scene capture. Our method leverages a generative image-based light decomposition model that factorizes complex indoor scene illumination into its constituent light sources. This factorization enables independent manipulation of...

💬 0 commentsarXiv:2601.15283v2PDF
0

Posted in cs.CV · 2026-01-21 · Yufan Deng, Zilin Pan, Hongyu Zhang, Xiaojie Li, Ruoqing Hu, Yufei Ding, Yiming Zou, Yan Zeng, Daquan Zhou

Rethinking Video Generation Model for the Embodied World

Video generation models have significantly advanced embodied intelligence, unlocking new possibilities for generating diverse robot data that capture perception, reasoning, and action in the physical world. However, synthesizing high-quality videos that accurately reflect real-world robotic interactions remains challenging, and the...

💬 0 commentsarXiv:2601.15282v1PDF
0

Posted in cs.CV · 2026-01-21 · Ying Yang, Zhengyao Lv, Tianlin Pan, Haofan Wang, Binxin Yang, Hubery Yin, Chen Li, Ziwei Liu, Chenyang Si

StableWorld: Towards Stable and Consistent Long Interactive Video Generation

In this paper, we explore the overlooked challenge of stability and temporal consistency in interactive video generation, which synthesizes dynamic and controllable video worlds through interactive behaviors such as camera movements and text prompts. Despite remarkable progress in world modeling, current methods still suffer from...

💬 0 commentsarXiv:2601.15281v1PDF
0

Posted in cs.HC · 2026-01-21 · Chloe Qianhui Zhao, Jie Cao, Jionghao Lin, Kenneth R. Koedinger

LLM-based Multimodal Feedback Produces Equivalent Learning and Better Student Perceptions than Educator Feedback

Providing timely, targeted, and multimodal feedback helps students quickly correct errors, build deep understanding and stay motivated, yet making it at scale remains a challenge. This study introduces a real-time AI-facilitated multimodal feedback system that integrates structured textual explanations with dynamic multimedia...

💬 0 commentsarXiv:2601.15280v1PDF
0

Posted in cs.LG · 2026-01-21 · Christoph Bartmann, Johannes Schimunek, Mykyta Ielanskyi, Philipp Seidl, Günter Klambauer, Sohvi Luukkonen

MolecularIQ: Characterizing Chemical Reasoning Capabilities Through Symbolic Verification on Molecular Graphs

A molecule's properties are fundamentally determined by its composition and structure encoded in its molecular graph. Thus, reasoning about molecular properties requires the ability to parse and understand the molecular graph. Large Language Models (LLMs) are increasingly applied to chemistry, tackling tasks such as molecular name...

💬 0 commentsarXiv:2601.15279v1PDF
0

Posted in cs.MM · 2026-01-21 · Mingyue Zha, Ho-Chun Herbert Chang

Interpreting Multimodal Communication at Scale in Short-Form Video: Visual, Audio, and Textual Mental Health Discourse on TikTok

Short-form video platforms integrate text, visuals, and audio into complex communicative acts, yet existing research analyzes these modalities in isolation, lacking scalable frameworks to interpret their joint contributions. This study introduces a pipeline combining automated multimodal feature extraction with Shapley value-based...

💬 0 commentsarXiv:2601.15278v1PDF
0

Posted in cs.CL · 2026-01-21 · Sahar Tahmasebi, Eric Müller-Budack, Ralph Ewerth

Robust Fake News Detection using Large Language Models under Adversarial Sentiment Attacks

Misinformation and fake news have become a pressing societal challenge, driving the need for reliable automated detection methods. Prior research has highlighted sentiment as an important signal in fake news detection, either by analyzing which sentiments are associated with fake news or by using sentiment and emotion features for...

💬 0 commentsarXiv:2601.15277v1PDF