Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through July 20, 2026 — 17:22:47 EST

0

Posted in cs.CV · 2026-01-10 · Kai Cheng, Ruoqi Wang, Qiong Luo

VVTRec: Radio Interferometric Reconstruction through Visual and Textual Modality Enrichment

Radio astronomy is an indispensable discipline for observing distant celestial objects. Measurements of wave signals from radio telescopes, called visibility, need to be transformed into images for astronomical observations. These dirty images blend information from real sources and artifacts. Therefore, astronomers usually perform...

💬 0 commentsarXiv:2601.06475v1PDF
0

Posted in cs.CV · 2026-01-10 · Chenxu Dang, Jie Wang, Guang Li, Zhiwen Hou, Zihan You, Hangjun Ye, Jie Ma, Long Chen, Yan Wang

SparseOccVLA: Bridging Occupancy and Vision-Language Models via Sparse Queries for Unified 4D Scene Understanding and Planning

In autonomous driving, Vision Language Models (VLMs) excel at high-level reasoning , whereas semantic occupancy provides fine-grained details. Despite significant progress in individual fields, there is still no method that can effectively integrate both paradigms. Conventional VLMs struggle with token explosion and limited...

💬 0 commentsarXiv:2601.06474v2PDF
0

Posted in cs.MA · 2026-01-10 · Sathish Sampath, Anuradha Baskaran

Adaptive Orchestration: Scalable Self-Evolving Multi-Agent Systems

As Large Language Models (LLMs) are increasingly deployed as autonomous agents, they face a critical scalability bottleneck known as the "Generalization-Specialization Dilemma." Monolithic agents equipped with extensive toolkits suffer from context pollution and attention decay, leading to hallucinations. Conversely, static...

💬 0 commentsarXiv:2601.09742v1PDF
0

Posted in cs.LG · 2026-01-10 · Chutian Huang, Chang Ma, Kaibo Wang, Yang Xiang

StablePDENet: Enhancing Stability of Operator Learning for Solving Differential Equations

Learning solution operators for differential equations with neural networks has shown great potential in scientific computing, but ensuring their stability under input perturbations remains a critical challenge. This paper presents a robust self-supervised neural operator framework that enhances stability through adversarial training...

💬 0 commentsarXiv:2601.06472v1PDF
0

Posted in cs.CL · 2026-01-10 · Junho Park, Dohoon Kim, Taesup Moon

PRISP: Privacy-Safe Few-Shot Personalization via Lightweight Adaptation

Large language model (LLM) personalization aims to adapt general-purpose models to individual users. Most existing methods, however, are developed under data-rich and resource-abundant settings, often incurring privacy risks. In contrast, realistic personalization typically occurs after deployment under (i) extremely limited user...

💬 0 commentsarXiv:2601.06471v1PDF
0

Posted in cs.CE · 2026-01-10 · Weipeng Xu, Ziyuan Xie, Haoju Lin, Xinyu Wang, Guangjin Mou, Tianju Xue

Style-constrained inverse design of microstructures with tailored mechanical properties using unconditional diffusion models

Deep generative models, particularly denoising diffusion models, have achieved remarkable success in high-fidelity generation of architected microstructures with desired properties and styles. Nevertheless, these recent methods typically rely on conditional training mechanisms and demand substantial computational effort to prepare the...

💬 0 commentsarXiv:2601.06469v1PDF
0

Posted in cs.LG · 2026-01-10 · Anh-Tuan Mai, Cam-Van Thi Nguyen, Duc-Trong Le

Divide and Refine: Enhancing Multimodal Representation and Explainability for Emotion Recognition in Conversation

Multimodal emotion recognition in conversation (MERC) requires representations that effectively integrate signals from multiple modalities. These signals include modality-specific cues, information shared across modalities, and interactions that emerge only when modalities are combined. In information-theoretic terms, these correspond...

💬 0 commentsarXiv:2601.14274v1PDF
0

Posted in cs.CR · 2026-01-10 · Imtiaz Ali Soomro, Hamood Ur Rehman, S. Jawad Hussain ID, Adeel Iqbal, Waqas Khalid, Heejung Yu ID

SecureDyn-FL: A Robust Privacy-Preserving Federated Learning Framework for Intrusion Detection in IoT Networks

The rapid proliferation of Internet of Things (IoT) devices across domains such as smart homes, industrial control systems, and healthcare networks has significantly expanded the attack surface for cyber threats, including botnet-driven distributed denial-of-service (DDoS), malware injection, and data exfiltration. Conventional...

💬 0 commentsarXiv:2601.06466v1PDF
0

Posted in cs.CV · 2026-01-10 · Chao Liu, Ngai-Man Cheung

On the Adversarial Robustness of 3D Large Vision-Language Models

3D Vision-Language Models (VLMs), such as PointLLM and GPT4Point, have shown strong reasoning and generalization abilities in 3D understanding tasks. However, their adversarial robustness remains largely unexplored. Prior work in 2D VLMs has shown that the integration of visual inputs significantly increases vulnerability to...

💬 0 commentsarXiv:2601.06464v1PDF
0

Posted in cs.LG · 2026-01-10 · Xuezhe Ma, Shicheng Wen, Linghao Jin, Bilge Acun, Ruihang Lai, Bohan Hou, Will Lin, Hao Zhang, Songlin Yang, Ryan Lee, Mengxi Wu, Jonathan May, Luke Zettlemoyer, Carole-Jean Wu

Gecko: An Efficient Neural Architecture Inherently Processing Sequences with Arbitrary Lengths

Designing a unified neural network to efficiently and inherently process sequential data with arbitrary lengths is a central and challenging problem in sequence modeling. The design choices in Transformer, including quadratic complexity and weak length extrapolation, have limited their ability to scale to long sequences. In this work,...

💬 0 commentsarXiv:2601.06463v1PDF
0

Posted in cs.CR · 2026-01-10 · Minfeng Qi, Dongyang He, Qin Wang, Lefeng Zhang

VIPER Strike: Defeating Visual Reasoning CAPTCHAs via Structured Vision-Language Inference

Visual Reasoning CAPTCHAs (VRCs) combine visual scenes with natural-language queries that demand compositional inference over objects, attributes, and spatial relations. They are increasingly deployed as a primary defense against automated bots. Existing solvers fall into two paradigms: vision-centric, which rely on template-specific...

💬 0 commentsarXiv:2601.06461v1PDF
0

Posted in cs.CV · 2026-01-10 · Weihao Hong, Zhiyuan Jiang, Bingyu Shen, Xinlei Guan, Yangyi Feng, Meng Xu, Boyang Li

Tone Matters: The Impact of Linguistic Tone on Hallucination in VLMs

Vision-Language Models (VLMs) are increasingly used in safety-critical applications that require reliable visual grounding. However, these models often hallucinate details that are not present in the image to satisfy user prompts. While recent datasets and benchmarks have been introduced to evaluate systematic hallucinations in VLMs,...

💬 0 commentsarXiv:2601.06460v1PDF
0

Posted in cs.IR · 2026-01-10 · Sayak Chakrabarty, Souradip Pal

PixRec: Leveraging Visual Context for Next-Item Prediction in Sequential Recommendation

Large Language Models (LLMs) have recently shown strong potential for usage in sequential recommendation tasks through text-only models, which combine advanced prompt design, contrastive alignment, and fine-tuning on downstream domain-specific data. While effective, these approaches overlook the rich visual information present in many...

💬 0 commentsarXiv:2601.06458v1PDF
0

Posted in cs.SE · 2026-01-10 · Shaunak Biswas, Hiya Bhatt, Karthik Vaidhyanathan

Architecting AgentOps Needs CHANGE

The emergence of Agentic AI systems has outpaced the architectural thinking required to operate them effectively. These agents differ fundamentally from traditional software: their behavior is not fixed at deployment but continuously shaped by experience, feedback, and context. Applying operational principles inherited from DevOps or...

💬 0 commentsarXiv:2601.06456v1PDF
0

Posted in cs.AI · 2026-01-10 · Hyungjun Yoon, Mohammad Malekzadeh, Sung-Ju Lee, Fahim Kawsar, Lorena Qendro

ConSensus: Multi-Agent Collaboration for Multimodal Sensing

Large language models (LLMs) are increasingly grounded in sensor data to perceive and reason about human physiology and the physical world. However, accurately interpreting heterogeneous multimodal sensor data remains a fundamental challenge. We show that a single monolithic LLM often fails to reason coherently across modalities,...

💬 0 commentsarXiv:2601.06453v2PDF
0

Posted in cs.RO · 2026-01-10 · Hyunseo Koh, Chang-Yong Song, Youngjae Choi, Misa Viveiros, David Hyde, Heewon Kim

CulinaryCut-VLAP: A Vision-Language-Action-Physics Framework for Food Cutting via a Force-Aware Material Point Method

Food cutting is a highly practical yet underexplored application at the intersection of vision and robotic manipulation. The task remains challenging because interactions between the knife and deformable materials are highly nonlinear and often entail large deformations, frequent contact, and topological change, which in turn hinder...

💬 0 commentsarXiv:2601.06451v1PDF
0

Posted in cs.CY · 2026-01-10 · WariNkwi K. Flores, KunTikzi Flores, Rosa M. Panama, KayaKanti Alta

Kara-Kichwa Data Sovereignty Framework: Reference Point for Indigenous Data Authority Renaissances in LAC

For Indigenous Peoples of the Apya Yala (or Abya Yala), particularly in the Kara and Kichwa citizens of the Pan-Andean-Amazonian biocultural region, data is not merely a knowledge or information resource, it is the extension of Khipu Panaka (Indigenous data authority), treading the data lifecycle, genealogical and relational memory...

💬 0 commentsarXiv:2601.06634v2PDF
0

Posted in cs.LG · 2026-01-10 · Zhangqi Duan, Nigel Fernandez, Andrew Lan

KASER: Knowledge-Aligned Student Error Simulator for Open-Ended Coding Tasks

Open-ended tasks, such as coding problems that are common in computer science education, provide detailed insights into student knowledge. However, training large language models (LLMs) to simulate and predict possible student errors in their responses to these problems can be challenging: they often suffer from mode collapse and fail...

💬 0 commentsarXiv:2601.06633v2PDF
0

Posted in cs.CL · 2026-01-10 · Mohammed Fayiz Parappan, Ricardo Henao

Labels have Human Values: Value Calibration of Subjective Tasks

Building NLP systems for subjective tasks requires one to ensure their alignment to contrasting human values. We propose the MultiCalibrated Subjective Task Learner framework (MC-STL), which clusters annotations into identifiable human value clusters by three approaches (similarity of annotator rationales, expert-value taxonomies or...

💬 0 commentsarXiv:2601.06631v1PDF
0

Posted in cs.DS · 2026-01-10 · Luis Alberto Croquevielle, Roman Sokolovskii, Thomas Heinis

Lower Bounds for the Algorithmic Complexity of Learned Indexes

Learned index structures aim to accelerate queries by training machine learning models to approximate the rank function associated with a database attribute. While effective in practice, their theoretical limitations are not fully understood. We present a general framework for proving lower bounds on query time for learned indexes,...

💬 0 commentsarXiv:2601.06629v1PDF
0

Posted in cs.CR · 2026-01-10 · Qiang Zhang, Elena Emma Wang, Jiaming Li, Xichun Wang

Burn-After-Use for Preventing Data Leakage through a Secure Multi-Tenant Architecture in Enterprise LLM

This study presents a Secure Multi-Tenant Architecture (SMTA) combined with a novel concept Burn-After-Use (BAU) mechanism for enterprise LLM environments to effectively prevent data leakage. As institutions increasingly adopt LLMs across departments, the risks of data leakage have become a critical security and compliance concern....

💬 0 commentsarXiv:2601.06627v3PDF
0

Posted in cs.CY · 2026-01-10 · Ira Wolfson

Informed Consent for AI Consciousness Research: A Talmudic Framework for Graduated Protections

Artificial intelligence research faces a critical ethical paradox: determining whether AI systems are conscious requires experiments that may harm entities whose moral status remains uncertain. Recent work proposes avoiding consciousness-uncertain AI systems entirely, yet this faces practical limitations-we cannot guarantee such...

💬 0 commentsarXiv:2601.08864v1PDF
0

Posted in cs.CL · 2026-01-10 · Marco Martinelli, Stefano Marchesin, Gianmaria Silvello

Efficient and Reliable Estimation of Named Entity Linking Quality: A Case Study on GutBrainIE

Named Entity Linking (NEL) is a core component of biomedical Information Extraction (IE) pipelines, yet assessing its quality at scale is challenging due to the high cost of expert annotations and the large size of corpora. In this paper, we present a sampling-based framework to estimate the NEL accuracy of large-scale IE corpora...

💬 0 commentsarXiv:2601.06624v1PDF
0

Posted in cs.RO · 2026-01-10 · Giovani Braglia, José Jair Alves Mendes Junior, Augusto Tetsuo Prado Inafuco, Federico Mariano, Leonardo S. Mattos

Robotic Tele-Operation for Upper Aerodigestive Tract Microsurgery: System Design and Validation

Upper aerodigestive tract (UADT) treatments frequently employ transoral laser microsurgery (TLM) for procedures such as the removal of tumors or polyps. In TLM, a laser beam is used to cut target tissue, while forceps are employed to grasp, manipulate, and stabilize tissue within the UADT. Although TLM systems may rely on different...

💬 0 commentsarXiv:2601.06617v3PDF
0

Posted in cs.HC · 2026-01-10 · Blessing Jerry, Lourdes Moreno, Virginia Francisco, Raquel Hervas

LLM-Driven Accessible Interface: A Model-Based Approach

The integration of Large Language Models (LLMs) into interactive systems opens new opportunities for adaptive user experiences, yet it also raises challenges regarding accessibility, explainability, and normative compliance. This paper presents an implemented model-driven architecture for generating personalised, multimodal, and...

💬 0 commentsarXiv:2601.06616v1PDF