Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through July 28, 2026 — 19:12:45 EST

0

Posted in cs.CV · 2026-01-10 · Hao Tang, Ting Huang, Zeyu Zhang

3D CoCa v2: Contrastive Learners with Test-Time Search for Generalizable Spatial Intelligence

Spatial intelligence refers to the ability to perceive, reason about, and describe objects and their relationships within three-dimensional environments, forming a foundation for embodied perception and scene understanding. 3D captioning aims to describe 3D scenes in natural language; however, it remains challenging due to the...

💬 0 commentsarXiv:2601.06496v1PDF
0

Posted in cs.IT · 2026-01-10 · Han Li, Xiang Wang, Fang-Wei Fu

On the Number of Subsequences in the Nonbinary Deletion Channel

In the deletion channel, an important problem is to determine the number of subsequences derived from a string $U$ of length $n$ when subjected to $t$ deletions. It is well-known that the number of subsequences in the setting exhibits a strong dependence on the number of runs in the string $U$, where a run is defined as a maximal...

💬 0 commentsarXiv:2601.06493v2PDF
0

Posted in cs.IT · 2026-01-10 · Chun-Neng Chu, Wei-Fu Tseng, Yen-Huan Li

Algorithms for Computing the Petz-Augustin Capacity

We propose the first algorithms with non-asymptotic convergence guarantees for computing the Petz-Augustin capacity, which generalizes the channel capacity and characterizes the optimal error exponent in classical-quantum channel coding. This capacity can be equivalently expressed as the maximization of two generalizations of mutual...

💬 0 commentsarXiv:2601.06492v1PDF
0

Posted in cs.MA · 2026-01-10 · Wenyu Mao, Haosong Tan, Shuchang Liu, Haoyang Liu, Yifan Xu, Huaxiang Ji, Xiang Wang

Bi-Mem: Bidirectional Construction of Hierarchical Memory for Personalized LLMs via Inductive-Reflective Agents

Constructing memory from users' long-term conversations overcomes LLMs' contextual limitations and enables personalized interactions. Recent studies focus on hierarchical memory to model users' multi-granular behavioral patterns via clustering and aggregating historical conversations. However, conversational noise and memory...

💬 0 commentsarXiv:2601.06490v1PDF
0

Posted in cs.LG · 2026-01-10 · Qiang Zhang, Boli Chen, Fanrui Zhang, Ruixue Ding, Shihang Wang, Qiuchen Wang, Yinfeng Huang, Haonan Zhang, Rongxiang Zhu, Pengyong Wang, Ailin Ren, Xin Li, Pengjun Xie, Jiawei Liu, Ning Guo, Jingren Zhou, Zheng-Jun Zha

ArenaRL: Scaling RL for Open-Ended Agents via Tournament-based Relative Ranking

Reinforcement learning has substantially improved the performance of LLM agents on tasks with verifiable outcomes, but it still struggles on open-ended agent tasks with vast solution spaces (e.g., complex travel planning). Due to the absence of objective ground-truth for these tasks, current RL algorithms largely rely on reward models...

💬 0 commentsarXiv:2601.06487v2PDF
0

Posted in cs.CV · 2026-01-10 · Yue Wang, Lawrence Amadi, Xiang Gao, Yazheng Chen, Yuanpeng Liu, Ning Lu, Xianfeng Gu

Learning Domain Agnostic Latent Embeddings of 3D Faces for Zero-shot Animal Expression Transfer

We present a zero-shot framework for transferring human facial expressions to 3D animal face meshes. Our method combines intrinsic geometric descriptors (HKS/WKS) with a mesh-agnostic latent embedding that disentangles facial identity and expression. The ID latent space captures species-independent facial structure, while the...

💬 0 commentsarXiv:2601.06484v1PDF
0

Posted in cs.CV · 2026-01-10 · JiaLin Zhang, Dong Li

SRFlow: A Dataset and Regularization Model for High-Resolution Facial Optical Flow via Splatting Rasterization

Facial optical flow supports a wide range of tasks in facial motion analysis. However, the lack of high-resolution facial optical flow datasets has hindered progress in this area. In this paper, we introduce Splatting Rasterization Flow (SRFlow), a high-resolution facial optical flow dataset, and Splatting Rasterization Guided FlowNet...

💬 0 commentsarXiv:2601.06479v1PDF
0

Posted in cs.CL · 2026-01-10 · Debasmita Panda, Akash Anil, Neelesh Kumar Shukla

IndRegBias: A Dataset for Studying Indian Regional Biases in English and Code-Mixed Social Media Comments

Warning: This paper consists of examples representing regional biases in Indian regions that might be offensive towards a particular region. While social biases corresponding to gender, race, socio-economic conditions, etc., have been extensively studied in the major applications of Natural Language Processing (NLP), biases...

💬 0 commentsarXiv:2601.06477v2PDF
0

Posted in cs.CV · 2026-01-10 · Kai Cheng, Ruoqi Wang, Qiong Luo

VVTRec: Radio Interferometric Reconstruction through Visual and Textual Modality Enrichment

Radio astronomy is an indispensable discipline for observing distant celestial objects. Measurements of wave signals from radio telescopes, called visibility, need to be transformed into images for astronomical observations. These dirty images blend information from real sources and artifacts. Therefore, astronomers usually perform...

💬 0 commentsarXiv:2601.06475v1PDF
0

Posted in cs.CV · 2026-01-10 · Chenxu Dang, Jie Wang, Guang Li, Zhiwen Hou, Zihan You, Hangjun Ye, Jie Ma, Long Chen, Yan Wang

SparseOccVLA: Bridging Occupancy and Vision-Language Models via Sparse Queries for Unified 4D Scene Understanding and Planning

In autonomous driving, Vision Language Models (VLMs) excel at high-level reasoning , whereas semantic occupancy provides fine-grained details. Despite significant progress in individual fields, there is still no method that can effectively integrate both paradigms. Conventional VLMs struggle with token explosion and limited...

💬 0 commentsarXiv:2601.06474v2PDF
0

Posted in cs.MA · 2026-01-10 · Sathish Sampath, Anuradha Baskaran

Adaptive Orchestration: Scalable Self-Evolving Multi-Agent Systems

As Large Language Models (LLMs) are increasingly deployed as autonomous agents, they face a critical scalability bottleneck known as the "Generalization-Specialization Dilemma." Monolithic agents equipped with extensive toolkits suffer from context pollution and attention decay, leading to hallucinations. Conversely, static...

💬 0 commentsarXiv:2601.09742v1PDF
0

Posted in cs.LG · 2026-01-10 · Chutian Huang, Chang Ma, Kaibo Wang, Yang Xiang

StablePDENet: Enhancing Stability of Operator Learning for Solving Differential Equations

Learning solution operators for differential equations with neural networks has shown great potential in scientific computing, but ensuring their stability under input perturbations remains a critical challenge. This paper presents a robust self-supervised neural operator framework that enhances stability through adversarial training...

💬 0 commentsarXiv:2601.06472v1PDF
0

Posted in cs.CL · 2026-01-10 · Junho Park, Dohoon Kim, Taesup Moon

PRISP: Privacy-Safe Few-Shot Personalization via Lightweight Adaptation

Large language model (LLM) personalization aims to adapt general-purpose models to individual users. Most existing methods, however, are developed under data-rich and resource-abundant settings, often incurring privacy risks. In contrast, realistic personalization typically occurs after deployment under (i) extremely limited user...

💬 0 commentsarXiv:2601.06471v1PDF
0

Posted in cs.CE · 2026-01-10 · Weipeng Xu, Ziyuan Xie, Haoju Lin, Xinyu Wang, Guangjin Mou, Tianju Xue

Style-constrained inverse design of microstructures with tailored mechanical properties using unconditional diffusion models

Deep generative models, particularly denoising diffusion models, have achieved remarkable success in high-fidelity generation of architected microstructures with desired properties and styles. Nevertheless, these recent methods typically rely on conditional training mechanisms and demand substantial computational effort to prepare the...

💬 0 commentsarXiv:2601.06469v1PDF
0

Posted in cs.LG · 2026-01-10 · Anh-Tuan Mai, Cam-Van Thi Nguyen, Duc-Trong Le

Divide and Refine: Enhancing Multimodal Representation and Explainability for Emotion Recognition in Conversation

Multimodal emotion recognition in conversation (MERC) requires representations that effectively integrate signals from multiple modalities. These signals include modality-specific cues, information shared across modalities, and interactions that emerge only when modalities are combined. In information-theoretic terms, these correspond...

💬 0 commentsarXiv:2601.14274v1PDF
0

Posted in cs.CR · 2026-01-10 · Imtiaz Ali Soomro, Hamood Ur Rehman, S. Jawad Hussain ID, Adeel Iqbal, Waqas Khalid, Heejung Yu ID

SecureDyn-FL: A Robust Privacy-Preserving Federated Learning Framework for Intrusion Detection in IoT Networks

The rapid proliferation of Internet of Things (IoT) devices across domains such as smart homes, industrial control systems, and healthcare networks has significantly expanded the attack surface for cyber threats, including botnet-driven distributed denial-of-service (DDoS), malware injection, and data exfiltration. Conventional...

💬 0 commentsarXiv:2601.06466v1PDF
0

Posted in cs.CV · 2026-01-10 · Chao Liu, Ngai-Man Cheung

On the Adversarial Robustness of 3D Large Vision-Language Models

3D Vision-Language Models (VLMs), such as PointLLM and GPT4Point, have shown strong reasoning and generalization abilities in 3D understanding tasks. However, their adversarial robustness remains largely unexplored. Prior work in 2D VLMs has shown that the integration of visual inputs significantly increases vulnerability to...

💬 0 commentsarXiv:2601.06464v1PDF
0

Posted in cs.LG · 2026-01-10 · Xuezhe Ma, Shicheng Wen, Linghao Jin, Bilge Acun, Ruihang Lai, Bohan Hou, Will Lin, Hao Zhang, Songlin Yang, Ryan Lee, Mengxi Wu, Jonathan May, Luke Zettlemoyer, Carole-Jean Wu

Gecko: An Efficient Neural Architecture Inherently Processing Sequences with Arbitrary Lengths

Designing a unified neural network to efficiently and inherently process sequential data with arbitrary lengths is a central and challenging problem in sequence modeling. The design choices in Transformer, including quadratic complexity and weak length extrapolation, have limited their ability to scale to long sequences. In this work,...

💬 0 commentsarXiv:2601.06463v1PDF
0

Posted in cs.CR · 2026-01-10 · Minfeng Qi, Dongyang He, Qin Wang, Lefeng Zhang

VIPER Strike: Defeating Visual Reasoning CAPTCHAs via Structured Vision-Language Inference

Visual Reasoning CAPTCHAs (VRCs) combine visual scenes with natural-language queries that demand compositional inference over objects, attributes, and spatial relations. They are increasingly deployed as a primary defense against automated bots. Existing solvers fall into two paradigms: vision-centric, which rely on template-specific...

💬 0 commentsarXiv:2601.06461v1PDF
0

Posted in cs.CV · 2026-01-10 · Weihao Hong, Zhiyuan Jiang, Bingyu Shen, Xinlei Guan, Yangyi Feng, Meng Xu, Boyang Li

Tone Matters: The Impact of Linguistic Tone on Hallucination in VLMs

Vision-Language Models (VLMs) are increasingly used in safety-critical applications that require reliable visual grounding. However, these models often hallucinate details that are not present in the image to satisfy user prompts. While recent datasets and benchmarks have been introduced to evaluate systematic hallucinations in VLMs,...

💬 0 commentsarXiv:2601.06460v1PDF
0

Posted in cs.IR · 2026-01-10 · Sayak Chakrabarty, Souradip Pal

PixRec: Leveraging Visual Context for Next-Item Prediction in Sequential Recommendation

Large Language Models (LLMs) have recently shown strong potential for usage in sequential recommendation tasks through text-only models, which combine advanced prompt design, contrastive alignment, and fine-tuning on downstream domain-specific data. While effective, these approaches overlook the rich visual information present in many...

💬 0 commentsarXiv:2601.06458v1PDF
0

Posted in cs.SE · 2026-01-10 · Shaunak Biswas, Hiya Bhatt, Karthik Vaidhyanathan

Architecting AgentOps Needs CHANGE

The emergence of Agentic AI systems has outpaced the architectural thinking required to operate them effectively. These agents differ fundamentally from traditional software: their behavior is not fixed at deployment but continuously shaped by experience, feedback, and context. Applying operational principles inherited from DevOps or...

💬 0 commentsarXiv:2601.06456v1PDF
0

Posted in cs.AI · 2026-01-10 · Hyungjun Yoon, Mohammad Malekzadeh, Sung-Ju Lee, Fahim Kawsar, Lorena Qendro

ConSensus: Multi-Agent Collaboration for Multimodal Sensing

Large language models (LLMs) are increasingly grounded in sensor data to perceive and reason about human physiology and the physical world. However, accurately interpreting heterogeneous multimodal sensor data remains a fundamental challenge. We show that a single monolithic LLM often fails to reason coherently across modalities,...

💬 0 commentsarXiv:2601.06453v2PDF
0

Posted in cs.RO · 2026-01-10 · Hyunseo Koh, Chang-Yong Song, Youngjae Choi, Misa Viveiros, David Hyde, Heewon Kim

CulinaryCut-VLAP: A Vision-Language-Action-Physics Framework for Food Cutting via a Force-Aware Material Point Method

Food cutting is a highly practical yet underexplored application at the intersection of vision and robotic manipulation. The task remains challenging because interactions between the knife and deformable materials are highly nonlinear and often entail large deformations, frequent contact, and topological change, which in turn hinder...

💬 0 commentsarXiv:2601.06451v1PDF