Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through July 20, 2026 — 03:18:17 EST

0

Posted in cs.CV · 2026-01-13 · Yan Zhu, Te Luo, Pei-Yao Fu, Zhen Zhang, Zi-Long Wang, Yi-Fan Qu, Zi-Han Geng, Jia-Qi Xu, Lu Yao, Li-Yun Ma, Wei Su, Wei-Feng Chen, Quan-Lin Li, Shuo Wang, Ping-Hong Zhou

GI-Bench: A Panoramic Benchmark Revealing the Knowledge-Experience Dissociation of Multimodal Large Language Models in Gastrointestinal Endoscopy Against Clinical Standards

Multimodal Large Language Models (MLLMs) show promise in gastroenterology, yet their performance against comprehensive clinical workflows and human benchmarks remains unverified. To systematically evaluate state-of-the-art MLLMs across a panoramic gastrointestinal endoscopy workflow and determine their clinical utility compared with...

💬 0 commentsarXiv:2601.08183v2PDF
0

Posted in cs.CV · 2026-01-13 · Jiamiao Lu, Dongbo Xie, Junjie Qiu, Lingkun Ma, Changming Sun, Weichuan Zhang

Second-order Gaussian directional derivative representations for image high-resolution corner detection

Corner detection is widely used in various computer vision tasks, such as image matching and 3D reconstruction. Our research indicates that there are theoretical flaws in Zhang et al.'s use of a simple corner model to obtain a series of corner characteristics, as the grayscale information of two adjacent corners can affect each other....

💬 0 commentsarXiv:2601.08182v2PDF
0

Posted in cs.LG · 2026-01-13 · Aviral Gupta, Armaan Sethi, Dhruv Kumar

TabPFN Through The Looking Glass: An interpretability study of TabPFN and its internal representations

Tabular foundational models are pre-trained models designed for a wide range of tabular data tasks. They have shown strong performance across domains, yet their internal representations and learned concepts remain poorly understood. This lack of interpretability makes it important to study how these models process and transform input...

💬 0 commentsarXiv:2601.08181v1PDF
0

Posted in cs.CV · 2026-01-13 · Anh H. Vo, Tae-Seok Kim, Hulin Jin, Soo-Mi Choi, Yong-Guk Kim

Instruction-Driven 3D Facial Expression Generation and Transition

A 3D avatar typically has one of six cardinal facial expressions. To simulate realistic emotional variation, we should be able to render a facial transition between two arbitrary expressions. This study presents a new framework for instruction-driven facial expression generation that produces a 3D face and, starting from an image of...

💬 0 commentsarXiv:2601.08179v1PDF
0

Posted in cs.HC · 2026-01-13 · Shangqian Li, Tianwa Chen, Gianluca Demartini

The Impact of AI Generated Content on Decision Making for Topics Requiring Expertise

Modelling users' online decision-making and opinion change is a complex issue that needs to consider users' personal determinants, the nature of the topic and the information retrieval activities. Furthermore, generative-AIbased products like ChatGPT gradually become an essential element for the retrieval of online information....

💬 0 commentsarXiv:2601.08178v1PDF
0

Posted in cs.CL · 2026-01-13 · Lavanya Prahallad, Sai Utkarsh Choudarypally, Pragna Prahallad, Pranathi Prahallad

Prompt-Based Clarity Evaluation and Topic Detection in Political Question Answering

Automatic evaluation of large language model (LLM) responses requires not only factual correctness but also clarity, particularly in political question-answering. While recent datasets provide human annotations for clarity and evasion, the impact of prompt design on automatic clarity evaluation remains underexplored. In this paper, we...

💬 0 commentsarXiv:2601.08176v1PDF
0

Posted in cs.CV · 2026-01-13 · Feiran Wang, Junyi Wu, Dawen Cai, Yuan Hong, Yan Yan

CogniMap3D: Cognitive 3D Mapping and Rapid Retrieval

We present CogniMap3D, a bioinspired framework for dynamic 3D scene understanding and reconstruction that emulates human cognitive processes. Our approach maintains a persistent memory bank of static scenes, enabling efficient spatial knowledge storage and rapid retrieval. CogniMap3D integrates three core capabilities: a multi-stage...

💬 0 commentsarXiv:2601.08175v1PDF
0

Posted in cs.CV · 2026-01-13 · Xiyan Feng, Wenbo Zhang, Lu Zhang, Yunzhi Zhuge, Huchuan Lu, You He

Towards Cross-Platform Generalization: Domain Adaptive 3D Detection with Augmentation and Pseudo-Labeling

This technical report represents the award-winning solution to the Cross-platform 3D Object Detection task in the RoboSense2025 Challenge. Our approach is built upon PVRCNN++, an efficient 3D object detection framework that effectively integrates point-based and voxel-based features. On top of this foundation, we improve...

💬 0 commentsarXiv:2601.08174v1PDF
0

Posted in cs.AI · 2026-01-13 · Daocheng Fu, Jianbiao Mei, Rong Wu, Xuemeng Yang, Jia Xu, Ding Wang, Pinlong Cai, Yong Liu, Licheng Wen, Botian Shi

The Agent's First Day: Benchmarking Learning, Exploration, and Scheduling in the Workplace Scenarios

The rapid evolution of Multi-modal Large Language Models (MLLMs) has advanced workflow automation; however, existing research mainly targets performance upper bounds in static environments, overlooking robustness for stochastic real-world deployment. We identify three key challenges: dynamic task scheduling, active exploration under...

💬 0 commentsarXiv:2601.08173v2PDF
0

Posted in cs.LG · 2026-01-13 · Farhad Mirkarimi

VBO-MI: A Fully Gradient-Based Bayesian Optimization Framework Using Variational Mutual Information Estimation

Many real-world tasks require optimizing expensive black-box functions accessible only through noisy evaluations, a setting commonly addressed with Bayesian optimization (BO). While Bayesian neural networks (BNNs) have recently emerged as scalable alternatives to Gaussian Processes (GPs), traditional BNN-BO frameworks remain burdened...

💬 0 commentsarXiv:2601.08172v1PDF
0

Posted in cs.CL · 2026-01-13 · Andrea Kang, Yingnian Wu, Hongjing Lu

Relational Knowledge Distillation Using Fine-tuned Function Vectors

Representing relations between concepts is a core prerequisite for intelligent systems to make sense of the world. Recent work using causal mediation analysis has shown that a small set of attention heads encodes task representation in in-context learning, captured in a compact representation known as the function vector. We show that...

💬 0 commentsarXiv:2601.08169v1PDF
0

Posted in cs.AI · 2026-01-13 · Mohammad Pivezhandi, Mahdi Banisharif, Abusayeed Saifullah, Ali Jannesari

ZeroDVFS: Zero-Shot LLM-Guided Core and Frequency Allocation for Embedded Platforms

Dynamic voltage and frequency scaling (DVFS) and task-to-core allocation are critical for thermal management and balancing energy and performance in embedded systems. Existing approaches either rely on utilization-based heuristics that overlook stall times, or require extensive offline profiling for table generation, preventing...

💬 0 commentsarXiv:2601.08166v2PDF
0

Posted in cs.CV · 2026-01-13 · Phuoc-Nguyen Bui, Toan Duc Nguyen, Junghyun Bum, Duc-Tai Le, Hyunseung Choo

Representation Learning with Semantic-aware Instance and Sparse Token Alignments

Medical contrastive vision-language pre-training (VLP) has demonstrated significant potential in improving performance on downstream tasks. Traditional approaches typically employ contrastive learning, treating paired image-report samples as positives and unpaired ones as negatives. However, in medical datasets, there can be...

💬 0 commentsarXiv:2601.08165v2PDF
0

Posted in cs.CV · 2026-01-13 · Jing Tao, Banglei Guan, Pengju Sun, Taihang Lei, Yang Shang, Qifeng Yu

A Hardware-Algorithm Co-Designed Framework for HDR Imaging and Dehazing in Extreme Rocket Launch Environments

Quantitative optical measurement of critical mechanical parameters -- such as plume flow fields, shock wave structures, and nozzle oscillations -- during rocket launch faces severe challenges due to extreme imaging conditions. Intense combustion creates dense particulate haze and luminance variations exceeding 120 dB, degrading image...

💬 0 commentsarXiv:2601.08162v1PDF
0

Posted in cs.RO · 2026-01-13 · Jing Tao, Banglei Guan, Yang Shang, Shunkun Liang, Qifeng Yu

Robust Subpixel Localization of Diagonal Markers in Large-Scale Navigation via Multi-Layer Screening and Adaptive Matching

This paper proposes a robust, high-precision positioning methodology to address localization failures arising from complex background interference in large-scale flight navigation and the computational inefficiency inherent in conventional sliding window matching techniques. The proposed methodology employs a three-tiered framework...

💬 0 commentsarXiv:2601.08161v1PDF
0

Posted in cs.CL · 2026-01-13 · Anxin Tian, Yiming Li, Xing Li, Hui-Ling Zhen, Lei Chen, Xianzhi Yu, Zhenhua Dong, Mingxuan Yuan

SwiftMem: Fast Agentic Memory via Query-aware Indexing

Agentic memory systems have become critical for enabling LLM agents to maintain long-term context and retrieve relevant information efficiently. However, existing memory frameworks suffer from a fundamental limitation: they perform exhaustive retrieval across the entire storage layer regardless of query characteristics. This...

💬 0 commentsarXiv:2601.08160v1PDF
0

Posted in cs.CL · 2026-01-13 · Yuqing Zhou, Zhuoer Wang, Jie Yuan, Hong Wang, Samson Koelle, Ziwei Zhu, Wei Niu

WISE-Flow: Workflow-Induced Structured Experience for Self-Evolving Conversational Service Agents

Large language model (LLM)-based agents are widely deployed in user-facing services but remain error-prone in new tasks, tend to repeat the same failure patterns, and show substantial run-to-run variability. Fixing failures via environment-specific training or manual patching is costly and hard to scale. To enable self-evolving agents...

💬 0 commentsarXiv:2601.08158v1PDF
0

Posted in cs.AI · 2026-01-13 · Arin Gopalan Yadav, Varad Dherange, Kumar Shivam

Project Synapse: A Hierarchical Multi-Agent Framework with Hybrid Memory for Autonomous Resolution of Last-Mile Delivery Disruptions

This paper introduces Project Synapse, a novel agentic framework designed for the autonomous resolution of last-mile delivery disruptions. Synapse employs a hierarchical multi-agent architecture in which a central Resolution Supervisor agent performs strategic task decomposition and delegates subtasks to specialized worker agents...

💬 0 commentsarXiv:2601.08156v1PDF
0

Posted in cs.CV · 2026-01-13 · Inpyo Song, Minjun Joo, Joonhyung Kwon, Eunji Jeon, Jangwon Lee

Instance-Aligned Captions for Explainable Video Anomaly Detection

Explainable video anomaly detection (VAD) is crucial for safety-critical applications, yet even with recent progress, much of the research still lacks spatial grounding, making the explanations unverifiable. This limitation is especially pronounced in multi-entity interactions, where existing explainable VAD methods often produce...

💬 0 commentsarXiv:2601.08155v1PDF
0

Posted in cs.NI · 2026-01-13 · Thakshila Perera, Amine Mezghani, Ekram Hossain

Multi-Objective Optimization for Joint Communication and Sensing in Multi-user MIMO Systems: Characterizing the Pareto Boundary

This paper investigates the Pareto boundary performance of a joint communication and sensing (JCAS) system that addresses both sensing and communication functions at the same time. In this scenario, a multiple-antenna base station (BS) transmits information to multiple single-antenna communication users while concurrently estimating...

💬 0 commentsarXiv:2601.08152v1PDF
0

Posted in cs.CV · 2026-01-13 · Shezheng Song, Shasha Li, Jie Yu

Where Does Vision Meet Language? Understanding and Refining Visual Fusion in MLLMs via Contrastive Attention

Multimodal Large Language Models (MLLMs) have achieved remarkable progress in vision-language understanding, yet how they internally integrate visual and textual information remains poorly understood. To bridge this gap, we perform a systematic layer-wise masking analysis across multiple architectures, revealing how visual-text fusion...

💬 0 commentsarXiv:2601.08151v1PDF
0

Posted in cs.IR · 2026-01-13 · Seokho Ahn, Sungbok Shin, Young-Duk Seo

Enriching Semantic Profiles into Knowledge Graph for Recommender Systems Using Large Language Models

Rich and informative profiling to capture user preferences is essential for improving recommendation quality. However, there is still no consensus on how best to construct and utilize such profiles. To address this, we revisit recent profiling-based approaches in recommender systems along four dimensions: 1) knowledge base, 2)...

💬 0 commentsarXiv:2601.08148v1PDF
0

Posted in cs.LG · 2026-01-13 · Chaoqun Fei, Huanjiang Liu, Tinglve Zhou, Yangyang Li, Tianyong Hao

Dynamic Graph Structure Learning via Resistance Curvature Flow

Geometric Representation Learning (GRL) aims to approximate the non-Euclidean topology of high-dimensional data through discrete graph structures, grounded in the manifold hypothesis. However, traditional static graph construction methods based on Euclidean distance often fail to capture the intrinsic curvature characteristics of the...

💬 0 commentsarXiv:2601.08149v1PDF
0

Posted in cs.CL · 2026-01-13 · Khumaisa Nur'aini, Ayu Purwarianti, Alham Fikri Aji, Derry Wijaya

Beyond Transfer Accuracy: Faithful Circuits for Controlled Low-Resource Adaptation

Existing circuit discovery methods rely on templated tasks with clean counterfactuals, limiting their use on diverse natural text. We adapt Contextual Decomposition for Transformers (CD-T) for unstructured settings via label-balanced activation means and task-directional relevance scoring, enabling counterfactual-free circuit...

💬 0 commentsarXiv:2601.08146v3PDF
0

Posted in cs.IT · 2026-01-13 · Junfeng Jia, Yanxun Chang

Cardinality-consistent flag codes with longer type vectors

Flag codes generalize constant dimension codes by considering sequences of nested subspaces with prescribed dimensions as codewords. A comprehensive construction, which unites cyclic orbit flag codes, yields two families of flag codes on $\mathbb{F}^n_q$ (where $n=sk+h$ with $s\geq 2$ and $0\leq h < k$): optimum distance flag codes of...

💬 0 commentsarXiv:2601.08144v1PDF