Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through July 20, 2026 — 06:46:41 EST

0

Posted in cs.CV · 2026-01-12 · Soumyaroop Nandi, Prem Natarajan

Rescind: Countering Image Misconduct in Biomedical Publications with Vision-Language and State-Space Modeling

Scientific image manipulation in biomedical publications poses a growing threat to research integrity and reproducibility. Unlike natural image forensics, biomedical forgery detection is uniquely challenging due to domain-specific artifacts, complex textures, and unstructured figure layouts. We present the first vision-language guided...

💬 0 commentsarXiv:2601.08040v1PDF
0

Posted in cs.LG · 2026-01-12 · Shaocong Ma, Heng Huang

Riemannian Zeroth-Order Gradient Estimation with Structure-Preserving Metrics for Geodesically Incomplete Manifolds

In this paper, we study Riemannian zeroth-order optimization in settings where the underlying Riemannian metric $g$ is geodesically incomplete, and the goal is to approximate stationary points with respect to this incomplete metric. To address this challenge, we construct structure-preserving metrics that are geodesically complete...

💬 0 commentsarXiv:2601.08039v2PDF
0

Posted in cs.SE · 2026-01-12 · Bonan Kou, Zijie Zhou, Muhao Chen, Tianyi Zhang

Automating API Documentation from Crowdsourced Knowledge

API documentation is crucial for developers to learn and use APIs. However, it is known that many official API documents are obsolete and incomplete. To address this challenge, we propose a new approach called AutoDoc that generates API documents with API knowledge extracted from online discussions on Stack Overflow (SO). AutoDoc...

💬 0 commentsarXiv:2601.08036v1PDF
0

Posted in cs.HC · 2026-01-12 · David Elsweiler

From Tool to Teacher: Rethinking Search Systems as Instructive Interfaces

Information access systems such as search engines and generative AI are central to how people seek, evaluate, and interpret information. Yet most systems are designed to optimise retrieval rather than to help users develop better search strategies or critical awareness. This paper introduces a pedagogical perspective on information...

💬 0 commentsarXiv:2601.08035v1PDF
0

Posted in cs.RO · 2026-01-12 · Cameron Smith, Basile Van Hoorick, Vitor Guizilini, Yue Wang

Fiducial Exoskeletons: Image-Centric Robot State Estimation

We introduce Fiducial Exoskeletons, an image-based reformulation of 3D robot state estimation that replaces cumbersome procedures and motor-centric pipelines with single-image inference. Traditional approaches - especially robot-camera extrinsic estimation - often rely on high-precision actuators and require time-consuming routines...

💬 0 commentsarXiv:2601.08034v1PDF
0

Posted in cs.LG · 2026-01-12 · Amir Eskandari, Aman Anand, Elyas Rashno, Farhana Zulkernine

InfGraND: An Influence-Guided GNN-to-MLP Knowledge Distillation

Graph Neural Networks (GNNs) are the go-to model for graph data analysis. However, GNNs rely on two key operations - aggregation and update, which can pose challenges for low-latency inference tasks or resource-constrained scenarios. Simple Multi-Layer Perceptrons (MLPs) offer a computationally efficient alternative. Yet, training an...

💬 0 commentsarXiv:2601.08033v1PDF
0

Posted in cs.IT · 2026-01-12 · Thomas F. Varley

The many faces of multivariate information

Extracting higher-order structures from multivariate data has become an area of intensive study in complex systems science, as these multipartite interactions can reveal insights into fundamental features of complex systems like emergent phenomena. Information theory provides a natural language for exploring these interactions, as it...

💬 0 commentsarXiv:2601.08030v2PDF
0

Posted in cs.CV · 2026-01-12 · Jifeng Song, Arun Das, Pan Wang, Hui Ji, Kun Zhao, Yufei Huang

FigEx2: Visual-Conditioned Panel Detection and Captioning for Scientific Compound Figures

Scientific compound figures combine multiple labeled panels into a single image. However, in a PMC-scale crawl of 346,567 compound figures, 16.3% have no caption and 1.8% only have captions shorter than ten words, causing them to be discarded by existing caption-decomposition pipelines. We propose FigEx2, a visual-conditioned...

💬 0 commentsarXiv:2601.08026v4PDF
0

Posted in cs.DC · 2026-01-12 · Adiba Masud, Nicholas Foley, Pragathi Durga Rajarajan, Palden Lama

Where to Split? A Pareto-Front Analysis of DNN Partitioning for Edge Inference

The deployment of deep neural networks (DNNs) on resource-constrained edge devices is frequently hindered by their significant computational and memory requirements. While partitioning and distributing a DNN across multiple devices is a well-established strategy to mitigate this challenge, prior research has largely focused on...

💬 0 commentsarXiv:2601.08025v1PDF
0

Posted in cs.CV · 2026-01-12 · Amin Abbasishahkoo, Mahboubeh Dadkhah, Lionel Briand

A Highly Efficient Diversity-based Input Selection for DNN Improvement Using VLMs

Maintaining or improving the performance of Deep Neural Networks (DNNs) through fine-tuning requires labeling newly collected inputs, a process that is often costly and time-consuming. To alleviate this problem, input selection approaches have been developed in recent years to identify small, yet highly informative subsets for...

💬 0 commentsarXiv:2601.08024v1PDF
0

Posted in cs.CV · 2026-01-12 · Samet Hicsonmez, Abd El Rahman Shabayek, Djamila Aouada

Training Free Zero-Shot Visual Anomaly Localization via Diffusion Inversion

Zero-Shot image Anomaly Detection (ZSAD) aims to detect and localise anomalies without access to any normal training samples of the target data. While recent ZSAD approaches leverage additional modalities such as language to generate fine-grained prompts for localisation, vision-only methods remain limited to image-level...

💬 0 commentsarXiv:2601.08022v1PDF
0

Posted in cs.CV · 2026-01-12 · Evžen Wybitul, Javier Rando, Florian Tramèr, Stanislav Fort

Representations of Text and Images Align From Layer One

We show that for a variety of concepts in adapter-based vision-language models, the representations of their images and their text descriptions are meaningfully aligned from the very first layer. This contradicts the established view that such image-text alignment only appears in late layers. We show this using a new synthesis-based...

💬 0 commentsarXiv:2601.08017v1PDF
0

Posted in cs.LG · 2026-01-12 · Nairui Liu, Fang He, Xindi Tang, Yineng Wang

Beyond the Next Port: A Multi-Task Transformer for Forecasting Future Voyage Segment Durations

Accurate forecasts of segment-level sailing durations are fundamental to enhancing maritime schedule reliability and optimizing long-term port operations. However, conventional estimated time of arrival (ETA) models are primarily designed for the immediate next port of call and rely heavily on real-time automatic identification system...

💬 0 commentsarXiv:2601.08013v2PDF
0

Posted in cs.SE · 2026-01-12 · Aarya Doshi, Yining Hong, Congying Xu, Eunsuk Kang, Alexandros Kapravelos, Christian Kästner

Towards Verifiably Safe Tool Use for LLM Agents

Large language model (LLM)-based AI agents extend LLM capabilities by enabling access to tools such as data sources, APIs, search engines, code sandboxes, and even other agents. While this empowers agents to perform complex tasks, LLMs may invoke unintended tool interactions and introduce risks, such as leaking sensitive data or...

💬 0 commentsarXiv:2601.08012v1PDF
0

Posted in cs.CV · 2026-01-12 · Xin Jin, Yichuan Zhong, Yapeng Tian

TP-Blend: Textual-Prompt Attention Pairing for Precise Object-Style Blending in Diffusion Models

Current text-conditioned diffusion editors handle single object replacement well but struggle when a new object and a new style must be introduced simultaneously. We present Twin-Prompt Attention Blend (TP-Blend), a lightweight training-free framework that receives two separate textual prompts, one specifying a blend object and the...

💬 0 commentsarXiv:2601.08011v4PDF
0

Posted in cs.CV · 2026-01-12 · Chaoyu Li, Fei Tao, Pooyan Fazli

CASHEW: Stabilizing Multimodal Reasoning via Iterative Trajectory Aggregation

Vision-language models achieve strong performance across a wide range of multimodal understanding and reasoning tasks, yet their multi-step reasoning remains unstable. Repeated sampling over the same input often produces divergent reasoning trajectories and inconsistent final predictions. To address this, we introduce two...

💬 0 commentsarXiv:2601.08010v3PDF
0

Posted in cs.AI · 2026-01-12 · Joe Kwon, Stephen Casper

Internal Deployment Gaps in AI Regulation

Frontier AI regulations primarily focus on systems deployed to external users, where deployment is more visible and subject to outside scrutiny. However, high-stakes applications can occur internally when companies deploy highly capable systems within their own organizations, such as for automating R&D, accelerating critical business...

💬 0 commentsarXiv:2601.08005v3PDF
0

Posted in cs.CL · 2026-01-12 · Weiyue Li, Mingxiao Song, Zhenda Shen, Dachuan Zhao, Yunfan Long, Yi Li, Yongce Li, Ruyi Yang, Mengyu Wang

LLM Review: Enhancing Creative Writing via Blind Peer Review Feedback

Large Language Models (LLMs) often struggle with creative generation, and multi-agent frameworks that improve reasoning through interaction can paradoxically hinder creativity by inducing content homogenization. We introduce LLM Review, a peer-review-inspired framework implementing Blind Peer Review: agents exchange targeted feedback...

💬 0 commentsarXiv:2601.08003v1PDF
0

Posted in cs.AI · 2026-01-12 · Can Jin, Rui Wu, Tong Che, Qixin Zhang, Hongwu Peng, Jiahui Zhao, Zhenting Wang, Wenqi Wei, Ligong Han, Zhao Zhang, Yuan Cao, Ruixiang Tang, Dimitris N. Metaxas

Reasoning over Precedents Alongside Statutes: Case-Augmented Deliberative Alignment for LLM Safety

Ensuring that Large Language Models (LLMs) adhere to safety principles without refusing benign requests remains a significant challenge. While OpenAI introduces deliberative alignment (DA) to enhance the safety of its o-series models through reasoning over detailed ``code-like'' safety rules, the effectiveness of this approach in...

💬 0 commentsarXiv:2601.08000v1PDF
0

Posted in cs.SD · 2026-01-12 · Tiantian Feng, Anfeng Xu, Jinkook Lee, Shrikanth Narayanan

VoxCog: Towards End-to-End Multilingual Cognitive Impairment Classification through Dialectal Knowledge

In this work, we present a novel perspective on cognitive impairment classification from speech by integrating speech foundation models that explicitly recognize speech dialects. Our motivation is based on the observation that individuals with Alzheimer's Disease (AD) or mild cognitive impairment (MCI) often produce measurable speech...

💬 0 commentsarXiv:2601.07999v1PDF
0

Posted in cs.CV · 2026-01-12 · Hongwei Lin, Diego Andrade, Mini Das, Howard C. Gifford

Predicting Region of Interest in Human Visual Search Based on Statistical Texture and Gabor Features

Understanding human visual search behavior is a fundamental problem in vision science and computer vision, with direct implications for modeling how observers allocate attention in location-unknown search tasks. In this study, we investigate the relationship between Gabor-based features and gray-level co-occurrence matrix (GLCM) based...

💬 0 commentsarXiv:2601.07998v1PDF
0

Posted in cs.CL · 2026-01-12 · Laurits Lyngbaek, Pascale Feldkamp, Yuri Bizzoni, Kristoffer L. Nielbo, Kenneth Enevoldsen

Is Sentiment Banana-Shaped? Exploring the Geometry and Portability of Sentiment Concept Vectors

Use cases of sentiment analysis in the humanities often require contextualized, continuous scores. Concept Vector Projections (CVP) offer a recent solution: by modeling sentiment as a direction in embedding space, they produce continuous, multilingual scores that align closely with human judgments. Yet the method's portability across...

💬 0 commentsarXiv:2601.07995v2PDF
0

Posted in cs.CL · 2026-01-12 · Nayoung Choi, Jonathan Zhang, Jinho D. Choi

DYCP: Dynamic Context Pruning for Long-Form Dialogue with LLMs

Large Language Models (LLMs) increasingly operate over long-form dialogues with frequent topic shifts. While recent LLMs support extended context windows, efficient management of dialogue history in practice is needed due to inference cost and latency constraints. We present DyCP, a lightweight context management method implemented...

💬 0 commentsarXiv:2601.07994v5PDF
0

Posted in cs.AI · 2026-01-11 · Michael Timothy Bennett

A Mind Cannot Be Smeared Across Time

Whether machines can be conscious depends not only on what they compute, but \emph{when} they compute it. Most deployed artificial systems realise their functions via sequential or time-multiplexed updates, yet a moment of conscious experience feels unified and simultaneous. I prove that this difference matters. I augment Stack Theory...

💬 0 commentsarXiv:2601.11620v2PDF
0

Posted in cs.CR · 2026-01-11 · Saleem Ishaq Tijjani, Bogdan Ghita, Nathan Clarke, Matthew Craven

Deep Recurrent Hidden Markov Learning Framework for Multi-Stage Advanced Persistent Threat Prediction

Advanced Persistent Threats (APTs) represent hidden, multi\-stage cyberattacks whose long term persistence and adaptive behavior challenge conventional intrusion detection systems (IDS). Although recent advances in machine learning and probabilistic modeling have improved APT detection performance, most existing approaches remain...

💬 0 commentsarXiv:2601.06734v2PDF