Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through July 28, 2026 — 09:22:02 EST

0

Posted in cs.RO · 2026-01-09 · Simon Archieri, Ahmet Cinar, Shu Pan, Jonatan Scharff Willners, Michele Grimaldi, Ignacio Carlucho, Yvan Petillot

InsSo3D: Inertial Navigation System and 3D Sonar SLAM for turbid environment inspection

This paper presents InsSo3D, an accurate and efficient method for large-scale 3D Simultaneous Localisation and Mapping (SLAM) using a 3D Sonar and an Inertial Navigation System (INS). Unlike traditional sonar, which produces 2D images containing range and azimuth information but lacks elevation information, 3D Sonar produces a 3D...

💬 0 commentsarXiv:2601.05805v2PDF
0

Posted in cs.CL · 2026-01-09 · Eilam Cohen, Itamar Bul, Danielle Inbar, Omri Loewenbach

Simplify-This: A Comparative Analysis of Prompt-Based and Fine-Tuned LLMs

Large language models (LLMs) enable strong text generation, and in general there is a practical tradeoff between fine-tuning and prompt engineering. We introduce Simplify-This, a comparative study evaluating both paradigms for text simplification with encoder-decoder LLMs across multiple benchmarks, using a range of evaluation...

💬 0 commentsarXiv:2601.05794v1PDF
0

Posted in cs.LG · 2026-01-09 · Manel Gil-Sorribes, Júlia Vilalta-Mor, Isaac Filella-Mercè, Robert Soliva, Álvaro Ciudad, Víctor Guallar, Alexis Molina

Tensor-DTI: Enhancing Biomolecular Interaction Prediction with Contrastive Embedding Learning

Accurate drug-target interaction (DTI) prediction is essential for computational drug discovery, yet existing models often rely on single-modality predefined molecular descriptors or sequence-based embeddings with limited representativeness. We propose Tensor-DTI, a contrastive learning framework that integrates multimodal embeddings...

💬 0 commentsarXiv:2601.05792v1PDF
0

Posted in cs.HC · 2026-01-09 · Tianwang Jia, Xiaoqing Chen, Dongrui Wu

SAFE: Secure and Accurate Federated Learning for Privacy-Preserving Brain-Computer Interfaces

Electroencephalogram (EEG)-based brain-computer interfaces (BCIs) are widely adopted due to their efficiency and portability; however, their decoding algorithms still face multiple challenges, including inadequate generalization, adversarial vulnerability, and privacy leakage. This paper proposes Secure and Accurate FEderated learning...

💬 0 commentsarXiv:2601.05789v1PDF
0

Posted in cs.AI · 2026-01-09 · Zezhou Wang, Ziyun Zhang, Xiaoyi Zhang, Zhuzhong Qian, Yan Lu

From Off-Policy to On-Policy: Enhancing GUI Agents via Bi-level Expert-to-Policy Assimilation

Vision-language models are increasingly deployed as computer-use agents (CUAs) that operate desktops and browsers. Top-performing CUAs are framework-based systems that decompose planning and execution, while end-to-end screenshot-to-action policies are easier to deploy but lag behind on benchmarks such as OSWorld-Verified. GUI...

💬 0 commentsarXiv:2601.05787v2PDF
0

Posted in cs.CV · 2026-01-09 · Quanjiang Li, Zhiming Liu, Tianxiang Xu, Tingjin Luo, Chenping Hou

Adaptive Disentangled Representation Learning for Incomplete Multi-View Multi-Label Classification

Multi-view multi-label learning frequently suffers from simultaneous feature absence and incomplete annotations, due to challenges in data acquisition and cost-intensive supervision. To tackle the complex yet highly practical problem while overcoming the existing limitations of feature recovery, representation disentanglement, and...

💬 0 commentsarXiv:2601.05785v1PDF
0

Posted in cs.SE · 2026-01-09 · Yaoqi Guo, Ying Xiao, Jie M. Zhang, Mark Harman, Yiling Lou, Yang Liu, Zhenpeng Chen

EET: Experience-Driven Early Termination for Cost-Efficient Software Engineering Agents

Software engineering (SE) agents powered by large language models are increasingly adopted in practice, yet they often incur substantial monetary cost. We introduce EET, an experience-driven early termination approach that reduces the cost of SE agents while preserving task performance. EET extracts structured experience from prior...

💬 0 commentsarXiv:2601.05777v2PDF
0

Posted in cs.CL · 2026-01-09 · Benedikt Ebing, Lennart Keller, Goran Glavaš

One Script Instead of Hundreds? On Pretraining Romanized Encoder Language Models

Exposing latent lexical overlap, script romanization has emerged as an effective strategy for improving cross-lingual transfer (XLT) in multilingual language models (mLMs). Most prior work, however, focused on setups that favor romanization the most: (1) transfer from high-resource Latin-script to low-resource non-Latin-script...

💬 0 commentsarXiv:2601.05776v1PDF
0

Posted in cs.SE · 2026-01-09 · Qingyuan Li, Chenchen Yu, Chuanyi Li, Xin-Cheng Wen, Cheryl Lee, Cuiyun Gao, Bin Luo

StriderSPD: Structure-Guided Joint Representation Learning for Binary Security Patch Detection

Vulnerabilities severely threaten software systems, making the timely application of security patches crucial for mitigating attacks. However, software vendors often silently patch vulnerabilities with limited disclosure, where Security Patch Detection (SPD) comes to protect software assets. Recently, most SPD studies have targeted...

💬 0 commentsarXiv:2601.05772v1PDF
0

Posted in cs.LG · 2026-01-09 · Yifan Zhang, Wei Bi, Kechi Zhang, Dongming Jin, Jie Fu, Zhi Jin

Weights to Code: Extracting Interpretable Algorithms from the Discrete Transformer

Algorithm extraction aims to synthesize executable programs directly from models trained on algorithmic tasks, enabling de novo recovery of executable mechanisms from weights without relying on human-written target programs. However, applying this paradigm to Transformer is complicated by representation entanglement (e.g.,...

💬 0 commentsarXiv:2601.05770v3PDF
0

Posted in cs.LG · 2026-01-09 · Nina Peire, Yupei Li, Björn Schuller

Affect and Effect: Limitations of regularisation-based continual learning in EEG-based emotion classification

Generalisation to unseen subjects in EEG-based emotion classification remains a challenge due to high inter-and intra-subject variability. Continual learning (CL) poses a promising solution by learning from a sequence of tasks while mitigating catastrophic forgetting. Regularisation-based CL approaches, such as Elastic Weight...

💬 0 commentsarXiv:2601.07858v1PDF
0

Posted in cs.CR · 2026-01-09 · Chandra Sekhar Kubam

Agentic AI Microservice Framework for Deepfake and Document Fraud Detection in KYC Pipelines

The rapid proliferation of synthetic media, presentation attacks, and document forgeries has created significant vulnerabilities in Know Your Customer (KYC) workflows across financial services, telecommunications, and digital-identity ecosystems. Traditional monolithic KYC systems lack the scalability and agility required to counter...

💬 0 commentsarXiv:2601.06241v1PDF
0

Posted in cs.CY · 2026-01-09 · Samuel Gerald Collins

Stuck in the Turing Matrix: Inauthenticity, Deception and the Social Life of AI

The Turing test may or may not be a valid test of machine intelligence. But in an age of generative AI, the test describes the positions we humans occupy. Judging whether or not something is human or machine produced is an everyday condition for many of us, one that involves taking a spectrum of positions along what the essay...

💬 0 commentsarXiv:2601.11613v1PDF
0

Posted in cs.SE · 2026-01-09 · James Uther

Putting green software principles into practice

The need and theoretical methods for measuring and reducing CO2 emitted by computing systems are well understood, but real-world examples are still limited. We describe a journey towards green software for a live product running on a public cloud. We discuss practical solutions found, in particular using the cost implications of...

💬 0 commentsarXiv:2601.09741v1PDF
0

Posted in cs.CV · 2026-01-09 · Chanchan Wang, Yuanfang Wang, Qing Xu, Guanxin Chen

WaveRNet: Wavelet-Guided Frequency Learning for Multi-Source Domain-Generalized Retinal Vessel Segmentation

Domain-generalized retinal vessel segmentation is critical for automated ophthalmic diagnosis, yet faces significant challenges from domain shift induced by non-uniform illumination and varying contrast, compounded by the difficulty of preserving fine vessel structures. While the Segment Anything Model (SAM) exhibits remarkable...

💬 0 commentsarXiv:2601.05942v1PDF
0

Posted in cs.CV · 2026-01-09 · Mehrdad Fazli, Bowen Wei, Ziwei Zhu

Context-Aware Decoding for Faithful Vision-Language Generation

Hallucinations, generating responses inconsistent with the visual input, remain a critical limitation of large vision-language models (LVLMs), especially in open-ended tasks such as image captioning and visual reasoning. In this work, we probe the layer-wise generation dynamics that drive hallucinations and propose a training-free...

💬 0 commentsarXiv:2601.05939v1PDF
0

Posted in cs.CV · 2026-01-09 · Pankaj Gupta, Priya Mudgil, Niharika Dutta, Kartik Bose, Nitish Kumar, Anupam Kumar, Jimil Shah, Vaneet Jearth, Jayanta Samanta, Vishal Sharma, Harshal Mandavdhare, Surinder Rana, Saroj K Sinha, Usha Dutta

Performance of a Deep Learning-Based Segmentation Model for Pancreatic Tumors on Public Endoscopic Ultrasound Datasets

Background: Pancreatic cancer is one of the most aggressive cancers, with poor survival rates. Endoscopic ultrasound (EUS) is a key diagnostic modality, but its effectiveness is constrained by operator subjectivity. This study evaluates a Vision Transformer-based deep learning segmentation model for pancreatic tumors. Methods: A...

💬 0 commentsarXiv:2601.05937v1PDF
0

Posted in cs.CL · 2026-01-09 · Jingsheng Zheng, Jintian Zhang, Yujie Luo, Yuren Mao, Yunjun Gao, Lun Du, Huajun Chen, Ningyu Zhang

Can We Predict Before Executing Machine Learning Agents?

Autonomous machine learning agents have revolutionized scientific discovery, yet they remain constrained by a Generate-Execute-Feedback paradigm. Previous approaches suffer from a severe Execution Bottleneck, as hypothesis evaluation relies strictly on expensive physical execution. To bypass these physical constraints, we internalize...

💬 0 commentsarXiv:2601.05930v2PDF
0

Posted in cs.LG · 2026-01-09 · Sidney Shapiro, Burhanuddin Panvelwala

Prophet as a Reproducible Forecasting Framework: A Methodological Guide for Business and Financial Analytics

Reproducibility remains a persistent challenge in forecasting research and practice, particularly in business and financial analytics, where forecasts inform high-stakes decisions. Traditional forecasting methods, while theoretically interpretable, often require extensive manual tuning and are difficult to replicate in proprietary...

💬 0 commentsarXiv:2601.05929v2PDF
0

Posted in cs.CL · 2026-01-09 · Nguyen Phuc Tran, Brigitte Jaumard, Oscar Delgado, Tristan Glatard, Karthikeyan Premkumar, Kun Ni

LLM-Augmented Knowledge Base Construction For Root Cause Analysis

Communications networks now form the backbone of our digital world, with fast and reliable connectivity. However, even with appropriate redundancy and failover mechanisms, it is difficult to guarantee "five 9s" (99.999 %) reliability, requiring rapid and accurate root cause analysis (RCA) during outages. In the event of an outage,...

💬 0 commentsarXiv:2604.06171v1PDF
0

Posted in cs.CV · 2026-01-09 · Yohann Perron, Vladyslav Sydorov, Christophe Pottier, Loic Landrieu

Adapting Vision Transformers to Ultra-High Resolution Semantic Segmentation with Relay Tokens

Current approaches for segmenting ultra high resolution images either slide a window, thereby discarding global context, or downsample and lose fine detail. We propose a simple yet effective method that brings explicit multi scale reasoning to vision transformers, simultaneously preserving local details and global awareness....

💬 0 commentsarXiv:2601.05927v1PDF
0

Posted in cs.CL · 2026-01-09 · Nora Graichen, Iria de-Dios-Flores, Gemma Boleda

The Grammar of Transformers: A Systematic Review of Interpretability Research on Syntactic Knowledge in Language Models

We present a systematic review of 337 articles evaluating the syntactic abilities of Transformer-based language models (TLMs), reporting on over 3,000 datapoints spanning a wide range of syntactic phenomena, languages, models, and methods. We take the data to collectively show that TLMs encode a non-trivial amount of syntactic...

💬 0 commentsarXiv:2601.19926v2PDF
0

Posted in cs.CR · 2026-01-09 · Tianshi Li

Agentic LLMs as Powerful Deanonymizers: Re-identification of Participants in the Anthropic Interviewer Dataset

On December 4, 2025, Anthropic released Anthropic Interviewer, an AI tool for running qualitative interviews at scale, along with a public dataset of 1,250 interviews with professionals, including 125 scientists, about their use of AI for research. Focusing on the scientist subset, I show that widely available LLMs with web search and...

💬 0 commentsarXiv:2601.05918v1PDF
0

Posted in cs.LG · 2026-01-09 · Pattarawat Chormai, Ali Hashemi, Klaus-Robert Müller, Grégoire Montavon

Distilling Lightweight Domain Experts from Large ML Models by Identifying Relevant Subspaces

Knowledge distillation involves transferring the predictive capabilities of large, high-performing AI models (teachers) to smaller models (students) that can operate in environments with limited computing power. In this paper, we address the scenario in which only a few classes and their associated intermediate concepts are relevant...

💬 0 commentsarXiv:2601.05913v1PDF
0

Posted in cs.CL · 2026-01-09 · Phuong-Hang Le, Valentin Pelloin, Arnault Chatelain, Maryem Bouziane, Mohammed Ghennai, Qianwen Guan, Kirill Milintsevich, Salima Mdhaffar, Aidan Mannion, Nils Defauw, Shuyue Gu, Alexandre Audibert, Marco Dinarelli, Yannick Estève, Lorraine Goeuriot, Steffen Lalande, Nicolas Hervé, Maximin Coavoux, François Portet, Étienne Ollion, Marie Candito, Maxime Peyrard, Solange Rossato, Benjamin Lecouteux, Aurélie Nardy, Gilles Sérasset, Vincent Segonne, Solène Evain, Diandra Fabre, Didier Schwab

Pantagruel: Unified Self-Supervised Encoders for French Text and Speech

We release Pantagruel models, a new family of self-supervised encoder models for French text and speech. Instead of predicting modality-tailored targets such as textual tokens or speech units, Pantagruel learns contextualized target representations in the feature space, allowing modality-specific encoders to capture linguistic and...

💬 0 commentsarXiv:2601.05911v2PDF