Qwen Councils
arXiv could not process that search. Try a simpler keyword search or an arXiv field query such as all:quantum.
Showing downloaded papers while arXiv is unavailable.

Computer Science

arXiv preprints from January 1, 2026 through September 24, 2026 — 12:26:29 EST

0

Posted in cs.CL · 2026-01-09 · Liu Zai, Iraklis Klampanos

Peek2: Regex-free Byte-level Byte-Pair Encoding Pretokenizer for LLM Inference on Edge Devices

Pretokenization is a crucial, sequential pass in Byte-level BPE tokenizers, yet little work has been done to optimize it for edge-side inference. Our proposed new implementation, Peek2, serves as a drop-in replacement for cl100k-like pretokenizers used in GPT-3, LLaMa-3, and Qwen-2.5. After breaking down and analyzing the logic of the...

💬 0 commentsarXiv:2601.05833v2PDF
0

Posted in cs.AI · 2026-01-09 · Weijie Li, Zhongqing Wang, Guodong Zhou

PCoKG: Personality-aware Commonsense Reasoning with Debate

Most commonsense reasoning models overlook the influence of personality traits, limiting their effectiveness in personalized systems such as dialogue generation. To address this limitation, we introduce the Personality-aware Commonsense Knowledge Graph (PCoKG), a structured dataset comprising 521,316 quadruples. We begin by employing...

💬 0 commentsarXiv:2601.06234v1PDF
0

Posted in cs.CR · 2026-01-09 · Manuel Brosch, Matthias Probst, Stefan Kögler, Georg Sigl

Influence of Parallelism in Vector-Multiplication Units on Correlation Power Analysis

The use of neural networks in edge devices is increasing, which introduces new security challenges related to the neural networks' confidentiality. As edge devices often offer physical access, attacks targeting the hardware, such as side-channel analysis, must be considered. To enhance the performance of neural network inference,...

💬 0 commentsarXiv:2601.05828v1PDF
0

Posted in cs.SE · 2026-01-09 · Zewei Lin, Jiachi Chen, Jingwen Zhang, Zexu Wang, Yuming Feng, Weizhe Zhang, Zibin Zheng

SSR: Safeguarding Staking Rewards by Defining and Detecting Logical Defects in DeFi Staking

Decentralized Finance (DeFi) staking is one of the most prominent applications within the DeFi ecosystem, where DeFi projects enable users to stake tokens on the platform and reward participants with additional tokens. However, logical defects in DeFi staking could enable attackers to claim unwarranted rewards by manipulating reward...

💬 0 commentsarXiv:2601.05827v1PDF
0

Posted in cs.CY · 2026-01-09 · Íris Damião, João Franco, Mariana Silva, Paulo Almeida, Pedro C. Magalhães, Joana Gonçalves-Sá

Cross-National Evidence of Disproportionate Media Visibility for the Radical Right in the 2024 European Elections

This study provides a systematic comparative analysis of media visibility of different political families during the 2024 European Parliament elections. We analyzed close to 21,500 unique news from leading national outlets in Austria, Germany, Ireland, Poland, and Portugal - countries with diverse political contexts and levels of...

💬 0 commentsarXiv:2601.05826v1PDF
0

Posted in cs.HC · 2026-01-09 · Lucija Mihić Zidar, Philipp Wicke, Praneel Bhatia, Rosa Lutz, Marius Klug, Thorsten O. Zander

Decoding Workload and Agreement From EEG During Spoken Dialogue With Conversational AI

Passive brain-computer interfaces offer a potential source of implicit feedback for alignment of large language models, but most mental state decoding has been done in controlled tasks. This paper investigates whether established EEG classifiers for mental workload and implicit agreement can be transferred to spoken human-AI dialogue....

💬 0 commentsarXiv:2601.05825v2PDF
0

Posted in cs.CV · 2026-01-09 · John Page, Xuesong Niu, Kai Wu, Kun Gai

Boosting Latent Diffusion Models via Disentangled Representation Alignment

Latent Diffusion Models (LDMs) rely heavily on the compressed latent space provided by Variational Autoencoders (VAEs) for high-quality image generation. Recent studies have attempted to obtain generation-friendly VAEs by directly adopting alignment strategies from LDM training, leveraging Vision Foundation Models (VFMs) as...

💬 0 commentsarXiv:2601.05823v2PDF
0

Posted in cs.HC · 2026-01-09 · Adarsh Pawar, Yuqiao Meng, Luoxi Tang, Zhaohan Xi

Improving Clinical Data Accessibility Through Automated FHIR Data Transformation Tools

The Fast Healthcare Interoperability Resources (FHIR) standard has emerged as a widely adopted specification for exchanging structured clinical data across healthcare systems. However, raw FHIR resources are often complex, verbose, and difficult for clinicians and analysts to interpret without specialized tooling. This paper presents...

💬 0 commentsarXiv:2601.05822v2PDF
0

Posted in cs.CL · 2026-01-09 · Milad Alshomary, Grace Li, Anubhav Jangra, Yufang Hou, Kathleen McKeown, Smaranda Muresan

LLMs as Science Journalists: Supporting Early-stage Researchers in Communicating Their Science to the Public

The scientific community needs tools that help early-stage researchers effectively communicate their findings and innovations to the public. Although existing general-purpose Large Language Models (LLMs) can assist in this endeavor, they are not optimally aligned for it. To address this, we propose a framework for training LLMs to...

💬 0 commentsarXiv:2601.05821v1PDF
0

Posted in cs.DC · 2026-01-09 · Shiting Long, Gustavo Ramirez-Hidalgo, Stepan Nassyr, Jose Jimenez-Merchan, Andreas Frommer, Dirk Pleiter

Performance-Portable Optimization and Analysis of Multiple Right-Hand Sides in a Lattice QCD Solver

Managing the high computational cost of iterative solvers for sparse linear systems is a known challenge in scientific computing. Moreover, scientific applications often face memory bandwidth constraints, making it critical to optimize data locality and enhance the efficiency of data transport. We extend the lattice QCD solver...

💬 0 commentsarXiv:2601.05816v1PDF
0

Posted in cs.LG · 2026-01-09 · Md Sultanul Islam Ovi, Muhsina Tarannum Munfa, G. M. M Miftahul Alam Adib, Syed Sabbir Hasan

A Dual Pipeline Machine Learning Framework for Automated Multi Class Sleep Disorder Screening Using Hybrid Resampling and Ensemble Learning

Accurate classification of sleep disorders, particularly insomnia and sleep apnea, is important for reducing long term health risks and improving patient quality of life. However, clinical sleep studies are resource intensive and are difficult to scale for population level screening. This paper presents a Dual Pipeline Machine...

💬 0 commentsarXiv:2601.05814v2PDF
0

Posted in cs.DB · 2026-01-09 · Enrique Feito-Casares, Ismael Gómez-Talal, José-Luis Rojo-Álvarez

Descriptor: Multi-Regional Cloud Honeypot Dataset (MURHCAD)

This data article introduces a comprehensive, high-resolution honeynet dataset designed to support standalone analyses of global cyberattack behaviors. Collected over a continuous 72-hour window (June 9 to 11, 2025) on Microsoft Azure, the dataset comprises 132,425 individual attack events captured by three honeypots (Cowrie, Dionaea,...

💬 0 commentsarXiv:2601.05813v1PDF
0

Posted in cs.LG · 2026-01-09 · Zhanpei Huang, Taochen chen, Fangqing Gu, Yiqun Zhang

Detecting Autism Spectrum Disorder with Deep Eye Movement Features

Autism Spectrum Disorder (ASD) is a neurodevelopmental disorder characterized by deficits in social communication and behavioral patterns. Eye movement data offers a non-invasive diagnostic tool for ASD detection, as it is inherently discrete and exhibits short-term temporal dependencies, reflecting localized gaze focus between...

💬 0 commentsarXiv:2601.05812v1PDF
0

Posted in cs.LG · 2026-01-09 · Enrique Feito-Casares, Francisco M. Melgarejo-Meseguer, José-Luis Rojo-Álvarez

Learning Reconstructive Embeddings in Reproducing Kernel Hilbert Spaces via the Representer Theorem

Motivated by the growing interest in representation learning approaches that uncover the latent structure of high-dimensional data, this work proposes new algorithms for reconstruction-based manifold learning within Reproducing-Kernel Hilbert Spaces (RKHS). Each observation is first reconstructed as a linear combination of the other...

💬 0 commentsarXiv:2601.05811v1PDF
0

Posted in cs.CV · 2026-01-09 · ChunTeng Chen, YiChen Hsu, YiWen Liu, WeiFang Sun, TsaiChing Ni, ChunYi Lee, Min Sun, YuanFu Yang

SceneFoundry: Generating Interactive Infinite 3D Worlds

The ability to automatically generate large-scale, interactive, and physically realistic 3D environments is crucial for advancing robotic learning and embodied intelligence. However, existing generative approaches often fail to capture the functional complexity of real-world interiors, particularly those containing articulated objects...

💬 0 commentsarXiv:2601.05810v2PDF
0

Posted in cs.CL · 2026-01-09 · Xiaoshuai Song, Haofei Chang, Guanting Dong, Yutao Zhu, Ji-Rong Wen, Zhicheng Dou

EnvScaler: Scaling Tool-Interactive Environments for LLM Agent via Programmatic Synthesis

Large language models (LLMs) are expected to be trained to act as agents in various real-world environments, but this process relies on rich and varied tool-interaction sandboxes. However, access to real systems is often restricted; LLM-simulated environments are prone to hallucinations and inconsistencies; and manually built...

💬 0 commentsarXiv:2601.05808v2PDF
0

Posted in cs.LG · 2026-01-09 · Mohamed Amine Hallam, Kuo-Kun Tseng

Fusion Matters: Length-Aware Analysis of Positional-Encoding Fusion in Transformers

Transformers require positional encodings to represent sequence order, yet most prior work focuses on designing new positional encodings rather than examining how positional information is fused with token embeddings. In this paper, we study whether the fusion mechanism itself affects performance, particularly in long-sequence...

💬 0 commentsarXiv:2601.05807v1PDF
0

Posted in cs.RO · 2026-01-09 · Marvin Seegert, Korbinian Moller, Johannes Betz

Modular Autonomy with Conversational Interaction: An LLM-driven Framework for Decision Making in Autonomous Driving

Recent advancements in Large Language Models (LLMs) offer new opportunities to create natural language interfaces for Autonomous Driving Systems (ADSs), moving beyond rigid inputs. This paper addresses the challenge of mapping the complexity of human language to the structured action space of modular ADS software. We propose a...

💬 0 commentsarXiv:2601.05806v1PDF
0

Posted in cs.RO · 2026-01-09 · Simon Archieri, Ahmet Cinar, Shu Pan, Jonatan Scharff Willners, Michele Grimaldi, Ignacio Carlucho, Yvan Petillot

InsSo3D: Inertial Navigation System and 3D Sonar SLAM for turbid environment inspection

This paper presents InsSo3D, an accurate and efficient method for large-scale 3D Simultaneous Localisation and Mapping (SLAM) using a 3D Sonar and an Inertial Navigation System (INS). Unlike traditional sonar, which produces 2D images containing range and azimuth information but lacks elevation information, 3D Sonar produces a 3D...

💬 0 commentsarXiv:2601.05805v2PDF
0

Posted in cs.CL · 2026-01-09 · Eilam Cohen, Itamar Bul, Danielle Inbar, Omri Loewenbach

Simplify-This: A Comparative Analysis of Prompt-Based and Fine-Tuned LLMs

Large language models (LLMs) enable strong text generation, and in general there is a practical tradeoff between fine-tuning and prompt engineering. We introduce Simplify-This, a comparative study evaluating both paradigms for text simplification with encoder-decoder LLMs across multiple benchmarks, using a range of evaluation...

💬 0 commentsarXiv:2601.05794v1PDF
0

Posted in cs.LG · 2026-01-09 · Manel Gil-Sorribes, Júlia Vilalta-Mor, Isaac Filella-Mercè, Robert Soliva, Álvaro Ciudad, Víctor Guallar, Alexis Molina

Tensor-DTI: Enhancing Biomolecular Interaction Prediction with Contrastive Embedding Learning

Accurate drug-target interaction (DTI) prediction is essential for computational drug discovery, yet existing models often rely on single-modality predefined molecular descriptors or sequence-based embeddings with limited representativeness. We propose Tensor-DTI, a contrastive learning framework that integrates multimodal embeddings...

💬 0 commentsarXiv:2601.05792v1PDF
0

Posted in cs.HC · 2026-01-09 · Tianwang Jia, Xiaoqing Chen, Dongrui Wu

SAFE: Secure and Accurate Federated Learning for Privacy-Preserving Brain-Computer Interfaces

Electroencephalogram (EEG)-based brain-computer interfaces (BCIs) are widely adopted due to their efficiency and portability; however, their decoding algorithms still face multiple challenges, including inadequate generalization, adversarial vulnerability, and privacy leakage. This paper proposes Secure and Accurate FEderated learning...

💬 0 commentsarXiv:2601.05789v1PDF
0

Posted in cs.AI · 2026-01-09 · Zezhou Wang, Ziyun Zhang, Xiaoyi Zhang, Zhuzhong Qian, Yan Lu

From Off-Policy to On-Policy: Enhancing GUI Agents via Bi-level Expert-to-Policy Assimilation

Vision-language models are increasingly deployed as computer-use agents (CUAs) that operate desktops and browsers. Top-performing CUAs are framework-based systems that decompose planning and execution, while end-to-end screenshot-to-action policies are easier to deploy but lag behind on benchmarks such as OSWorld-Verified. GUI...

💬 0 commentsarXiv:2601.05787v2PDF
0

Posted in cs.CV · 2026-01-09 · Quanjiang Li, Zhiming Liu, Tianxiang Xu, Tingjin Luo, Chenping Hou

Adaptive Disentangled Representation Learning for Incomplete Multi-View Multi-Label Classification

Multi-view multi-label learning frequently suffers from simultaneous feature absence and incomplete annotations, due to challenges in data acquisition and cost-intensive supervision. To tackle the complex yet highly practical problem while overcoming the existing limitations of feature recovery, representation disentanglement, and...

💬 0 commentsarXiv:2601.05785v1PDF
0

Posted in cs.SE · 2026-01-09 · Yaoqi Guo, Ying Xiao, Jie M. Zhang, Mark Harman, Yiling Lou, Yang Liu, Zhenpeng Chen

EET: Experience-Driven Early Termination for Cost-Efficient Software Engineering Agents

Software engineering (SE) agents powered by large language models are increasingly adopted in practice, yet they often incur substantial monetary cost. We introduce EET, an experience-driven early termination approach that reduces the cost of SE agents while preserving task performance. EET extracts structured experience from prior...

💬 0 commentsarXiv:2601.05777v2PDF