Qwen Councils
arXiv could not process that search. Try a simpler keyword search or an arXiv field query such as all:quantum.
Showing downloaded papers while arXiv is unavailable.

Computer Science

arXiv preprints from January 1, 2026 through September 24, 2026 — 00:58:34 EST

0

Posted in cs.CL · 2026-01-12 · Rei Taniguchi, Yuyang Dong, Makoto Onizuka, Chuan Xiao

Adaptive Layer Selection for Layer-Wise Token Pruning in LLM Inference

Due to the prevalence of large language models (LLMs), key-value (KV) cache reduction for LLM inference has received remarkable attention. Among numerous works that have been proposed in recent years, layer-wise token pruning approaches, which select a subset of tokens at particular layers to retain in KV cache and prune others, are...

💬 0 commentsarXiv:2601.07667v2PDF
0

Posted in cs.CV · 2026-01-12 · Dang Dinh Nguyen, Decky Aspandi Latif, Titus Zaharia

Variational Contrastive Learning for Skeleton-based Action Recognition

In recent years, self-supervised representation learning for skeleton-based action recognition has advanced with the development of contrastive learning methods. However, most of contrastive paradigms are inherently discriminative and often struggle to capture the variability and uncertainty intrinsic to human motion. To address this...

💬 0 commentsarXiv:2601.07666v1PDF
0

Posted in cs.AI · 2026-01-12 · William Walden, Miriam Wanner

Reasoning Models Will Sometimes Lie About Their Reasoning

Hint-based faithfulness evaluations have established that Large Reasoning Models (LRMs) may not say what they think: they do not always volunteer information about how key parts of the input (e.g. answer hints) influence their reasoning. Yet, these evaluations also fail to specify what models should do when confronted with hints or...

💬 0 commentsarXiv:2601.07663v4PDF
0

Posted in cs.CV · 2026-01-12 · Yuze He, Yanning Zhou, Wang Zhao, Jingwen Ye, Zhongkai Wu, Ran Yi, Yong-Jin Liu

StdGEN++: A Comprehensive System for Semantic-Decomposed 3D Character Generation

We present StdGEN++, a novel and comprehensive system for generating high-fidelity, semantically decomposed 3D characters from diverse inputs. Existing 3D generative methods often produce monolithic meshes that lack the structural flexibility required by industrial pipelines in gaming and animation. Addressing this gap, StdGEN++ is...

💬 0 commentsarXiv:2601.07660v1PDF
0

Posted in cs.CR · 2026-01-12 · Elliot Jones, William Knottenbelt

Towards Automating Blockchain Consensus Verification with IsabeLLM

Consensus protocols are crucial for a blockchain system as they are what allow agreement between the system's nodes in a potentially adversarial environment. For this reason, it is paramount to ensure their correct design and implementation to prevent such adversaries from carrying out malicious behaviour. Formal verification allows...

💬 0 commentsarXiv:2601.07654v1PDF
0

Posted in cs.AI · 2026-01-12 · Marc Lanctot, Kate Larson, Ian Gemp, Michael Kaisers

Active Evaluation of General Agents: Problem Definition and Comparison of Baseline Algorithms

As intelligent agents become more generally-capable, i.e. able to master a wide variety of tasks, the complexity and cost of properly evaluating them rises significantly. Tasks that assess specific capabilities of the agents can be correlated and stochastic, requiring many samples for accurate comparisons, leading to added costs. In...

💬 0 commentsarXiv:2601.07651v2PDF
0

Posted in cs.CL · 2026-01-12 · Jing Yang, Nils Feldhus, Salar Mohtaj, Leonhard Hennig, Qianli Wang, Eleni Metheniti, Sherzod Hakimov, Charlott Jakob, Veronika Solopova, Konrad Rieck, David Schlangen, Sebastian Möller, Vera Schmitt

What Are We Measuring in NLG? A Meta-Analysis of Evaluation Trends 2020-2025

As Natural Language Generation (NLG) dominates modern NLP, scalable evaluation remains a critical bottleneck. Consequently, LLM-as-a-judge (LaaJ) adoption has accelerated rapidly, appearing in more papers than human evaluation in 2025. This pivotal shift motivates a critical analysis of current evaluation practices. Overcoming the...

💬 0 commentsarXiv:2601.07648v2PDF
0

Posted in cs.CL · 2026-01-12 · Zijing Wang, Yongkang Liu, Mingyang Wang, Ercong Nie, Deyuan Chen, Zhengjie Zhao, Shi Feng, Daling Wang, Xiaocui Yang, Yifei Zhang, Hinrich Schütze

PlaM: Training-Free Plateau-Guided Model Merging for Better Visual Grounding in MLLMs

Multimodal Large Language Models (MLLMs) rely on strong linguistic reasoning inherited from their base language models. However, multimodal instruction fine-tuning paradoxically degrades this text's reasoning capability, undermining multimodal performance. To address this issue, we propose a training-free framework to mitigate this...

💬 0 commentsarXiv:2601.07645v1PDF
0

Posted in cs.CR · 2026-01-12 · Eckehard Hermann, Harald Lampesberger

Hagenberg Risk Management Process (Part 1): Multidimensional Polar Heatmaps for Context-Sensitive Risk Analysis

Traditional two-dimensional risk matrices (heatmaps) are widely used to model and visualize likelihood and impact relationships, but they face fundamental methodological limitations when applied to complex infrastructures. In particular, regulatory frameworks such as NIS2 and DORA call for more context-sensitive and system-oriented...

💬 0 commentsarXiv:2601.07644v1PDF
0

Posted in cs.AI · 2026-01-12 · Jiaxuan Lu, Ziyu Kong, Yemin Wang, Rong Fu, Haiyuan Wan, Cheng Yang, Wenjie Lou, Haoran Sun, Lilong Wang, Yankai Jiang, Xiaosong Wang, Xiao Sun, Dongzhan Zhou

Beyond Static Tools: Test-Time Tool Evolution for Scientific Reasoning

The central challenge of AI for Science is not reasoning alone, but the ability to create computational methods in an open-ended scientific world. Existing LLM-based agents rely on static, pre-defined tool libraries, a paradigm that fundamentally fails in scientific domains where tools are sparse, heterogeneous, and intrinsically...

💬 0 commentsarXiv:2601.07641v1PDF
0

Posted in cs.AI · 2026-01-12 · Isaiah Onando Mulang, Felix Sasaki, Tassilo Klein, Jonas Kolk, Nikolay Grechanov, Johannes Hoffart

SALT-KG: A Benchmark for Semantics-Aware Learning on Enterprise Tables

Building upon the SALT benchmark for relational prediction (Klein et al., 2024), we introduce SALT-KG, a benchmark for semantics-aware learning on enterprise tables. SALT-KG extends SALT by linking its multi-table transactional data with a structured Operational Business Knowledge represented in a Metadata Knowledge Graph (OBKG) that...

💬 0 commentsarXiv:2601.07638v1PDF
0

Posted in cs.MS · 2026-01-12 · Jan Brandejs, Niklas Hörnblad, Edward F. Valeev, Alexander Heinecke, Jeff Hammond, Devin Matthews, Paolo Bientinesi

Tensor Algebra Processing Primitives (TAPP): Towards a Standard for Tensor Operations

To address the absence of a universal standard interface for tensor operations, we introduce the Tensor Algebra Processing Primitives (TAPP), a C-based interface designed to decouple the application layer from hardware-specific implementations. We provide a mathematical formulation of tensor contractions and a reference implementation...

💬 0 commentsarXiv:2601.07827v1PDF
0

Posted in cs.RO · 2026-01-12 · Huanyu Li, Kun Lei, Sheng Zang, Kaizhe Hu, Yongyuan Liang, Bo An, Xiaoli Li, Huazhe Xu

Failure-Aware RL: Reliable Offline-to-Online Reinforcement Learning with Self-Recovery for Real-World Manipulation

Post-training algorithms based on deep reinforcement learning can push the limits of robotic models for specific objectives, such as generalizability, accuracy, and robustness. However, Intervention-requiring Failures (IR Failures) (e.g., a robot spilling water or breaking fragile glass) during real-world exploration happen...

💬 0 commentsarXiv:2601.07821v1PDF
0

Posted in cs.CL · 2026-01-12 · Manar Ali, Judith Sieker, Sina Zarrieß, Hendrik Buschmeier

Reference Games as a Testbed for the Alignment of Model Uncertainty and Clarification Requests

In human conversation, both interlocutors play an active role in maintaining mutual understanding. When listeners are uncertain about what speakers mean, for example, they can request clarification. It is an open question for language models whether they can assume a similar listener role, recognizing and expressing their own...

💬 0 commentsarXiv:2601.07820v2PDF
0

Posted in cs.RO · 2026-01-12 · Francisco Leiva, Claudio Canales, Michelle Valenzuela, Javier Ruiz-del-Solar

Data-driven control of hydraulic impact hammers under strict operational and control constraints

This paper presents a data-driven methodology for the control of static hydraulic impact hammers, also known as rock breakers, which are commonly used in the mining industry. The task addressed in this work is that of controlling the rock-breaker so its end-effector reaches arbitrary target poses, which is required in normal operation...

💬 0 commentsarXiv:2601.07813v1PDF
0

Posted in cs.CV · 2026-01-12 · Anurag Das, Adrian Bulat, Alberto Baldrati, Ioannis Maniadis Metaxas, Bernt Schiele, Georgios Tzimiropoulos, Brais Martinez

More Images, More Problems? A Controlled Analysis of VLM Failure Modes

Large Vision Language Models (LVLMs) have demonstrated remarkable capabilities, yet their proficiency in understanding and reasoning over multiple images remains largely unexplored. While existing benchmarks have initiated the evaluation of multi-image models, a comprehensive analysis of their core weaknesses and their causes is still...

💬 0 commentsarXiv:2601.07812v1PDF
0

Posted in cs.CV · 2026-01-12 · Thomas Snyder, H. Lexie Yang, Stefan Schnake, Steffen Schotthöfer

Compressing Vision Transformers in Geospatial Transfer Learning with Manifold-Constrained Optimization

Deploying geospatial foundation models on resource-constrained edge devices demands compact architectures that maintain high downstream performance. However, their large parameter counts and the accuracy loss often induced by compression limit practical adoption. In this work, we leverage manifold-constrained optimization framework...

💬 0 commentsarXiv:2601.08882v1PDF
0

Posted in cs.CL · 2026-01-12 · Ahmed Sabir, Markus Kängsepp, Rajesh Sharma

The Confidence Trap: Gender Bias and Predictive Certainty in LLMs

The increased use of Large Language Models (LLMs) in sensitive domains leads to growing interest in how their confidence scores correspond to fairness and bias. This study examines the alignment between LLM-predicted confidence and human-annotated bias judgments. Focusing on gender bias, the research investigates probability...

💬 0 commentsarXiv:2601.07806v1PDF
0

Posted in cs.CV · 2026-01-12 · Sijun Dong, Siming Fu, Kaiyu Li, Xiangyong Cao, Xiaoliang Meng, Bo Du

Exchange Is All You Need for Remote Sensing Change Detection

Remote sensing change detection fundamentally relies on the effective fusion and discrimination of bi-temporal features. Prevailing paradigms typically utilize Siamese encoders bridged by explicit difference computation modules, such as subtraction or concatenation, to identify changes. In this work, we challenge this complexity with...

💬 0 commentsarXiv:2601.07805v1PDF
0

Posted in cs.CL · 2026-01-12 · Jiongchi Yu, Yuhan Ma, Xiaoyu Zhang, Junjie Wang, Qiang Hu, Chao Shen, Xiaofei Xie

PTCBENCH: Benchmarking Contextual Stability of Personality Traits in LLM Systems

With the increasing deployment of large language models (LLMs) in affective agents and AI systems, maintaining a consistent and authentic LLM personality becomes critical for user trust and engagement. However, existing work overlooks a fundamental psychological consensus that personality traits are dynamic and context-dependent. To...

💬 0 commentsarXiv:2602.00016v1PDF
0

Posted in cs.IT · 2026-01-12 · Yiqi Chen, Holger Boche, Marc Geitz

Lossy Source Coding with Broadcast Side Information

This paper considers the source coding problem with broadcast side information. The side information is sent to two receivers through a noisy broadcast channel. We provide an outer bound of the rate--distortion--bandwidth (RDB) quadruples and achievable RDB quadruples when the helper uses a separation-based scheme. Some special cases...

💬 0 commentsarXiv:2601.07797v2PDF
0

Posted in cs.CL · 2026-01-12 · Shaz Furniturewala, Gerard Christopher Yeo, Kokil Jaidka

Learning Through Dialogue: Engagement and Efficacy Matter More Than Explanations

Large language models (LLMs) are increasingly used as conversational partners for learning, yet the interactional dynamics supporting users' learning and engagement are understudied. We analyze the linguistic and interactional features from both LLM and participant chats across 397 human-LLM conversations about socio-political issues...

💬 0 commentsarXiv:2601.07796v2PDF
0

Posted in cs.CV · 2026-01-12 · Patrick Bauer, Marius Schwinning, Florian Renk, Andreas Weinmann, Hichem Snoussi

Vision-Language Model for Accurate Crater Detection

The European Space Agency (ESA), driven by its ambitions on planned lunar missions with the Argonaut lander, has a profound interest in reliable crater detection, since craters pose a risk to safe lunar landings. This task is usually addressed with automated crater detection algorithms (CDA) based on deep learning techniques. It is...

💬 0 commentsarXiv:2601.07795v1PDF
0

Posted in cs.CL · 2026-01-12 · Tianda Sun, Dimitar Kazakov

Kinship Data Benchmark for Multi-hop Reasoning

Large language models (LLMs) are increasingly evaluated on their ability to perform multi-hop reasoning, i.e., to combine multiple pieces of information into a coherent inference. We introduce KinshipQA, a benchmark designed to probe this capability through reasoning over kinship relations. The central contribution of our work is a...

💬 0 commentsarXiv:2601.07794v1PDF
0

Posted in cs.AI · 2026-01-12 · Yahya Masri, Emily Ma, Zifu Wang, Joseph Rogers, Chaowei Yang

Benchmarking Small Language Models and Small Reasoning Language Models on System Log Severity Classification

System logs are crucial for monitoring and diagnosing modern computing infrastructure, but their scale and complexity require reliable and efficient automated interpretation. Since severity levels are predefined metadata in system log messages, having a model merely classify them offers limited standalone practical value, revealing...

💬 0 commentsarXiv:2601.07790v1PDF