Qwen Councils
arXiv could not process that search. Try a simpler keyword search or an arXiv field query such as all:quantum.
Showing downloaded papers while arXiv is unavailable.

Computer Science

arXiv preprints from January 1, 2026 through September 24, 2026 — 01:14:17 EST

0

Posted in cs.HC · 2026-01-12 · Liberty Kent, Nilufer Tuptuk, Ingolf Becker

Passing the Baton: Shift Handovers within Cybersecurity Incident Response Teams

Effective shift transitions are crucial for cybersecurity incident response teams, yet there is limited guidance on managing these handovers. This exploratory study aimed to develop guidelines for such transitions through the analysis of existing literature and consultation with practitioners. Two draft guidelines (A and B) were...

💬 0 commentsarXiv:2601.07788v1PDF
0

Posted in cs.SE · 2026-01-12 · Abdullah Al Mujahid, Mia Mohammad Imran

"TODO: Fix the Mess Gemini Created": Towards Understanding GenAI-Induced Self-Admitted Technical Debt

As large language models (LLMs) such as ChatGPT, Copilot, Claude, and Gemini become integrated into software development workflows, developers increasingly leave traces of AI involvement in their code comments. Among these, some comments explicitly acknowledge both the use of generative AI and the presence of technical shortcomings....

💬 0 commentsarXiv:2601.07786v1PDF
0

Posted in cs.CL · 2026-01-12 · Mariana Costa, Alberlucia Rafael Soarez, Daniel Kim, Camila Ferreira

Enhancing Self-Correction in Large Language Models through Multi-Perspective Reflection

While Chain-of-Thought (CoT) prompting advances LLM reasoning, challenges persist in consistency, accuracy, and self-correction, especially for complex or ethically sensitive tasks. Existing single-dimensional reflection methods offer insufficient improvements. We propose MyGO Poly-Reflective Chain-of-Thought (PR-CoT), a novel...

💬 0 commentsarXiv:2601.07780v1PDF
0

Posted in cs.MA · 2026-01-12 · Bowen Yang, Kaiming Jin, Zhenyu Wu, Zhaoyang Liu, Qiushi Sun, Zehao Li, JingJing Xie, Zhoumianze Liu, Fangzhi Xu, Kanzhi Cheng, Qingyun Li, Yian Wang, Yu Qiao, Zun Wang, Zichen Ding

OS-Symphony: A Holistic Framework for Robust and Generalist Computer-Using Agent

While Vision-Language Models (VLMs) have significantly advanced Computer-Using Agents (CUAs), current frameworks struggle with robustness in long-horizon workflows and generalization in novel domains. These limitations stem from a lack of granular control over historical visual context curation and the absence of visual-aware tutorial...

💬 0 commentsarXiv:2601.07779v1PDF
0

Posted in cs.LG · 2026-01-12 · Wen Guo

DT-ICU: Towards Explainable Digital Twins for ICU Patient Monitoring via Multi-Modal and Multi-Task Iterative Inference

We introduce DT-ICU, a multimodal digital twin framework for continuous risk estimation in intensive care. DT-ICU integrates variable-length clinical time series with static patient information in a unified multitask architecture, enabling predictions to be updated as new observations accumulate over the ICU stay. We evaluate DT-ICU...

💬 0 commentsarXiv:2601.07778v1PDF
0

Posted in cs.GT · 2026-01-12 · Sarvin Bahmani, Rasmus Ibsen-Jensen, Soumyajit Paul, Sven Schewe, Friedrich Slivovsky, Qiyi Tang, Dominik Wojtczak, Shufang Zhu

The Complexity of Games with Randomised Control

We study the complexity of solving two-player infinite duration games played on a fixed finite graph, where the control of a node is not predetermined but rather assigned randomly. In classic random-turn games, control of each node is assigned randomly every time the node is visited during a play. In this work, we study two natural...

💬 0 commentsarXiv:2601.07775v1PDF
0

Posted in cs.CV · 2026-01-12 · Lingchen Sun, Rongyuan Wu, Zhengqiang Zhang, Ruibin Li, Yujing Sun, Shuaizheng Liu, Lei Zhang

Self-transcendence: Is External Feature Guidance Indispensable for Accelerating Diffusion Transformer Training?

Recent works such as REPA have shown that guiding diffusion models with external semantic features (e.g., DINO) can significantly accelerate the training of diffusion transformers (DiTs). However, the use of pretrained external features as guidance signals introduces additional dependencies. We argue that DiTs actually have the power...

💬 0 commentsarXiv:2601.07773v3PDF
0

Posted in cs.RO · 2026-01-12 · Alex Huang, Akshay Karthik

THETA: Triangulated Hand-State Estimation for Teleoperation and Automation in Robotic Hand Control

The teleoperation of robotic hands is limited by the high costs of depth cameras and sensor gloves, commonly used to estimate hand relative joint positions (XYZ). We present a novel, cost-effective approach using three webcams for triangulation-based tracking to approximate relative joint angles (theta) of human fingers. We also...

💬 0 commentsarXiv:2601.07768v1PDF
0

Posted in cs.LG · 2026-01-12 · Jiawei Wang, Yanfei Zhou, Siddartha Devic, Deqing Fu

Are LLM Decisions Faithful to Verbal Confidence?

Large Language Models (LLMs) can produce surprisingly sophisticated estimates of their own uncertainty. However, it remains unclear to what extent this expressed confidence is tied to the reasoning, knowledge, or decision making of the model. To test this, we introduce $\textbf{RiskEval}$: a framework designed to evaluate whether...

💬 0 commentsarXiv:2601.07767v1PDF
0

Posted in cs.CL · 2026-01-12 · Igor Sterner, Alex Lascarides, Frank Keller

Contrastive Learning with Narrative Twins for Modeling Story Salience

Understanding narratives requires identifying which events are most salient for a story's progression. We present a contrastive learning framework for modeling narrative salience that learns story embeddings from narrative twins: stories that share the same plot but differ in surface form. Our model is trained to distinguish a story...

💬 0 commentsarXiv:2601.07765v1PDF
0

Posted in cs.AI · 2026-01-12 · Sahil Rajesh Dhayalkar

Reasoning Stabilization Point: A Training-Time Signal for Stable Evidence and Shortcut Reliance

Fine-tuning pretrained language models can improve task performance while subtly altering the evidence a model relies on. We propose a training-time interpretability view that tracks token-level attributions across finetuning epochs. We define explanation driftas the epoch-to-epoch change in normalized token attributions on a fixed...

💬 0 commentsarXiv:2601.11625v1PDF
0

Posted in cs.GT · 2026-01-12 · Tatiana Belova, Yuriy Dementiev, Artur Ignatiev, Danil Sagunov

Structural Approach to Guiding a Present-Biased Agent

Time-inconsistent behavior, such as procrastination or abandonment of long-term goals, arises when agents evaluate immediate outcomes disproportionately higher than future ones. This leads to globally suboptimal behavior, where plans are frequently revised or abandoned entirely. In the influential model of Kleinberg and Oren (2014)...

💬 0 commentsarXiv:2601.07763v1PDF
0

Posted in cs.CV · 2026-01-12 · Yanxiang Huang, Guohua Gao, Zhaoyang Wei, Jianyuan Ni

Video Evidence to Reasoning Efficient Video Understanding via Explicit Evidence Grounding

Large Vision-Language Models (LVLMs) face a fundamental dilemma in video reasoning: they are caught between the prohibitive computational costs of verbose reasoning and the hallucination risks of efficient, ungrounded approaches. To resolve this, we introduce the Chain of Evidence (CoE), a novel framework that architecturally...

💬 0 commentsarXiv:2601.07761v1PDF
0

Posted in cs.LG · 2026-01-12 · Shao-Ting Chiu, Siu Wun Cheung, Ulisses Braga-Neto, Chak Shing Lee, Rui Peng Li

Free-RBF-KAN: Kolmogorov-Arnold Networks with Adaptive Radial Basis Functions for Efficient Function Learning

Kolmogorov-Arnold Networks (KANs) offer a promising framework for approximating complex nonlinear functions, yet the original B-spline formulation suffers from significant computational overhead due to De Boor algorithm. While recent RBF-based variants improve efficiency, they often sacrifice the approximation accuracy inherent in the...

💬 0 commentsarXiv:2601.07760v3PDF
0

Posted in cs.CL · 2026-01-12 · Aryan Mishra, Akash Anil

Structure First, Reason Next: Enhancing a Large Language Model using Knowledge Graph for Numerical Reasoning in Financial Documents

Numerical reasoning is an important task in the analysis of financial documents. It helps in understanding and performing numerical predictions with logical conclusions for the given query seeking answers from financial texts. Recently, Large Language Models (LLMs) have shown promising results in multiple Question-Answering (Q-A)...

💬 0 commentsarXiv:2601.07754v1PDF
0

Posted in cs.CV · 2026-01-12 · Agnieszka Kaliszewska, Monika Syga

On the application of the Wasserstein metric to 2D curves classification

In this work we analyse a number of variants of the Wasserstein distance which allow to focus the classification on the prescribed parts (fragments) of classified 2D curves. These variants are based on the use of a number of discrete probability measures which reflect the importance of given fragments of curves. The performance of...

💬 0 commentsarXiv:2601.07749v1PDF
0

Posted in cs.LG · 2026-01-12 · Robert Lewis, Katie Matton, Rosalind W. Picard, John Guttag

Improving Domain Generalization in Contrastive Learning using Adaptive Temperature Control

Self-supervised pre-training with contrastive learning is a powerful method for learning from sparsely labeled data. However, performance can drop considerably when there is a shift in the distribution of data from training to test time. We study this phenomenon in a setting in which the training data come from multiple domains, and...

💬 0 commentsarXiv:2601.07748v1PDF
0

Posted in cs.CV · 2026-01-12 · Chen Ling, Tongwei Zhang, Hanqian Li, Nai Ding

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes

Multimodal Large Language Models (MLLMs) have demonstrated remarkable performance in mainstream visual understanding tasks, but their ability to process action scenes that contradict everyday common sense remains undertested. To address this gap, we introduce CAIT, a benchmark comprising 400 high-fidelity synthetic scenes focused on...

💬 0 commentsarXiv:2601.07737v2PDF
0

Posted in cs.CY · 2026-01-12 · Arianna Burzacchi, Marco Pistore

Evaluating Impacts of Traffic Regulations in Complex Mobility Systems Using Scenario-Based Simulations

Urban traffic regulation policies are increasingly used to address congestion, emissions, and accessibility in cities, yet their impacts are difficult to assess due to the socio-technical complexity of urban mobility systems. Recent advances in data availability and computational power enable new forms of model-driven,...

💬 0 commentsarXiv:2601.07735v3PDF
0

Posted in cs.IT · 2026-01-12 · Abdelaziz Bounhar, Mireille Sarkiss, Michèle Wigger

Distributed Detection under Stringent Resource Constraints

This paper identifies the Stein-exponent of distributed detection when the sensor communicates to the decision center over a discrete memoryless channel (DMC) subject to one of three stringent communication constraints: 1) The number of channel uses of the DMC grows sublinearly in the number of source observations n; 2) The number of...

💬 0 commentsarXiv:2601.07989v1PDF
0

Posted in cs.CL · 2026-01-12 · Adithya V Ganesan, Vasudha Varadarajan, Oscar NE Kjell, Whitney R Ringwald, Scott Feltman, Benjamin J Luft, Roman Kotov, Ryan L Boyd, H Andrew Schwartz

From Word Sequences to Behavioral Sequences: Adapting Modeling and Evaluation Paradigms for Longitudinal NLP

While NLP typically treats documents as independent and unordered samples, in longitudinal studies, this assumption rarely holds: documents are nested within authors and ordered in time, forming person-indexed, time-ordered $\textit{behavioral sequences}$. Here, we demonstrate the need for and propose a longitudinal modeling and...

💬 0 commentsarXiv:2601.07988v2PDF
0

Posted in cs.CL · 2026-01-12 · Haorui Yu, Diji Yang, Hang He, Fengrui Zhang, Qiufeng Yi

VULCA-Bench: A Multicultural Vision-Language Benchmark for Evaluating Cultural Understanding

We introduce VULCA-Bench, a multicultural art-critique benchmark for evaluating Vision-Language Models' (VLMs) cultural understanding beyond surface-level visual perception. Existing VLM benchmarks predominantly measure L1-L2 capabilities (object recognition, scene description, and factual question answering) while under-evaluate...

💬 0 commentsarXiv:2601.07986v3PDF
0

Posted in cs.CL · 2026-01-12 · Z. Melce Hüsünbeyi, Virginie Mouilleron, Leonie Uhling, Daniel Foppe, Tatjana Scheffler, Djamé Seddah

Multilingual, Multimodal Pipeline for Creating Authentic and Structured Fact-Checked Claim Dataset

The rapid proliferation of misinformation across online platforms underscores the urgent need for robust, up-to-date, explainable, and multilingual fact-checking resources. However, existing datasets are limited in scope, often lacking multimodal evidence, structured annotations, and detailed links between claims, evidence, and...

💬 0 commentsarXiv:2601.07985v3PDF
0

Posted in cs.AI · 2026-01-12 · Alfred Shen, Aaron Shen

Gated Sparse Attention: Combining Computational Efficiency with Training Stability for Long-Context Language Models

The computational burden of attention in long-context language models has motivated two largely independent lines of work: sparse attention mechanisms that reduce complexity by attending to selected tokens, and gated attention variants that improve training sta-bility while mitigating the attention sink phenomenon. We observe that...

💬 0 commentsarXiv:2601.15305v1PDF