Qwen Councils
arXiv is taking too long to respond. Please try again or narrow your search.
Showing downloaded papers while arXiv is unavailable.

Computer Science

arXiv preprints from January 1, 2026 through July 20, 2026 — 16:07:39 EST

0

Posted in cs.AI · 2026-01-13 · Tengjun Jin, Yoojin Choi, Yuxuan Zhu, Daniel Kang

Pervasive Annotation Errors Break Text-to-SQL Benchmarks and Leaderboards

Researchers have proposed numerous text-to-SQL techniques to streamline data analytics and accelerate the development of data-driven applications. To compare these techniques and select the best one for deployment, the community depends on public benchmarks and their leaderboards. Since these benchmarks heavily rely on human...

💬 0 commentsarXiv:2601.08778v3PDF
0

Posted in cs.LG · 2026-01-13 · Yang Cai, Weiqiang Zheng

Asymptotic Universal Alignment: A New Alignment Framework via Test-Time Scaling

Aligning large language models (LLMs) to serve users with heterogeneous and potentially conflicting preferences is a central challenge for personalized and trustworthy AI. We formalize an ideal notion of universal alignment through test-time scaling: for each prompt, the model produces $k\ge 1$ candidate responses and a user selects...

💬 0 commentsarXiv:2601.08777v1PDF
0

Posted in cs.CV · 2026-01-13 · Yanhua Zhao

An Example for Domain Adaptation Using CycleGAN

Cycle-Consistent Adversarial Network (CycleGAN) is very promising in domain adaptation. In this report, an example in medical domain will be explained. We present struecture of a CycleGAN model for unpaired image-to-image translation from microscopy to pseudo H\&E stained histopathology images.

💬 0 commentsarXiv:2601.08776v3PDF
0

Posted in cs.SE · 2026-01-13 · Manideep Reddy Chinthareddy

Reliable Graph-RAG for Codebases: AST-Derived Graphs vs LLM-Extracted Knowledge Graphs

Retrieval-Augmented Generation for software engineering often relies on vector similarity search, which captures topical similarity but can fail on multi-hop architectural reasoning such as controller to service to repository chains, interface-driven wiring, and inheritance. This paper benchmarks three retrieval pipelines on Java...

💬 0 commentsarXiv:2601.08773v1PDF
0

Posted in cs.HC · 2026-01-13 · Yejoon Song, Bandi Kim, Yeju Kwon, Sung Park

Exploring the Effects of Generative AI Assistance on Writing Self-Efficacy

Generative AI (GenAI) is increasingly used in academic writing, yet its effects on students' writing self-efficacy remain contingent on how assistance is configured. This pilot study investigates how ideation-level, sentence-level, full-process, and no AI support differentially shape undergraduate writers' self-efficacy using a 2 by 2...

💬 0 commentsarXiv:2601.09033v3PDF
0

Posted in cs.AI · 2026-01-13 · Logan Ritchie, Sushant Mehta, Nick Heiner, Mason Yu, Edwin Chen

The Hierarchy of Agentic Capabilities: Evaluating Frontier Models on Realistic RL Environments

The advancement of large language model (LLM) based agents has shifted AI evaluation from single-turn response assessment to multi-step task completion in interactive environments. We present an empirical study evaluating frontier AI models on 150 workplace tasks within a realistic e-commerce RL environment from Surge. Our analysis...

💬 0 commentsarXiv:2601.09032v1PDF
0

Posted in cs.RO · 2026-01-13 · Xuetao Li, Wenke Huang, Mang Ye, Jifeng Xuan, Bo Du, Sheng Liu, Miao Li

Generalizable Geometric Prior and Recurrent Spiking Feature Learning for Humanoid Robot Manipulation

Humanoid robot manipulation is a crucial research area for executing diverse human-level tasks, involving high-level semantic reasoning and low-level action generation. However, precise scene understanding and sample-efficient learning from human demonstrations remain critical challenges, severely hindering the applicability and...

💬 0 commentsarXiv:2601.09031v1PDF
0

Posted in cs.CR · 2026-01-13 · Aniesh Chawla, Udbhav Prasad

Proactively Detecting Threats: A Novel Approach Using LLMs

Enterprise security faces escalating threats from sophisticated malware, compounded by expanding digital operations. This paper presents the first systematic evaluation of large language models (LLMs) to proactively identify indicators of compromise (IOCs) from unstructured web-based threat intelligence sources, distinguishing it from...

💬 0 commentsarXiv:2601.09029v1PDF
0

Posted in cs.CL · 2026-01-13 · Fengran Mo, Zhan Su, Yuchen Hui, Jinghan Zhang, Jia Ao Sun, Zheyuan Liu, Chao Zhang, Tetsuya Sakai, Jian-Yun Nie

OpenDecoder: Open Large Language Model Decoding to Incorporate Document Quality in RAG

The development of large language models (LLMs) has achieved superior performance in a range of downstream tasks, including LLM-based retrieval-augmented generation (RAG). The quality of generated content heavily relies on the usefulness of the retrieved information and the capacity of LLMs' internal information processing mechanism...

💬 0 commentsarXiv:2601.09028v2PDF
0

Posted in cs.LG · 2026-01-13 · Shuai Jiang, Marc Salvadó-Benasco, Eric C. Cyr, Alena Kopaničáková, Rolf Krause, Jacob B. Schroder

Layer-Parallel Training for Transformers

We present a new training methodology for transformers using a multilevel, layer-parallel approach. Through a neural ODE formulation of transformers, our application of a multilevel parallel-in-time algorithm for the forward and backpropagation phases of training achieves parallel acceleration over the layer dimension. This...

💬 0 commentsarXiv:2601.09026v2PDF
0

Posted in cs.LG · 2026-01-13 · Samuel Myren, Nidhi Parikh, Natalie Klein

Meta-learning to Address Data Shift in Time Series Classification

Across engineering and scientific domains, traditional deep learning (TDL) models perform well when training and test data share the same distribution. However, the dynamic nature of real-world data, broadly termed \textit{data shift}, renders TDL models prone to rapid performance degradation, requiring costly relabeling and...

💬 0 commentsarXiv:2601.09018v1PDF
0

Posted in cs.CL · 2026-01-13 · Haryo Akbarianto Wibowo, Alaa Elsetohy, Qinrong Cui, Alham Fikri Aji

Multicultural Spyfall: Assessing LLMs through Dynamic Multilingual Social Deduction Game

The rapid advancement of Large Language Models (LLMs) has necessitated more robust evaluation methods that go beyond static benchmarks, which are increasingly prone to data saturation and leakage. In this paper, we propose a dynamic benchmarking framework for evaluating multilingual and multicultural capabilities through the social...

💬 0 commentsarXiv:2601.09017v1PDF
0

Posted in cs.CL · 2026-01-13 · Mara Finkelstein, Isaac Caswell, Tobias Domhan, Jan-Thorsten Peter, Juraj Juraska, Parker Riley, Daniel Deutsch, Geza Kovacs, Cole Dilanni, Colin Cherry, Eleftheria Briakou, Elizabeth Nielsen, Jiaming Luo, Kat Black, Ryan Mullins, Sweta Agrawal, Wenda Xu, Erin Kats, Stephane Jaskiewicz, Markus Freitag, David Vilar

TranslateGemma Technical Report

We present TranslateGemma, a suite of open machine translation models based on the Gemma 3 foundation models. To enhance the inherent multilingual capabilities of Gemma 3 for the translation task, we employ a two-stage fine-tuning process. First, supervised fine-tuning is performed using a rich mixture of high-quality large-scale...

💬 0 commentsarXiv:2601.09012v3PDF
0

Posted in cs.CV · 2026-01-13 · Amar Kavuri, Howard C. Gifford, Mini Das

Changes in Visual Attention Patterns for Detection Tasks due to Dependencies on Signal and Background Spatial Frequencies

We aim to investigate the impact of image and signal properties on visual attention mechanisms during a signal detection task in digital images. The application of insight yielded from this work spans many areas of digital imaging where signal or pattern recognition is involved in complex heterogenous background. We used simulated...

💬 0 commentsarXiv:2601.09008v1PDF
0

Posted in cs.CV · 2026-01-13 · Xiaoyu Ji, Chenhao Zhang, Tyler James Downard, Zoltan Nagy, Ali Shakouri, Fengqing Zhu

Instance camera focus prediction for crystal agglomeration classification

Agglomeration refers to the process of crystal clustering due to interparticle forces. Crystal agglomeration analysis from microscopic images is challenging due to the inherent limitations of two-dimensional imaging. Overlapping crystals may appear connected even when located at different depth layers. Because optical microscopes have...

💬 0 commentsarXiv:2601.09004v1PDF
0

Posted in cs.AR · 2026-01-13 · Peter M. Kogge

Annotated PIM Bibliography

Processing in Memory (PIM) and similar terms such as Compute In Memory (CIM), Logic in Memory (LIM), In Memory Computing (IMC), and Near Memory Computing (NMC) have gained attention recently as a potentially ``revolutionary new'' technique. The truth, however, is that many examples of the technology go back over 60 years. This...

💬 0 commentsarXiv:2601.09002v1PDF
0

Posted in cs.CL · 2026-01-13 · Pedro Memoli Buffa, Luciano Del Corro

Entropy Sentinel: Continuous LLM Accuracy Monitoring from Decoding Entropy Traces in STEM

Deploying LLMs raises two coupled challenges: (1) monitoring--estimating where a model underperforms as traffic and domains drift--and (2) improvement--prioritizing data acquisition to close the largest performance gaps. We test whether an inference-time signal can estimate slice-level accuracy under domain shift. For each response,...

💬 0 commentsarXiv:2601.09001v4PDF
0

Posted in cs.LG · 2026-01-13 · Annalisa Belloni, Lorenzo Noci, Antonio Orvieto

Universal Dynamics of Warmup Stable Decay: understanding WSD beyond Transformers

The Warmup Stable Decay (WSD) learning rate scheduler has recently become popular, largely due to its good performance and flexibility when training large language models. It remains an open question whether the remarkable performance of WSD - using a decaying learning rate for only a fraction of training compared to cosine decay - is...

💬 0 commentsarXiv:2601.09000v1PDF
0

Posted in cs.LG · 2026-01-13 · Pranjal Patil, Anli Ji, Berkay Aydin

Physics-Guided Counterfactual Explanations for Large-Scale Multivariate Time Series: Application in Scalable and Interpretable SEP Event Prediction

Accurate prediction of solar energetic particle events is vital for safeguarding satellites, astronauts, and space-based infrastructure. Modern space weather monitoring generates massive volumes of high-frequency, multivariate time series (MVTS) data from sources such as the Geostationary perational Environmental Satellites (GOES)....

💬 0 commentsarXiv:2601.08999v1PDF
0

Posted in cs.SE · 2026-01-13 · Alexander Berndt, Thomas Bach, Rainer Gemulla, Marcus Kessel, Sebastian Baltes

On the Flakiness of LLM-Generated Tests for Industrial and Open-Source Database Management Systems

Flaky tests are a common problem in software testing. They produce inconsistent results when executed multiple times on the same code, invalidating the assumption that a test failure indicates a software defect. Recent work on LLM-based test generation has identified flakiness as a potential problem with generated tests. However, its...

💬 0 commentsarXiv:2601.08998v1PDF
0

Posted in cs.SE · 2026-01-13 · Brent Pappas, Paul Gazzillo

Build Code is Still Code: Finding the Antidote for Pipeline Poisoning

Open source C code underpins society's computing infrastructure. Decades of work has helped harden C code against attackers, but C projects do not consist of only C code. C projects also contain build system code for automating development tasks like compilation, testing, and packaging. These build systems are critcal to software...

💬 0 commentsarXiv:2601.08995v1PDF
0

Posted in cs.LG · 2026-01-13 · Emile Dos Santos Ferreira, Andrei Paleyes, Neil D. Lawrence

Optimising for Energy Efficiency and Performance in Machine Learning

The ubiquity of machine learning (ML) and the demand for ever-larger models bring an increase in energy consumption and environmental impact. However, little is known about the energy scaling laws in ML, and existing research focuses on training cost -- ignoring the larger cost of inference. Furthermore, tools for measuring the energy...

💬 0 commentsarXiv:2601.08991v2PDF
0

Posted in cs.DS · 2026-01-13 · Matteo Caporrella, Stefano Leucci

An Almost-Optimal Upper Bound on the Push Number of the Torus Puzzle

We study the Torus Puzzle, a solitaire game in which the elements of an input $m \times n$ matrix need to be rearranged into a target configuration via a sequence of unit rotations (i.e., circular shifts) of rows and/or columns. Amano et al. proposed a more permissive variant of the above puzzle, where each row and column rotation can...

💬 0 commentsarXiv:2601.08989v3PDF
0

Posted in cs.AI · 2026-01-13 · Ananya Mantravadi, Shivali Dalmia, Abhishek Mukherji

ART: Action-based Reasoning Task Benchmarking for Medical AI Agents

Reliable clinical decision support requires medical AI agents capable of safe, multi-step reasoning over structured electronic health records (EHRs). While large language models (LLMs) show promise in healthcare, existing benchmarks inadequately assess performance on action-based tasks involving threshold evaluation, temporal...

💬 0 commentsarXiv:2601.08988v1PDF
0

Posted in cs.CR · 2026-01-13 · Mohammad Waquas Usmani, Susmit Shannigrahi, Michael Zink

ABE-VVS: Attribute-Based Encrypted Volumetric Video Streaming

This work introduces ABE-VVS, a framework that performs attribute based selective coordinate encryption for point cloud based volumetric video streaming, enabling lightweight yet effective digital rights management (DRM). Rather than encrypting entire point cloud frames, our approach encrypts only selected subsets of coordinates ($X,...

💬 0 commentsarXiv:2601.08987v2PDF