Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through July 20, 2026 — 14:29:57 EST

0

Posted in cs.CV · 2026-01-20 · Hunter Heidenreich, Ben Elliott, Olivia Dinica, Yosheb Getachew

GutenOCR: A Grounded Vision-Language Front-End for Documents

GutenOCR is a family of grounded OCR front-ends obtained by fine-tuning Qwen2.5-VL-3B and Qwen2.5-VL-7B. The resulting single-checkpoint vision-language models expose reading, detection, and grounding through a unified, prompt-based interface. Trained on business documents, scientific articles, and synthetic grounding data, the models...

💬 0 commentsarXiv:2601.14490v2PDF
0

Posted in cs.LG · 2026-01-20 · Krish Tadigotla

Translational Gaps in Graph Transformers for Longitudinal EHR Prediction: A Critical Appraisal of GT-BEHRT

Transformer-based models have improved predictive modeling on longitudinal electronic health records through large-scale self-supervised pretraining. However, most EHR transformer architectures treat each clinical encounter as an unordered collection of codes, which limits their ability to capture meaningful relationships within a...

💬 0 commentsarXiv:2603.13231v1PDF
0

Posted in cs.LG · 2026-01-20 · Mrigank Dhingra, Omer San

Stabilizing autoregressive forecasts in chaotic systems via multi-rate latent recurrence

Long-horizon autoregressive forecasting of chaotic dynamical systems remains challenging due to rapid error amplification and distribution shift: small one-step inaccuracies compound into physically inconsistent rollouts and collapse of large-scale statistics. We introduce MSR-HINE, a hierarchical implicit forecaster that augments...

💬 0 commentsarXiv:2601.14487v1PDF
0

Posted in cs.AI · 2026-01-20 · Yuan Tian, Yi Mei, Mengjie Zhang

Scalable Knee-Point Guided Activity Group Selection in Multi-Tree Genetic Programming for Dynamic Multi-Mode Project Scheduling

The dynamic multi-mode resource-constrained project scheduling problem is a challenging scheduling problem that requires making decisions on both the execution order of activities and their corresponding execution modes. Genetic programming has been widely applied as a hyper-heuristic to evolve priority rules that guide the selection...

💬 0 commentsarXiv:2601.14485v1PDF
0

Posted in cs.NI · 2026-01-20 · Egemen Erbayat, Gustavo B. Figueiredo, Shih-Chun Lin, Motoharu Matsuura, Hiroshi Hasegawa, Suresh Subramaniam

A benchmarking framework for PON-based fronthaul network design

As mobile networks transition toward 5G and 6G RAN architectures, Passive Optical Networks (PONs) offer a critical solution for cost-effective fronthaul transport. However, the lack of standardized evaluation models in current literature makes an objective comparison of diverse optimization strategies difficult. This paper addresses...

💬 0 commentsarXiv:2601.14480v2PDF
0

Posted in cs.CL · 2026-01-20 · Crish Nagarkar, Leonid Bogachev, Serge Sharoff

Can LLM Reasoning Be Trusted? A Comparative Study: Using Human Benchmarking on Statistical Tasks

This paper investigates the ability of large language models (LLMs) to solve statistical tasks, as well as their capacity to assess the quality of reasoning. While state-of-the-art LLMs have demonstrated remarkable performance in a range of NLP tasks, their competence in addressing even moderately complex statistical challenges is not...

💬 0 commentsarXiv:2601.14479v1PDF
0

Posted in cs.CL · 2026-01-20 · Sasha Ronaghi, Emma-Louise Aveling, Maria Levis, Rachel Lauren Ross, Emily Alsentzer, Sara Singer

Large Language Models for Large-Scale, Rigorous Qualitative Analysis in Applied Health Services Research

Large language models (LLMs) show promise for improving the efficiency of qualitative analysis in large, multi-site health-services research. Yet methodological guidance for LLM integration into qualitative analysis and evidence of their impact on real-world research methods and outcomes remain limited. We developed a model- and...

💬 0 commentsarXiv:2601.14478v1PDF
0

Posted in cs.CV · 2026-01-20 · Frank Bieder, Hendrik Königshof, Haohao Hu, Fabian Immel, Yinzhe Shen, Jan-Hendrik Pauls, Christoph Stiller

XD-MAP: Cross-Modal Domain Adaptation via Semantic Parametric Maps for Scalable Training Data Generation

Until open-world foundation models match the performance of specialized approaches, deep learning systems remain dependent on task- and sensor-specific data availability. To bridge the gap between available datasets and deployment domains, domain adaptation strategies are widely used. In this work, we propose XD-MAP, a novel approach...

💬 0 commentsarXiv:2601.14477v2PDF
0

Posted in cs.LG · 2026-01-20 · Naoya Onizawa, Takahiro Hanyu

GPU-accelerated simulated annealing based on p-bits with real-world device-variability modeling

Probabilistic computing using probabilistic bits (p-bits) presents an efficient alternative to traditional CMOS logic for complex problem-solving, including simulated annealing and machine learning. Realizing p-bits with emerging devices such as magnetic tunnel junctions (MTJs) introduces device variability, which was expected to...

💬 0 commentsarXiv:2601.14476v1PDF
0

Posted in cs.CL · 2026-01-20 · Maral Doctorarastoo, Katherine A. Flanigan, Mario Bergés, Christopher McComb

Evaluating Few-Shot Temporal Reasoning of LLMs for Human Activity Prediction in Smart Environments

Anticipating human activities and their durations is essential in applications such as smart-home automation, simulation-based architectural and urban design, activity-based transportation system simulation, and human-robot collaboration, where adaptive systems must respond to human activities. Existing data-driven agent-based...

💬 0 commentsarXiv:2602.11176v1PDF
0

Posted in cs.CV · 2026-01-20 · Yajvan Ravan, Aref Malek, Chester Dolph, Nikhil Behari

Real-Time Wildfire Localization on the NASA Autonomous Modular Sensor using Deep Learning

High-altitude, multi-spectral, aerial imagery is scarce and expensive to acquire, yet it is necessary for algorithmic advances and application of machine learning models to high-impact problems such as wildfire detection. We introduce a human-annotated dataset from the NASA Autonomous Modular Sensor (AMS) using 12-channel, medium to...

💬 0 commentsarXiv:2601.14475v1PDF
0

Posted in cs.LG · 2026-01-20 · Danny Butvinik, Nana Boateng, Achi Hackmon

Adaptive KDE for Real-Time Thresholding: Prioritized Queues for Financial Crime Investigation

We study the problem of converting a continuous stream of risk scores into stable decision thresholds under non-stationary score distributions. This problem arises in a wide range of detection systems where scores must be partitioned into prioritized processing regions while preserving semantic consistency over time.

💬 0 commentsarXiv:2601.14473v2PDF
0

Posted in cs.SD · 2026-01-20 · Mohammed Salah Al-Radhi, Riad Larbi, Mátyás Bartalis, Géza Németh

Prosody-Guided Harmonic Attention for Phase-Coherent Neural Vocoding in the Complex Spectrum

Neural vocoders are central to speech synthesis; despite their success, most still suffer from limited prosody modeling and inaccurate phase reconstruction. We propose a vocoder that introduces prosody-guided harmonic attention to enhance voiced segment encoding and directly predicts complex spectral components for waveform synthesis...

💬 0 commentsarXiv:2601.14472v1PDF
0

Posted in cs.SE · 2026-01-20 · Mohamad Salim, Jasmine Latendresse, SayedHassan Khatoonabadi, Emad Shihab

Tokenomics: Quantifying Where Tokens Are Used in Agentic Software Engineering

LLM-based Multi-Agent (LLM-MA) systems are increasingly applied to automate complex software engineering tasks such as requirements engineering, code generation, and testing. However, their operational efficiency and resource consumption remain poorly understood, hindering practical adoption due to unpredictable costs and...

💬 0 commentsarXiv:2601.14470v1PDF
0

Posted in cs.DC · 2026-01-20 · Roeland Wiersema

JAXMg: A multi-GPU linear solver in JAX

Solving large dense linear systems and eigenvalue problems is a core requirement in many areas of scientific computing, but scaling these operations beyond a single GPU remains challenging within modern programming frameworks. While highly optimized multi-GPU solver libraries exist, they are typically difficult to integrate into...

💬 0 commentsarXiv:2601.14466v1PDF
0

Posted in cs.IR · 2026-01-20 · Weronika Łajewska, Krisztian Balog

Trust Me on This: A User Study of Trustworthiness for RAG Responses

The integration of generative AI into information access systems often presents users with synthesized answers that lack transparency. This study investigates how different types of explanations can influence user trust in responses from retrieval-augmented generation systems. We conducted a controlled, two-stage user study where...

💬 0 commentsarXiv:2601.14460v1PDF
0

Posted in cs.AI · 2026-01-20 · Valerio Belcamino, Nicholas Attolino, Alessio Capitanelli, Fulvio Mastrogiovanni

On the Generalization Gap in LLM Planning: Tests and Verifier-Reward RL

Recent work shows that fine-tuned Large Language Models (LLMs) can achieve high valid plan rates on PDDL planning tasks. However, it remains unclear whether this reflects transferable planning competence or domain-specific memorization. In this work, we fine-tune a 1.7B-parameter LLM on 40,000 domain-problem-plan tuples from 10 IPC...

💬 0 commentsarXiv:2601.14456v1PDF
0

Posted in cs.SE · 2026-01-20 · Madjda Fares, Yogya Gamage, Benoit Baudry

Unpacking Security Scanners for GitHub Actions Workflows

GitHub Actions is a widely used platform to automate the build and deployment of software projects through configurable workflows. As the platform's popularity grows, it also becomes a target of choice for software supply chain attacks. These attacks exploit excessive permissions, ambiguous versions or the absence of artifact...

💬 0 commentsarXiv:2601.14455v2PDF
0

Posted in cs.CL · 2026-01-19 · Kriti Bhattarai, Vipina K. Keloth, Donald Wright, Andrew Loza, Yang Ren, Hua Xu

BioPulse-QA: A Dynamic Biomedical Question-Answering Benchmark for Evaluating Factuality, Robustness, and Bias in Large Language Models

Objective: Large language models (LLMs) are increasingly applied in biomedical settings, and existing benchmark datasets have played an important role in supporting model development and evaluation. However, these benchmarks often have limitations. Many rely on static or outdated datasets that fail to capture the dynamic,...

💬 0 commentsarXiv:2601.12632v1PDF
0

Posted in cs.SI · 2026-01-19 · Abdul Sittar, Miha Cesnovar, Alenka Gucek, Marko Grobelnik

Constructing a Dataset to Support Agent-Based Modeling of Online Interactions: Users, Topics, and Interaction Networks

Agent-based modeling (ABM) provides a powerful framework for exploring how individual behaviors and interactions give rise to collective social dynamics. However, most ABMs rely on handcrafted or parameterized agent rules that are not empirically grounded, thereby limiting their realism and validation against observed data. To address...

💬 0 commentsarXiv:2601.12628v1PDF
0

Posted in cs.HC · 2026-01-19 · Jingshu Li, Tianqi Song, Nattapat Boonprakong, Zicheng Zhu, Yitian Yang, Yi-Chieh Lee

AI-exhibited Personality Traits Can Shape Human Self-concept through Conversations

Recent Large Language Model (LLM) based AI can exhibit recognizable and measurable personality traits during conversations to improve user experience. However, as human understandings of their personality traits can be affected by their interaction partners' traits, a potential risk is that AI traits may shape and bias users'...

💬 0 commentsarXiv:2601.12727v1PDF
0

Posted in cs.IT · 2026-01-19 · Rishabh Iyer

Explicit Entropic Constructions for Coverage, Facility Location, and Graph Cuts

Shannon entropy is a polymatroidal set function and lies at the foundation of information theory, yet the class of entropic polymatroids is strictly smaller than the class of all submodular functions. In parallel, submodular and combinatorial information measures (SIMs) have recently been proposed as a principled framework for...

💬 0 commentsarXiv:2601.12724v1PDF
0

Posted in cs.NE · 2026-01-19 · Yuhiro Ono, Tomohiro Harada, Yukiya Miura

An Evolutionary Framework for Automatic Optimization Benchmark Generation via Large Language Models

Optimization benchmarks play a fundamental role in assessing algorithm performance; however, existing artificial benchmarks often fail to capture the diversity and irregularity of real-world problem structures, while benchmarks derived from real-world problems are costly and difficult to construct. To address these challenges, we...

💬 0 commentsarXiv:2601.12723v2PDF
0

Posted in cs.AI · 2026-01-19 · Hanbin Wang, Jingwei Song, Jinpeng Li, Qi Zhu, Fei Mi, Ganqu Cui, Yasheng Wang, Lifeng Shang

Teaching Large Reasoning Models Effective Reflection

Large Reasoning Models (LRMs) have recently shown impressive performance on complex reasoning tasks, often by engaging in self-reflective behaviors such as self-critique and backtracking. However, not all reflections are beneficial-many are superficial, offering little to no improvement over the original answer and incurring...

💬 0 commentsarXiv:2601.12720v1PDF
0

Posted in cs.CV · 2026-01-19 · Lin Zhao, Yushu Wu, Aleksei Lebedev, Dishani Lahiri, Meng Dong, Arpit Sahni, Michael Vasilkovsky, Hao Chen, Ju Hu, Aliaksandr Siarohin, Sergey Tulyakov, Yanzhi Wang, Anil Kag, Yanyu Li

S2DiT: Sandwich Diffusion Transformer for Mobile Streaming Video Generation

Diffusion Transformers (DiTs) have recently improved video generation quality. However, their heavy computational cost makes real-time or on-device generation infeasible. In this work, we introduce S2DiT, a Streaming Sandwich Diffusion Transformer designed for efficient, high-fidelity, and streaming video generation on mobile...

💬 0 commentsarXiv:2601.12719v2PDF