Qwen Councils
arXiv could not process that search. Try a simpler keyword search or an arXiv field query such as all:quantum.
Showing downloaded papers while arXiv is unavailable.

Computer Science

arXiv preprints from January 1, 2026 through September 24, 2026 — 06:08:29 EST

0

Posted in cs.DB · 2026-01-10 · Daan de Graaf, Robert Brijder, Soham Chakraborty, George Fletcher, Bram van de Wall, Nikolay Yakovets

Algorithm Support for Graph Databases, Done Right

Graph database query languages cannot express algorithms like PageRank, forcing costly data wrangling, while existing solutions such as algorithm libraries, vertex-centric APIs, and recursive CTEs lack the necessary combination of expressiveness, performance, and usability. We present GraphAlg: a domain-specific language for graph...

💬 0 commentsarXiv:2601.06705v1PDF
0

Posted in cs.LG · 2026-01-10 · Dushan N. Wadduwage, Dineth Jayakody, Leonidas Zimianitis

Beyond Perfect Scores: Proof-by-Contradiction for Trustworthy Machine Learning

Machine learning (ML) models show strong promise for new biomedical prediction tasks, but concerns about trustworthiness have hindered their clinical adoption. In particular, it is often unclear whether a model relies on true clinical cues or on spurious hierarchical correlations in the data. This paper introduces a simple yet broadly...

💬 0 commentsarXiv:2601.06704v1PDF
0

Posted in cs.CY · 2026-01-10 · Seung Jun Choi

Mapping and Comparing Climate Equity Policy Practices Using RAG LLM-Based Semantic Analysis and Recommendation Systems

This study investigates the use of large language models to enhance the policymaking process. We first analyze planning-related job postings to revisit the evolving roles of planners in the era of AI. We then examine climate equity plans across the U.S. and apply ChatGPT to conduct semantic analysis, extracting policy, strategy, and...

💬 0 commentsarXiv:2601.06703v1PDF
0

Posted in cs.CL · 2026-01-10 · Besher Hassan, Xiuying Chen

GRASP LoRA: GRPO Guided Adapter Sparsity Policy for Cross Lingual Transfer

Parameter efficient fine tuning is a way to adapt LLMs to new languages when compute or data are limited, yet adapter pipelines usually choose a global prune ratio by grid search. This practice is computationally expensive and development set intensive, since it repeats training, freezes sparsity, and misses fractional optima. We...

💬 0 commentsarXiv:2601.06702v1PDF
0

Posted in cs.LG · 2026-01-10 · Poushali Sengupta, Rabindra Khadka, Sabita Maharjan, Frank Eliassen, Yan Zhang, Shashi Raj Pandey, Pedro G. Lind, Anis Yazidi

Explainability of Complex AI Models with Correlation Impact Ratio

Complex AI systems make better predictions but often lack transparency, limiting trustworthiness, interpretability, and safe deployment. Common post hoc AI explainers, such as LIME, SHAP, HSIC, and SAGE, are model agnostic but are too restricted in one significant regard: they tend to misrank correlated features and require costly...

💬 0 commentsarXiv:2601.06701v1PDF
0

Posted in cs.CL · 2026-01-10 · Zhiyao Zhang, Yazan Mash'Al, Yuhan Wu

Characterising Toxicity in Generative Large Language Models

In recent years, the advent of the attention mechanism has significantly advanced the field of natural language processing (NLP), revolutionizing text processing and text generation. This has come about through transformer-based decoder-only architectures, which have become ubiquitous in NLP due to their impressive text processing and...

💬 0 commentsarXiv:2601.06700v1PDF
0

Posted in cs.CR · 2026-01-10 · Boutaina Jebari, Khalil Ibrahimi, Hamidou Tembine, Mounir Ghogho

Incentive Mechanism Design for Privacy-Preserving Decentralized Blockchain Relayers

Public blockchains, though renowned for their transparency and immutability, suffer from significant privacy concerns. Network-level analysis and long-term observation of publicly available transactions can often be used to infer user identities. To mitigate this, several blockchain applications rely on relayers, which serve as...

💬 0 commentsarXiv:2601.06699v2PDF
0

Posted in cs.CR · 2026-01-10 · Saleem Ishaq Tijjani, Bogdan Ghita, Nathan Clarke, Matthew Craven

S-DAPT-2026: A Stage-Aware Synthetic Dataset for Advanced Persistent Threat Detection

The detection of advanced persistent threats (APTs) remains a crucial challenge due to their stealthy, multistage nature and the limited availability of realistic, labeled datasets for systematic evaluation. Synthetic dataset generation has emerged as a practical approach for modeling APT campaigns; however, existing methods often...

💬 0 commentsarXiv:2601.06690v2PDF
0

Posted in cs.SE · 2026-01-10 · Mateus Costa Lucena

An Exploratory Pilot Survey on Technical Quality Control Practices in Agile R&D Projects

Managing technical quality in agile Research and Development (R&D) software projects represents a persistent challenge, particularly in contexts characterized by high technical uncertainty and experimental pressure. This exploratory pilot survey explores how agile R&D software teams report the use of practices and metrics related to...

💬 0 commentsarXiv:2601.06689v1PDF
0

Posted in cs.IT · 2026-01-10 · Terence Viaud, Ioannis Kontoyiannis

The Sample Complexity of Lossless Data Compression

A new framework is introduced for examining and evaluating the fundamental limits of lossless data compression, that emphasizes genuinely non-asymptotic results. The {\em sample complexity} of compressing a given source is defined as the smallest blocklength at which it is possible to compress that source at a specifically constrained...

💬 0 commentsarXiv:2601.06688v5PDF
0

Posted in cs.CY · 2026-01-10 · Stefaan Verhulst

The Case for Strategic Data Stewardship: Re-imagining Data Governance to Make Responsible Data Re-use Possible

As societal challenges grow more complex, access to data for public interest use is paradoxically becoming more constrained. This emerging data winter is not simply a matter of scarcity, but of shrinking legitimate and trusted pathways for responsible data reuse. Concerns over misuse, regulatory uncertainty, and the competitive race...

💬 0 commentsarXiv:2601.06687v1PDF
0

Posted in cs.SE · 2026-01-10 · Vignesh Alagappan

A Governance Model for IoT Data in Global Manufacturing

Industrial IoT platforms in global manufacturing environments generate continuous operational data across production assets, utilities, and connected products. While data ingestion and storage capabilities have matured significantly, enterprises continue to face systemic challenges in governing IoT data at scale. These challenges are...

💬 0 commentsarXiv:2601.09744v1PDF
0

Posted in cs.DB · 2026-01-10 · Isabelle Mohr, Joao Gandarela, John Dujany, Andre Freitas

Reflective Reasoning for SQL Generation

Robust text-to-SQL over complex, real-world databases remains brittle even with modern LLMs: iterative refinement often introduces syntactic and semantic drift, corrections tend to be non-transferable across queries, and naive use of large context windows scales poorly. We propose a controlled text-to-SQL framework built around...

💬 0 commentsarXiv:2601.06678v1PDF
0

Posted in cs.LG · 2026-01-10 · Zohaib Khan, Omer Tafveez, Zoha Hayat Bhatti

Plasticity vs. Rigidity: The Impact of Low-Rank Adapters on Reasoning on a Micro-Budget

Recent advances in mathematical reasoning typically rely on massive scale, yet the question remains: can strong reasoning capabilities be induced in small language models ($\leq1.5\text{B}$) under extreme constraints? We investigate this by training models on a single A40 GPU (48GB) for under 24 hours using Reinforcement Learning with...

💬 0 commentsarXiv:2601.06677v1PDF
0

Posted in cs.CL · 2026-01-10 · Yingchaojie Feng, Qiang Huang, Xiaoya Xie, Zhaorui Yang, Jun Yu, Wei Chen, Anthony K. H. Tung

One Interaction Is Worth a Thousand Guesses: Benchmarking the Interactive Capabilities of Deep Research Agents

Deep research agents powered by Large Language Models (LLMs) can perform multi-step reasoning, web exploration, and long-form report generation. However, existing systems remain largely autonomous, assuming fully specified user intent and evaluating only final outputs. In practice, research goals are often underspecified and evolve...

💬 0 commentsarXiv:2601.06676v2PDF
0

Posted in cs.CL · 2026-01-10 · Tyler Lizzo, Larry Heck

Evaluating Cross-Lingual Unlearning in Multilingual Language Models

We present the first comprehensive evaluation of cross-lingual unlearning in multilingual LLMs. Using translated TOFU benchmarks in seven language/script variants, we test major unlearning algorithms and show that most fail to remove facts outside the training language, even when utility remains high. However, subspace-projection...

💬 0 commentsarXiv:2601.06675v1PDF
0

Posted in cs.CV · 2026-01-10 · Sanjay Pradeep, Chen Wang, Matthew M. Dahm, Jeff D. Eldredge, Candace S. J. Tsai

Quantification and Classification of Carbon Nanotubes in Electron Micrographs using Vision Foundation Models

Accurate characterization of carbon nanotube morphologies in electron microscopy images is vital for exposure assessment and toxicological studies, yet current workflows rely on slow, subjective manual segmentation. This work presents a unified framework leveraging vision foundation models to automate the quantification and...

💬 0 commentsarXiv:2601.06673v2PDF
0

Posted in cs.CL · 2026-01-10 · Adir Rahamim, Asaf Yehudai, Boaz Carmeli, Leshem Choshen, Yosi Mass, Yonatan Belinkov

Will it Merge? On The Causes of Model Mergeability

Model merging has emerged as a promising technique for combining multiple fine-tuned models into a single multitask model without retraining. However, the factors that determine whether merging will succeed or fail remain poorly understood. In this work, we investigate why specific models are merged better than others. To do so, we...

💬 0 commentsarXiv:2601.06672v1PDF
0

Posted in cs.CY · 2026-01-10 · Francisco Glaubos Nunes Clímaco, Jorge Lucas Silva Cavalcante

Otimizando A Alocação De Salas De Aula Com Foco Na Acessibilidade Para Pessoas Com Deficiência

This paper addresses the challenge of classroom allocation in higher education institutions, with an explicit emphasis on accessibility for Persons with Disabilities (PwDs). Employing a case study of a university's computer science department, the paper proposes an Integer Linear Programming (ILP)-based optimization model, which is...

💬 0 commentsarXiv:2601.06670v1PDF
0

Posted in cs.CR · 2026-01-10 · Xinyu Hou, Yang Lu, Rabimba Karanjai, Lei Xu, Weidong Shi

zkRansomware: Proof-of-Data Recoverability and Multi-round Game Theoretic Modeling of Ransomware Decisions

Ransomware is still one of the most serious cybersecurity threats. Victims often pay but fail to regain access to their data, while also facing the danger of losing data privacy. These uncertainties heavily shape the attacker-victim dynamics in decision-making. In this paper, we introduce and analyze zkRansomware. This new ransomware...

💬 0 commentsarXiv:2601.06667v1PDF
0

Posted in cs.CL · 2026-01-10 · Yuzhuo Bai, Shuzheng Si, Kangyang Luo, Qingyi Wang, Wenhao Li, Gang Chen, Fanchao Qi, Maosong Sun

InFi-Check: Interpretable and Fine-Grained Fact-Checking of LLMs

Large language models (LLMs) often hallucinate, yet most existing fact-checking methods treat factuality evaluation as a binary classification problem, offering limited interpretability and failing to capture fine-grained error types. In this paper, we introduce InFi-Check, a framework for interpretable and fine-grained fact-checking...

💬 0 commentsarXiv:2601.06666v1PDF
0

Posted in cs.LG · 2026-01-10 · Harshil Vejendla

RewriteNets: End-to-End Trainable String-Rewriting for Generative Sequence Modeling

Dominant sequence models like the Transformer represent structure implicitly through dense attention weights, incurring quadratic complexity. We propose RewriteNets, a novel neural architecture built on an alternative paradigm: explicit, parallel string rewriting. Each layer in a RewriteNet contains a set of learnable rules. For each...

💬 0 commentsarXiv:2601.07868v1PDF
0

Posted in cs.LG · 2026-01-10 · Md Nafees Fuad Rafi, Samiul Hasan

Reinforcement Learning-Guided Dynamic Multi-Graph Fusion for Evacuation Traffic Prediction

Real-time traffic prediction is critical for managing transportation systems during hurricane evacuations. Although data-driven graph-learning models have demonstrated strong capabilities in capturing the complex spatiotemporal dynamics of evacuation traffic at a network level, they mostly consider a single dimension (e.g.,...

💬 0 commentsarXiv:2601.06664v1PDF
0

Posted in cs.AI · 2026-01-10 · Kaiwen Zhou, Shreedhar Jangam, Ashwin Nagarajan, Tejas Polu, Suhas Oruganti, Chengzhi Liu, Ching-Chen Kuo, Yuting Zheng, Sravana Narayanaraju, Xin Eric Wang

SafePro: Evaluating the Safety of Professional-Level AI Agents

Large language model-based agents are rapidly evolving from simple conversational assistants into autonomous systems capable of performing complex, professional-level tasks in various domains. While these advancements promise significant productivity gains, they also introduce critical safety risks that remain under-explored. Existing...

💬 0 commentsarXiv:2601.06663v2PDF