Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through July 28, 2026 — 13:16:33 EST

0

Posted in cs.LG · 2026-01-01 · Ali Rahimi

The Hessian of tall-skinny networks is easy to invert

We describe an exact algorithm to solve linear systems of the form $Hx=b$ where $H$ is the Hessian of a deep net. The method computes Hessian-inverse-vector products without storing the Hessian or its inverse. It requires time and storage that scale linearly in the number of layers. This is in contrast to the naive approach of first...

💬 0 commentsarXiv:2601.06096v3PDF
0

Posted in cs.AI · 2026-01-01 · Sankar B, Srinidhi Ranjini Girish, Aadya Bharti, Dibakar Sen

Progressive Ideation using an Agentic AI Framework for Human-AI Co-Creation

The generation of truly novel and diverse ideas is important for contemporary engineering design, yet it remains a significant cognitive challenge for novice designers. Current 'single-spurt' AI systems exacerbate this challenge by producing a high volume of semantically clustered ideas. We propose MIDAS (Meta-cognitive Ideation...

💬 0 commentsarXiv:2601.00475v1PDF
0

Posted in cs.LG · 2026-01-01 · Abhisek Ganguly, Santosh Ansumali, Sauro Succi

Deep Neural Networks as Discrete Dynamical Systems: Implications for Physics-Informed Learning

We revisit the analogy between feed-forward deep neural networks (DNNs) and discrete dynamical systems derived from neural integral equations and their corresponding partial differential equation (PDE) forms. A comparative analysis between the numerical/exact solutions of the Burgers' and Eikonal equations, and the same obtained via...

💬 0 commentsarXiv:2601.00473v4PDF
0

Posted in cs.CE · 2026-01-01 · L. Castro, Y. Navidtehrani. C. Betegón, E. Martínez-Pañeda

Coupled thermo-chemo-mechanical phase field-based modelling of hydrogen-assisted cracking in girth welds

A new computational framework is presented to predict the structural integrity of welds in hydrogen transmission pipelines. The framework combines: (i) a thermo-mechanical weld process model, and (ii) a coupled deformation-diffusion-fracture phase field-based model that accounts for plasticity and hydrogen trapping, considering...

💬 0 commentsarXiv:2601.00471v1PDF
0

Posted in cs.SE · 2026-01-01 · Negin Ayoughi, David Dewar, Shiva Nejati, Mehrdad Sabetzadeh

DSL or Code? Evaluating the Quality of LLM-Generated Algebraic Specifications: A Case Study in Optimization at Kinaxis

Model-driven engineering (MDE) provides abstraction and analytical rigour, but industrial adoption in many domains has been limited by the cost of developing and maintaining models. Large language models (LLMs) can help shift this cost balance by supporting direct generation of models from natural-language (NL) descriptions. For...

💬 0 commentsarXiv:2601.00469v2PDF
0

Posted in cs.RO · 2026-01-01 · Dennis Christmann, Juan F. Gutierrez, Sthiti Padhi, Patrick Plörer, Aditya Takur, Simona Silvestri, Andres Gomez

Space Debris Removal using Nano-Satellites controlled by Low-Power Autonomous Agents

Space debris is an ever-increasing problem in space travel. There are already many old, no longer functional spacecraft and debris orbiting the earth, which endanger both the safe operation of satellites and space travel. Small nano-satellite swarms can address this problem by autonomously de-orbiting debris safely into the Earth's...

💬 0 commentsarXiv:2601.00465v1PDF
0

Posted in cs.CE · 2026-01-01 · Chandrasekhar Gokavarapu, Komala Lakshmi Chinnam

Harmonic Analysis on Directed Networks via a Biorthogonal Laplacian Calculus for Non-Normal Digraphs

Spectral graph signal processing is traditionally built on self-adjoint Laplacians, where orthogonal eigenbases yield an energy-preserving Fourier transform and a variational frequency ordering via a real Dirichlet form. Directed networks break self-adjointness: the combinatorial directed Laplacian $L=D_{\mathrm{out}}-A$ is generally...

💬 0 commentsarXiv:2601.00464v2PDF
0

Posted in cs.LG · 2026-01-01 · Shuang Wu, Arash A. Amini

Laplacian Kernelized Bandit

We study multi-user contextual bandits where users are related by a graph and their reward functions exhibit both non-linear behavior and graph homophily. We introduce a principled joint penalty for the collection of user reward functions $\{f_u\}$, combining a graph smoothness term based on RKHS distances with an individual roughness...

💬 0 commentsarXiv:2601.00461v1PDF
0

Posted in cs.LG · 2026-01-01 · Saurav Sengupta, Scott Kilianski, Suchetha Sharma, Sakina Lashkeri, Ashley McHugh, Mark Beenhakker, Donald E. Brown

Combining Residual U-Net and Data Augmentation for Dense Temporal Segmentation of Spike Wave Discharges in Single-Channel EEG

Manual annotation of spike-wave discharges (SWDs), the electrographic hallmark of absence seizures, is labor-intensive for long-term electroencephalography (EEG) monitoring studies. While machine learning approaches show promise for automated detection, they often struggle with cross-subject generalization due to high inter-individual...

💬 0 commentsarXiv:2601.00459v2PDF
0

Posted in cs.LG · 2026-01-01 · Hyunjun Kim

Geometric Regularization in Mixture-of-Experts: The Disconnect Between Weights and Activations

Mixture-of-Experts (MoE) models achieve efficiency through sparse activation, but the role of geometric regularization in expert specialization remains unclear. We apply orthogonality loss to enforce expert diversity and find it fails on multiple fronts: it does not reduce weight-space overlap (MSO actually increases by up to 114%),...

💬 0 commentsarXiv:2601.00457v1PDF
0

Posted in cs.AR · 2026-01-01 · Elham Cheshmikhani, Hamed Farbeh, Hossein Asadi

ROBIN: Incremental Oblique Interleaved ECC for Reliability Improvement in STT-MRAM Caches

Spin-Transfer Torque Magnetic RAM} (STT-MRAM) is a promising alternative for SRAMs in on-chip cache memories. Besides all its advantages, high error rate in STT-MRAM is a major limiting factor for on-chip cache memories. In this paper, we first present a comprehensive analysis that reveals that the conventional Error-Correcting Codes...

💬 0 commentsarXiv:2601.00456v1PDF
0

Posted in cs.LG · 2026-01-01 · Amit Daniely

Deep Networks Learn Deep Hierarchical Models

We consider supervised learning with $n$ labels and show that layerwise SGD on residual networks can efficiently learn a class of hierarchical models. This model class assumes the existence of an (unknown) label hierarchy $L_1 \subseteq L_2 \subseteq \dots \subseteq L_r = [n]$, where labels in $L_1$ are simple functions of the input,...

💬 0 commentsarXiv:2601.00455v1PDF
0

Posted in cs.CL · 2026-01-01 · Hyunjun Kim

Entropic Context Shaping: Information-Theoretic Filtering for Context-Aware LLM Agents

Context engineering for large language model (LLM) agents requires distinguishing pragmatically useful information from misleading distractors. We introduce Entropic Context Shaping (ECS), an information-theoretic framework that measures context utility via the shift in the model's answer distribution toward the correct answer. Unlike...

💬 0 commentsarXiv:2601.11585v1PDF
0

Posted in cs.CL · 2026-01-01 · Hyunjun Kim

Defensive M2S: Training Guardrail Models on Compressed Multi-turn Conversations

Guardrail models are essential for ensuring the safety of Large Language Model (LLM) deployments, but processing full multi-turn conversation histories incurs significant computational cost. We propose Defensive M2S, a training paradigm that fine-tunes guardrail models on Multi-turn to Single-turn (M2S) compressed conversations rather...

💬 0 commentsarXiv:2601.00454v1PDF
0

Posted in cs.LG · 2026-01-01 · Yongtao Qu, Shangzhe Li, Weitong Zhang

Imitation from Observations with Trajectory-Level Generative Embeddings

We consider the offline imitation learning from observations (LfO) where the expert demonstrations are scarce and the available offline suboptimal data are far from the expert behavior. Many existing distribution-matching approaches struggle in this regime because they impose strict support constraints and rely on brittle one-step...

💬 0 commentsarXiv:2601.00452v2PDF
0

Posted in cs.LG · 2026-01-01 · Hongbin Lin, Chenyang Ren, Juangui Xu, Zhengyu Hu, Cheng-Long Wang, Yao Shu, Hui Xiong, Jingfeng Zhang, Di Wang, Lijie Hu

Controllable Concept Bottleneck Models

Concept Bottleneck Models (CBMs) have garnered much attention for their ability to elucidate the prediction process through a human-understandable concept layer. However, most previous studies focused on static scenarios where the data and concepts are assumed to be fixed and clean. In real-world applications, deployed models require...

💬 0 commentsarXiv:2601.00451v1PDF
0

Posted in cs.AR · 2026-01-01 · Elham Cheshmikhani, Hamed Farbeh, Hossein Asadi

Enhancing Reliability of STT-MRAM Caches by Eliminating Read Disturbance Accumulation

Spin-Transfer Torque Magnetic RAM (STT-MRAM) as one of the most promising replacements for SRAMs in on-chip cache memories benefits from higher density and scalability, near-zero leakage power, and non-volatility, but its reliability is threatened by high read disturbance error rate. Error-Correcting Codes (ECCs) are conventionally...

💬 0 commentsarXiv:2601.00450v1PDF
0

Posted in cs.CL · 2026-01-01 · Dimitris Vartziotis

Language as Mathematical Structure: Examining Semantic Field Theory Against Language Games

Large language models (LLMs) offer a new empirical setting in which long-standing theories of linguistic meaning can be examined. This paper contrasts two broad approaches: social constructivist accounts associated with language games, and a mathematically oriented framework we call Semantic Field Theory. Building on earlier work by...

💬 0 commentsarXiv:2601.00448v1PDF
0

Posted in cs.GT · 2026-01-01 · Benjamin Cookson, Nisarg Shah, Ziqi Yu

Unifying Proportional Fairness in Centroid and Non-Centroid Clustering

Proportional fairness criteria inspired by democratic ideals of proportional representation have received growing attention in the clustering literature. Prior work has investigated them in two separate paradigms. Chen et al. [ICML 2019] study centroid clustering, in which each data point's loss is determined by its distance to a...

💬 0 commentsarXiv:2601.00447v1PDF
0

Posted in cs.LG · 2026-01-01 · Miseon Park, Kijung Yoon

A Comparative Study of Adaptation Strategies for Time Series Foundation Models in Anomaly Detection

Time series anomaly detection is essential for the reliable operation of complex systems, but most existing methods require extensive task-specific training. We explore whether time series foundation models (TSFMs), pretrained on large heterogeneous data, can serve as universal backbones for anomaly detection. Through systematic...

💬 0 commentsarXiv:2601.00446v1PDF
0

Posted in cs.CL · 2026-01-01 · Muhammad Shahmeer Khan

Comparative Efficiency Analysis of Lightweight Transformer Models: A Multi-Domain Empirical Benchmark for Enterprise NLP Deployment

In the rapidly evolving landscape of enterprise natural language processing (NLP), the demand for efficient, lightweight models capable of handling multi-domain text automation tasks has intensified. This study conducts a comparative analysis of three prominent lightweight Transformer models - DistilBERT, MiniLM, and ALBERT - across...

💬 0 commentsarXiv:2601.00444v1PDF
0

Posted in cs.CV · 2026-01-01 · I-Hsien Ting, Yi-Jun Tseng, Yu-Sheng Lin

Application of deep learning techniques in non-contrast computed tomography pulmonary angiogram for pulmonary embolism diagnosis

Pulmonary embolism is a life-threatening disease, early detection and treatment can significantly reduce mortality. In recent years, many studies have been using deep learning in the diagnosis of pulmonary embolism with contrast medium computed tomography pulmonary angiography, but the contrast medium is likely to cause acute kidney...

💬 0 commentsarXiv:2601.00925v1PDF
0

Posted in cs.IT · 2026-01-01 · Gabriel Sac Himelfarb, Moshe Schwartz

On the burst-covering radius of binary cyclic codes

We define and study burst-covering codes. We provide some general bounds connecting the parameters of a code with its burst-covering radius. We then provide stronger bounds on the burst-covering radius of cyclic codes, by employing linear-feedback shift-register (LFSR) sequences. For the case of BCH codes we prove a new bound on...

💬 0 commentsarXiv:2601.00435v2PDF
0

Posted in cs.CL · 2026-01-01 · Kian Ahrabian, Eric Boxer, Jay Pujara

Toward Better Temporal Structures for Geopolitical Events Forecasting

Forecasting on geopolitical temporal knowledge graphs (TKGs) through the lens of large language models (LLMs) has recently gained traction. While TKGs and their generalization, hyper-relational temporal knowledge graphs (HTKGs), offer a straightforward structure to represent simple temporal relationships, they lack the expressive...

💬 0 commentsarXiv:2601.00430v2PDF
0

Posted in cs.SE · 2026-01-01 · Rares Folea, Emil Slusanschi

On Plagiarism and Software Plagiarism

This paper explores the complexities of automatic detection of software similarities, in relation to the unique challenges of digital artifacts, and introduces Project Martial, an open-source software solution for detecting code similarity. This research enumerates some of the existing approaches to counter software plagiarism by...

💬 0 commentsarXiv:2601.00429v1PDF