Qwen Councils

Statistics

arXiv preprints from January 1, 2026 through September 20, 2026 — 17:46:23 EST

0

Posted in stat.AP · 2026-01-19 · Matthew Martin

Improving Geopolitical Forecasts with Bayesian Networks

This study explores how Bayesian networks (BNs) can improve forecast accuracy compared to logistic regression and recalibration and aggregation methods, using data from the Good Judgment Project. Regularized logistic regression models and a baseline recalibrated aggregate were compared to two types of BNs: structure-learned BNs with...

💬 0 commentsarXiv:2601.13362v1PDF
0

Posted in stat.ME · 2026-01-19 · Deep Ghoshal, Xiaofeng Shao

Resampling-free Inference for Time Series via RKHS Embedding

In this article, we study nonparametric inference problems in the context of multivariate or functional time series, including testing for goodness-of-fit, the presence of a change point in the marginal distribution, and the independence of two time series, among others. Most methodologies available in the existing literature address...

💬 0 commentsarXiv:2601.13468v2PDF
0

Posted in stat.ML · 2026-01-19 · Zihan Dong, Xiaotian Hou, Ruijia Wu, Linjun Zhang

Labels or Preferences? Budget-Constrained Learning with Human Judgments over AI-Generated Outputs

The increasing reliance on human preference feedback to judge AI-generated pseudo labels has created a pressing need for principled, budget-conscious data acquisition strategies. We address the crucial question of how to optimally allocate a fixed annotation budget between ground-truth labels and pairwise preferences in AI. Our...

💬 0 commentsarXiv:2601.13458v2PDF
0

Posted in stat.ME · 2026-01-19 · Qingyang Zhang

Categorical distance correlation under general encodings and its application to high-dimensional feature screening

In this paper, we extend distance correlation to categorical data with general encodings, such as one-hot encoding for nominal variables and semicircle encoding for ordinal variables. Unlike existing methods, our approach leverages the spacing information between categories, which enhances the performance of distance correlation. Two...

💬 0 commentsarXiv:2601.13454v1PDF
0

Posted in stat.ME · 2026-01-19 · Youmi Suk, Weicong Lyu

Identifying Causes of Test Unfairness: Manipulability and Separability

Differential item functioning (DIF) is a widely used statistical notion for identifying items that may disadvantage specific groups of test-takers. These groups are often defined by non-manipulable characteristics, e.g., gender, race/ethnicity, or English-language learner (ELL) status. While DIF can be framed as a causal fairness...

💬 0 commentsarXiv:2601.13449v1PDF
0

Posted in stat.ML · 2026-01-19 · Szabolcs Szentpéteri, Balázs Csanád Csáji

Distribution-Free Confidence Ellipsoids for Ridge Regression with PAC Bounds

Linearly parametrized models are widely used in control and signal processing, with the least-squares (LS) estimate being the archetypical solution. When the input is insufficiently exciting, the LS problem may be unsolvable or numerically unstable. This issue can be resolved through regularization, typically with ridge regression....

💬 0 commentsarXiv:2601.13436v1PDF
0

Posted in stat.ME · 2026-01-19 · Xinyuan Chen, Fan Li

Optimal estimation of generalized causal effects in cluster-randomized trials with multiple outcomes

Cluster-randomized trials (CRTs) are widely used to evaluate group-level interventions and increasingly collect multiple outcomes capturing complementary dimensions of benefit and risk. Investigators often seek a single global summary of treatment effect, yet existing methods largely focus on single-outcome estimands or rely on...

💬 0 commentsarXiv:2601.13428v2PDF
0

Posted in stat.ME · 2026-01-19 · Lorenzo Mauri, Federica Stolf, Amy H. Herring, Cameron Miller, David B. Dunson

Pathway-based Bayesian factor models for 'omics data

Interpreting RNA-sequencing data requires identifying coordinated gene expression patterns that correspond to biological pathways. Standard factor models provide useful dimension reduction but typically ignore existing pathway knowledge or incorporate it through restrictive assumptions, limiting interpretability, and reproducibility....

💬 0 commentsarXiv:2601.13419v2PDF
0

Posted in stat.ME · 2026-01-19 · Jianbin Tan, Pixu Shi

Associating High-Dimensional Longitudinal Datasets through an Efficient Cross-Covariance Decomposition

Understanding associations between paired high-dimensional longitudinal datasets is a fundamental yet challenging problem that arises across scientific domains, including longitudinal multi-omic studies. The difficulty stems from the complex, time-varying cross-covariance structure coupled with high dimensionality, which complicates...

💬 0 commentsarXiv:2601.13405v1PDF
0

Posted in stat.AP · 2026-01-19 · Abdullah M. Braik, Maria Koliou

A Two-Stage Bayesian Framework for Multi-Fidelity Online Updating of Spatial Fragility Fields

This paper addresses a long-standing gap in natural hazard modeling by unifying physics-based fragility functions with real-time post-disaster observations. It introduces a Bayesian framework that continuously refines regional vulnerability estimates as new data emerges. The framework reformulates physics-informed fragility estimates...

💬 0 commentsarXiv:2601.13396v1PDF
0

Posted in stat.ML · 2026-01-18 · Sharan Sahu, Cameron J. Hogan, Martin T. Wells

On the Provable Suboptimality of Momentum SGD in Nonstationary Stochastic Optimization

In this paper, we provide a comprehensive theoretical analysis of Stochastic Gradient Descent (SGD) and its momentum variants (Polyak Heavy-Ball and Nesterov) for tracking time-varying optima under strong convexity and smoothness. Our finite-time bounds reveal a sharp decomposition of tracking error into transient, noise-induced, and...

💬 0 commentsarXiv:2601.12238v4PDF
0

Posted in stat.AP · 2026-01-18 · Zhicheng Chen, Wenyu Chen, Xinyi Lei

A warping function-based control chart for detecting distributional changes in damage-sensitive features for structural condition assessment

Data-driven damage detection methods achieve damage identification by analyzing changes in damage-sensitive features (DSFs) derived from structural health monitoring (SHM) data. The core reason for their effectiveness lies in the fact that damage or structural state transition can be manifested as changes in the distribution of DSF...

💬 0 commentsarXiv:2601.12221v1PDF
0

Posted in stat.AP · 2026-01-18 · Sijie Zheng

A Machine Learning--Based Surrogate EKMA Framework for Diagnosing Urban Ozone Formation Regimes: Evidence from Los Angeles

Surface ozone pollution remains a persistent challenge in many metropolitan regions worldwide, as the nonlinear dependence of ozone formation on nitrogen oxides and volatile organic compounds (VOCs) complicates the design of effective emission control strategies. While chemical transport models provide mechanistic insights, they rely...

💬 0 commentsarXiv:2601.12321v1PDF
0

Posted in stat.ME · 2026-01-18 · Peterson Mambondimumwe, Sphiwe B. Skhosana, Najmeh Nakhaei Rad

Robust semi-parametric mixtures of linear experts using the contaminated Gaussian distribution

Semi- and non-parametric mixture of regressions are a very useful flexible class of mixture of regressions in which some or all of the parameters are non-parametric functions of the covariates. These models are, however, based on the Gaussian assumption of the component error distributions. Thus, their estimation is sensitive to...

💬 0 commentsarXiv:2601.12425v1PDF
0

Posted in stat.ME · 2026-01-18 · Xiaoru Huang, Tonghui Yu, Xiaoyu Liu

Single-index Semiparametric Transformation Cure Models with Interval-censored Data

Interval censored data commonly arise in medical studies when the event time of interest is only known to lie within an interval. In the presence of a cure subgroup, conventional mixture cure models typically assume a logistic model for the uncure probability and a proportional hazards model for the susceptible subjects. However, in...

💬 0 commentsarXiv:2601.12370v1PDF
0

Posted in stat.CO · 2026-01-18 · Ning Ning, Amin Wu

Bayesian Inference for Partially Observed McKean-Vlasov SDEs with Full Distribution Dependence

McKean-Vlasov stochastic differential equations (MVSDEs) describe systems whose dynamics depend on both individual states and the population distribution, and they arise widely in neuroscience, finance, and epidemiology. In many applications the system is only partially observed, making inference very challenging when both drift and...

💬 0 commentsarXiv:2601.12515v1PDF
0

Posted in stat.AP · 2026-01-18 · Shanshan Luo, Wei Li, Xueli Wang, Shaojie Wei, Zhi Geng

Assessing Interactive Causes of an Occurred Outcome Due to Two Binary Exposures

In contrast to evaluating treatment effects, causal attribution analysis focuses on identifying the key factors responsible for an observed outcome. For two binary exposure variables and a binary outcome variable, researchers need to assess not only the likelihood that an observed outcome was caused by a particular exposure, but also...

💬 0 commentsarXiv:2601.12478v1PDF
0

Posted in stat.ML · 2026-01-18 · Frank Cole, Yulong Lu, Shaurya Sehgal

A Theory of Diversity for Random Matrices with Applications to In-Context Learning of Schrödinger Equations

We address the following question: given a collection $\{\mathbf{A}^{(1)}, \dots, \mathbf{A}^{(N)}\}$ of independent $d \times d$ random matrices drawn from a common distribution $\mathbb{P}$, what is the probability that the centralizer of $\{\mathbf{A}^{(1)}, \dots, \mathbf{A}^{(N)}\}$ is trivial? We provide lower bounds on this...

💬 0 commentsarXiv:2601.12587v1PDF
0

Posted in stat.AP · 2026-01-18 · Mustafa Cavus, Przemysław Biecek, Julian Tejada, Fernando Marmolejo-Ramos, Andre Faro

Analyzing the Temporal Factors for Anxiety and Depression Symptoms with the Rashomon Perspective

This paper introduces a new modeling perspective in the public mental health domain to provide a robust interpretation of the relations between anxiety and depression, and the demographic and temporal factors. This perspective particularly leverages the Rashomon Effect, where multiple models exhibit similar predictive performance but...

💬 0 commentsarXiv:2601.20874v1PDF
0

Posted in stat.AP · 2026-01-18 · Dennis Christensen, Geir Petter Novik

Stop using limiting stimuli as a measure of sensitivities of energetic materials

Accurately estimating the sensitivity of explosive materials is a potentially life-saving task which requires standardised protocols across nations. One of the most widely applied procedures worldwide is the so-called '1-In-6' test from the United Nations (UN) Manual of Tests in Criteria, which estimates a 'limiting stimulus' for a...

💬 0 commentsarXiv:2601.12552v1PDF
0

Posted in stat.ME · 2026-01-18 · Tingxuan Han, Yuhao Wang

Rerandomization for quantile treatment effects

Although complete randomization is widely regarded as the gold standard for causal inference, covariate imbalance can still arise by chance in finite samples. Rerandomization has emerged as an effective tool to improve covariate balance across treatment groups and enhance the precision of causal effect estimation. While existing work...

💬 0 commentsarXiv:2601.12540v1PDF
0

Posted in stat.AP · 2026-01-17 · Xin Xiong, Zijian Guo, Haobo Zhu, Chuan Hong, Jordan W Smoller, Tianxi Cai, Molei Liu

Adversarial Drift-Aware Predictive Transfer: Toward Durable Clinical AI

Clinical AI systems frequently suffer performance decay post-deployment due to temporal data shifts, such as evolving populations, diagnostic coding updates (e.g., ICD-9 to ICD-10), and systemic shocks like the COVID-19 pandemic. Addressing this ``aging'' effect via frequent retraining is often impractical due to computational costs...

💬 0 commentsarXiv:2601.11860v2PDF
0

Posted in stat.ML · 2026-01-17 · Longlin Yu, Ziheng Cheng, Shiyue Zhang, Cheng Zhang

A Kernel Approach for Semi-implicit Variational Inference

Semi-implicit variational inference (SIVI) enhances the expressiveness of variational families through hierarchical semi-implicit distributions, but the intractability of their densities makes standard ELBO-based optimization biased. Recent score-matching approaches to SIVI (SIVI-SM) address this issue via a minimax formulation, at...

💬 0 commentsarXiv:2601.12023v1PDF