Qwen Councils

Statistics

arXiv preprints from January 1, 2026 through September 18, 2026 — 18:35:25 EST

0

Posted in stat.ME · 2026-09-15 · Takes Fujita, Nobutaka Hattori

When AI Generates Covariates: Causal Typing and Estimand Drift in Sequential Experiments

AI-generated covariates from notes, conversations, images, and wearable streams can change the causal question when their roles are left unspecified. A generated feature may represent a treatment version, pre-action state, history, design variable, mediator, outcome proxy, observation process, or intercurrent event; these roles are...

💬 0 commentsarXiv:2609.17772v1PDF
0

Posted in stat.ME · 2026-09-16 · Juejue Wang, Pedro H. C. Sant'Anna, Victor Chernozhukov, Carlos Cinelli

Omitted Variable Bias in Difference-in-Differences Designs

We study the omitted variable bias (OVB) problem in canonical difference-in-differences (DiD) designs when unobserved confounding induces departures from the parallel trends assumption. Our results provide a novel characterization of the OVB formula for the average treatment effect on the treated (ATT), which is of independent...

💬 0 commentsarXiv:2609.19386v1PDF
0

Posted in stat.ML · 2026-09-17 · Sho Kawano, Zehang Richard Li, Paul A. Parker

Prediction-Powered Smoothing and Validation for Disaggregated AI Evaluation

Evaluating an AI system requires disaggregated assessment, as performance varies across domains such as benchmark task types or conversation types in deployed agents. Exhaustive testing is expensive, so evaluation rests on a sample of labeled units. We treat the evaluation set as a finite population and seek accurate point and...

💬 0 commentsarXiv:2609.20758v1PDF
0

Posted in stat.CO · 2026-09-17 · Zhiliang Deng, Xiaomei Yang

Hankel-Christoffel-Nevai Screening of Posterior Relevance in Bayesian Inverse Problems

We introduce a Hankel--Christoffel--Nevai framework for screening posterior-relevant candidates in Bayesian inverse problems. A likelihood-weighted moment matrix records how Bayesian updating changes the geometry of the prior, and Christoffel and Nevai constructions convert this information into inexpensive relevance scores. The...

💬 0 commentsarXiv:2609.20711v1PDF
0

Posted in stat.ME · 2026-09-17 · Penghui Fu, Xiaoxian Ding, Chunlin Ji, Jianhua Z. Huang, C. F. Jeff Wu

Sampling-Based Batch Sequential Design by Stein Variational Gradient Descent

Many real-world experimental design problems require a batch of experimental runs across stages, in which multiple points are selected and evaluated at each stage. However, most work in the design literature is focused on fully sequential (point-by-point) methods. This paper proposes a sampling-based framework to systematically...

💬 0 commentsarXiv:2609.20583v1PDF
0

Posted in stat.ML · 2026-09-17 · Jingbo Liu, Zhiyuan Yu

TAP Accuracy Below the Fluctuation Scale and Universal Posterior Geometry in Spherical Linear Models

We study the Bayes-optimal spherical linear model as the ambient dimension and sample size grow proportionally, under a quantitative Marchenko--Pastur spectral-regularity condition on the design. This condition is satisfied by normalized i.i.d. designs with standardized entries of finite fourth moment, but does not require entrywise...

💬 0 commentsarXiv:2609.20577v1PDF
0

Posted in stat.ME · 2026-09-17 · Claudia Di Caterina, Luigi Pace, Alessandra Salvan, Nicola Sartori

Efficient computation of mixture confidence sequences in generalized linear models

Classical confidence intervals, when repeatedly obtained on accumulating data at different sample sizes, produce contradictory inferences with high probability. We propose a simple and efficient strategy for computing, instead, mixture confidence sequences for regression coefficients in generalized linear models under this sequential...

💬 0 commentsarXiv:2609.20496v1PDF
0

Posted in stat.ML · 2026-09-17 · Zhenlin Yao, Wei Xiong

Online Supervised Dimension Reduction with Random Features: Diagnostics and Computational Trade-offs

Accurate optimization of a supervised spectral objective need not produce an accurate population subspace or a better predictive representation. We investigate these distinctions for Online Kernel Supervised Principal Component Analysis (OKSPCA), which combines a centered cross-moment in finite random-feature coordinates with an...

💬 0 commentsarXiv:2609.20454v1PDF
0

Posted in stat.ME · 2026-09-17 · Steve Lawford

Gaussian Boundary Inference in a Hypergeometric Heavy-Tailed Family

This paper develops inference for a Gaussian-nested hypergeometric family of distribution functions. The family \[ G_c(z) = \frac12 + z\,\frac{Γ(c-1/2)}{2\sqrt2\,Γ(c)}\,{}_1F_1\!\left(\frac12;c;-\frac{z^2}{2}\right),\quad c\ge\frac32, \] contains the standard normal distribution at the boundary $c=3/2$. Away from the boundary, the...

💬 0 commentsarXiv:2609.20393v1PDF
0

Posted in stat.ML · 2026-09-17 · Weiwei Wang, Yuqiang Li, Xianyi Wu, Bingyi Jing

Model-based Bootstrap for Offline Policy Evaluation in Tabular Reinforcement Learning

Offline policy evaluation (OPE) is crucial in high-stakes reinforcement learning applications, where new policies must be assessed reliably before deployment. In such settings, point estimates alone are insufficient; principled uncertainty quantification, such as confidence intervals and variance estimates, is essential for safe and...

💬 0 commentsarXiv:2609.20389v1PDF
0

Posted in stat.ME · 2026-09-17 · Jieru Shi, Rajen D. Shah

Conditional Independence Testing in Time Series

We consider the problem of testing Granger causality in time series, specifically, whether the future outcome $Y_{t+1}$ and the exposure history $\bar X_t$ are conditionally independent given the history of $\bar Y_t$ and a set of confounding variables $\bar Z_t$ up to time $t$. This testing procedure distinguishes true causal effects...

💬 0 commentsarXiv:2609.20772v1PDF
0

Posted in stat.ME · 2026-09-17 · Faria Rauf Ria, Tarikul Islam, Mahbub A. H. M. Latif

Identification and Estimation of Causal Estimands with Missing Not at Random Data

Missing not at random (MNAR) data pose significant challenges for causal inference, particularly when both confounders and the outcome are partially observed. Without additional assumptions beyond those required for causal inference, causal estimands are generally not identifiable under MNAR mechanisms. This paper first develops...

💬 0 commentsarXiv:2609.20113v1PDF
0

Posted in stat.ME · 2026-09-17 · Sultan Amed, Sayantan Banerjee

Noise-adjusted turnover in estimated networks

Economic networks are often estimated separately over two periods, and changes in their edge sets are interpreted as structural rewiring. Since both networks are estimated, observed turnover also reflects graph-selection error. We study the two-snapshot Hamming-turnover functional under a homogeneous edge-misclassification model. With...

💬 0 commentsarXiv:2609.20044v1PDF
0

Posted in stat.ME · 2026-09-17 · Chidiogo Joy Agboeke, Hamidreza Maleki Almani, Dario Gasbarra, Foad Shokrollahi, Tommi Sottinen

Parameter Estimation for the Mixed Fractional Merton Jump Diffusion Model with EM Algorithm

This paper proposes an Expectation--Maximization algorithm with Metropolis--Hastings sampling for parameter estimation in a Mixed Fractional Merton Jump Diffusion model. The model combines fractional Brownian motion to capture long-range dependence with a compound Poisson jump process to describe abrupt movements in financial returns....

💬 0 commentsarXiv:2609.20041v1PDF
0

Posted in stat.AP · 2026-09-17 · Caelan McNamara, Fion Tan, Ella White, Elizaveta Semenova, Marta Blangiardo

Comparing statistical learning models in wastewater-based epidemiology: An application to norovirus

Wastewater-based epidemiology (WBE) is an increasingly important tool for infectious disease surveillance, but there has been limited direct comparison of modelling approaches for predicting pathogen concentrations across space and time. We compare the predictive performance of six modelling approaches using norovirus in England as a...

💬 0 commentsarXiv:2609.20038v1PDF
0

Posted in stat.ME · 2026-09-17 · Aaron Coats, Vinny Davies, Mayetri Gupta

A Bayesian Bi-Directional Splitting Framework for Variable Selection in Large Datasets

Modern tabular datasets are becoming increasingly large, both in the number of samples and covariates, posing significant challenges for Bayesian variable selection due to the resulting computational burden. While there is extensive literature on scaling Bayesian inference to large numbers of observations or high-dimensional covariate...

💬 0 commentsarXiv:2609.20000v1PDF
0

Posted in stat.ME · 2026-09-17 · Alexander D. V. Spiers, Michael J. Grayling, Graham M. Wheeler, Adrian P. Mander

Gain-function optimisation of graphical multiple testing procedures for confirmatory clinical trials

Graphical multiple testing procedures are a flexible and transparent way to control the family-wise error rate when a confirmatory trial pursues several label claims, but they leave open the question of which graph to use. In practice sponsors often fall back on fixed-sequence or Holm procedures that may poorly reflect what the trial...

💬 0 commentsarXiv:2609.19994v1PDF
0

Posted in stat.ML · 2026-09-17 · Xianjun Li, Yunfei Yang

Error bounds in Sobolev norms for approximations with norm constrained ReLU neural networks

Recent studies have shown that smooth functions can be well approximated by ReLU neural networks with path norm constraint on the weights. We extend these results from uniform approximation to approximation in Sobolev norm. Specifically, we analyze how well Sobolev functions in $W^{n,p}$ can be approximated by neural networks with...

💬 0 commentsarXiv:2609.19937v1PDF
0

Posted in stat.ME · 2026-09-17 · Zhihao Qiao, Budhi Surya, Azam Asanjarani, Yoni Nazarathy

Multi-Absorbing Phase-Type Distributions for Right-Censored Competing Risks Data

Phase-type (PH) distributions are versatile semi-parametric models for lifetime duration and can be used in survival and reliability analysis. In this paper we put forward methods and software for using PH distributions in a competing-risks model. The resulting multi-absorbing phase-type (MAPH) distribution records both the time until...

💬 0 commentsarXiv:2609.19921v1PDF
0

Posted in stat.AP · 2026-09-17 · Jingyi Li, Huaming Wu, Haixiang Zhang

Identifying Damage Pathways Linking Sequence Composition to Storage Failure in DNA Data Storage via High-Dimensional Mediation Analysis

DNA data storage offers extraordinary information density and long-term durability, but its reliability is limited by sequence-dependent errors introduced during synthesis and accumulated during storage. It remains unclear how sequence composition is associated with storage failure through specific molecular damage components. We...

💬 0 commentsarXiv:2609.19822v1PDF
0

Posted in stat.ME · 2026-09-17 · Santeri Holopainen, Jari Metsämuuronen, Mikko-Jussi Laakso, Janne V. Kujala

Mutual Information as a Tool for Optimal Classification: Application to Identifying Rapid-Responding Behaviour

Existing methods for identifying rapid-responding behaviour in large-scale assessments require parametric assumptions about the population. In this study, we propose a novel, non-parametric, mutual information-based framework of methods as an alternative. The methods within this framework compute the mutual information of the observed...

💬 0 commentsarXiv:2609.19781v1PDF
0

Posted in stat.ME · 2026-09-17 · Hyewon Kim, Seongoh Park

Matrix Graphical Model Via Joint Estimation of Partial Correlations

Matrix graphical models aim to characterize conditional dependence structures in matrix-variate data under a separable covariance assumption. In this framework, the precision matrix is decomposed as a Kronecker product, enabling separate modeling of undirected graphs across row and column domains. Existing methods have been developed...

💬 0 commentsarXiv:2609.19718v1PDF
0

Posted in stat.ME · 2026-09-17 · Prasanjit Dubey, Xiaoming Huo

Report resolution in federated multiple testing under family-wise error control

Several institutions test one family of hypotheses under family-wise error control but cannot pool their data, so each site releases, for each hypothesis, only a report of its own p-value. The report's resolution is the number of values it can take. We quantify the power lost to such reports relative to the most powerful centralized...

💬 0 commentsarXiv:2609.19708v1PDF
0

Posted in stat.ME · 2026-09-17 · Xiaofei Wu, Jian Qing Shi

Heterogeneity-calibrated Byzantine-robust distributed composite quantile regression

We study sparse composite quantile regression (CQR) for distributed data with heterogeneous honest sites and Byzantine workers. Honest sites share a common slope but may differ in their covariate distributions, error laws, and quantile intercepts. The proposed heterogeneity-calibrated robust CQR (HC-RCQR) profiles local intercepts and...

💬 0 commentsarXiv:2609.19701v1PDF
0

Posted in stat.ME · 2026-09-17 · Haochen Lei, Qian Zhang, Hongyuan Cao

Selective Inference in Growth Curve Models

Growth curve models are widely used in psychological research, and variable selection can help identify baseline characteristics associated with longitudinal heterogeneity. However, conventional inference after data-driven variable selection can be invalid because the same outcome data are used for both selection and inference. We...

💬 0 commentsarXiv:2609.19573v1PDF