Qwen Councils
arXiv is taking too long to respond. Please try again or narrow your search.
Showing downloaded papers while arXiv is unavailable.

Statistics

arXiv preprints from January 1, 2026 through September 19, 2026 — 02:17:37 EST

0

Posted in stat.ML · 2026-08-31 · Fariborz Setoudehtazang, Geoffrey J. McLachlan

Informative Label Missingness in Multiclass Classification Information Geometry and Excess Risk

Informative label missingness can change the usual efficiency ordering between completely and partially labelled classifiers because the pattern of missing labels may itself carry information about the classification model. We develop a general likelihood-based theory for this phenomenon in parametric multiclass classification. An...

💬 0 commentsarXiv:2608.30561v1PDF
0

Posted in stat.ME · 2026-08-31 · Bosen Cui, Yuhong Yang, Fan Yang

Power and sample size calculations for causal mediation analysis with a binary mediator in randomized trials

Mediation analyses are increasingly conducted in randomized trials, but a sample size adequate for the total treatment effect may leave the natural indirect effect (NIE) or natural direct effect (NDE) substantially underpowered. Randomization does not extend to the mediator, so precision depends on the conditional mediator...

💬 0 commentsarXiv:2608.30412v1PDF
0

Posted in stat.ML · 2026-08-31 · Mingzhi Song

Estimating Population-Risk Curves Along Nonconvex Gradient Flows from the Training Sample

We estimate the conditional population-risk curve of a realized smooth nonconvex gradient flow from the training sample. Flow approximate leave-one-out (Flow-ALO) propagates a deletion response and evaluates omitted observations at approximate deleted paths. The risk-curve error decomposes into response approximation, exact-LOO...

💬 0 commentsarXiv:2608.30261v1PDF
0

Posted in stat.CO · 2026-08-31 · Takato Ueno, Shuji Kijima

GPU-Parallelization of Markov Chain Pool Decoding with Unbiased MCMC

Markov chain pool decoding (MCPD) devised by Knill et al. (1996) identifies likely positive clones from noisy pooled-test results. The standard MCPD estimates clone-wise posterior probabilities using Gibbs sampling, but it may allocate excessive computational effort to low-scoring clones. This paper focuses on parallelizing MCPD on...

💬 0 commentsarXiv:2608.30239v1PDF
0

Posted in stat.ML · 2026-08-31 · Darinka Dentcheva, Xiangyu Tian

Fairness in multi-class multi-group classification problems via contextial coherent risk measures

We propose a new design of fair classifiers for multi-class classification problems in the presence of vector-valued sensitive attributes. In that scenario each sensitive attribute has multiple values and forms several groups relevant to the fairness consideration. Naturally those groups are overlapping and one should also analyze the...

💬 0 commentsarXiv:2608.30223v1PDF
0

Posted in stat.ME · 2026-08-31 · Lorenzo Gasparollo, Mats J. Stensrud

Causal inference with staggered entries and effects that change over calendar time

Studies with staggered entry, in which individuals enroll at different calendar times, are ubiquitous in medicine and related disciplines. Because these studies usually have a fixed administrative end of follow-up, identification of the estimand of interest relies on assumptions about the right-censoring mechanism. The assumptions are...

💬 0 commentsarXiv:2608.30099v1PDF
0

Posted in stat.CO · 2026-08-31 · Jongmin Mun

Multifidelity Computer Model Emulation Via Diffusion Model Steering and Targeted Maximum Likelihood

We develop a multifidelity method for fusing low-resolution simulations with computationally expensive high-resolution simulations, which are run infrequently and are therefore prone to bias. We formulate this fusion as a constrained optimization under missing-not-at-random (MNAR) selection bias. This formulation searches for the...

💬 0 commentsarXiv:2608.30096v1PDF
0

Posted in stat.ME · 2026-08-30 · Anirban Mondal, Paromita Banerjee, Abhijit Mandal

Robust K-means Clustering using the Density Power Divergence Measure

We introduce a robust clustering method, MK-means DPD, that estimates cluster centers and covariance matrices using density power divergence (DPD) measures combined with Mahalanobis distance, making it resistant to outliers and adaptable to heterogeneous, elliptical clusters, unlike the classical K-means algorithm. Since Mahalanobis...

💬 0 commentsarXiv:2608.30093v1PDF
0

Posted in stat.ML · 2026-08-30 · Shulei Wang

Learning Representations through Token Prediction: Geometry, Approximation, and Downstream Guarantees

Token prediction is a central pre-training objective for modern language models. Despite its empirical success, why token prediction learns broadly useful representations remains incompletely understood. We develop a statistical framework connecting token prediction with representation geometry, encoder approximation, and downstream...

💬 0 commentsarXiv:2608.30072v1PDF
0

Posted in stat.ML · 2026-08-30 · Yasin Khadem Charvadeh, Grace Y. Yi, Mithat Gönen, Pouya Faroughi

A Deep Latent Variable Framework for Jointly Modeling Missingness, Measurement Error, and Heterogeneity

Missing data, measurement error, and population heterogeneity are pervasive challenges in analyzing data arising from modern observational studies and machine learning applications. Although these problems frequently coexist and interact, they are often treated separately in existing works. We propose a unified probabilistic framework...

💬 0 commentsarXiv:2608.30040v1PDF
0

Posted in stat.AP · 2026-08-30 · Zongyue Teng, Ningkun Zhou, Xinyu Zhang, Robert Wallis, Qingyan Xiang

Nonlinear trajectories of lung function recovery in patients with pulmonary disease: empirical evaluation of longitudinal modeling approaches

Introduction: Longitudinal lung function recovery after pulmonary disease commonly follows nonlinear trajectories, and failure to adequately model these trajectories can lead to biased or misleading estimates of treatment effects. However, an important methodological gap remains as there is limited assessment of statistical methods...

💬 0 commentsarXiv:2608.30015v1PDF
0

Posted in stat.ME · 2026-08-30 · Omar Alzeley, Michail Tsagris

Modelling compositional data with structural zero values

Compositional data are positive multivariate data whose sum equals 1. A popular method to analyze such data is via log--ratio transformations, which are however not applicable when zero values are present. In this paper we present a conditional logistic normal distribution suitable for compositional data with structural zero values....

💬 0 commentsarXiv:2608.29954v1PDF
0

Posted in stat.ME · 2026-08-30 · Eric Goldman, Fushing Hsieh

Design of Experiment in Complex Systems based on Computational Taxonomy

Via Computational Taxonomy (CT), we develop Design of Experiment(DoE) based on rigorously redefined constituting ingredients of complex system dynamics: randomness, nonlinearity and even class, through a data-driven constructed Taxonomic Hierarchy. As an opposite quest of Classification without man-made assumptions and structures, we...

💬 0 commentsarXiv:2608.29883v1PDF
0

Posted in stat.ME · 2026-08-30 · Dongxu Yang, Wanfeng Liang, Le Zhou, Long Feng

Spatial-sign-based multilinear principal component analysis for tensor data

Multilinear principal component analysis (MPCA) reduces the dimension of tensor-valued data while preserving their mode-specific structure, but its quadratic scatter criterion can be unstable under heavy-tailed distributions and contamination. We propose spatial-sign-based multilinear principal component analysis (SMPCA), a robust...

💬 0 commentsarXiv:2608.29862v1PDF
0

Posted in stat.ME · 2026-08-30 · Kexuan Li, Xue Fan, Lingli Yang

Worst-Case Win Ratios Under Partially Specified Outcome Hierarchies

Win statistics require a prespecified outcome hierarchy. Clinical teams sometimes agree only on the highest priority outcome, leaving the order of lower priority outcomes unresolved, and clinically meaningful thresholds may be specified as ranges. Separate sensitivity analyses describe how the results change. A single inference for...

💬 0 commentsarXiv:2608.29857v1PDF
0

Posted in stat.ME · 2026-08-30 · Shintaro Yoshizawa

A Generalized Ridge Regression and Convolutional LASSO

We derive the complete duality theory underlying the Hodrick--Prescott filter, whose rank-deficient second-difference penalty admits infinitely many equivalent trend representations via generalized inverses. Constructing two canonical choices---the Moore--Penrose-based \emph{B-representation} and an alternative...

💬 0 commentsarXiv:2608.29821v1PDF
0

Posted in stat.AP · 2026-08-30 · Robert Dalton, Aidan O'Sullivan

Decarbonising price formation: unit-level evidence on battery storage and the imbalance price in the GB Balancing Mechanism

Renewables now dominate Great Britain's generation mix but rarely occupy the marginal price-setting position, which raises the question of which flexible technologies translate a renewable-rich system into real-time price formation. This study reconstructs the price-ranked edge of the eligible bid or offer stack in the GB Balancing...

💬 0 commentsarXiv:2608.29818v1PDF
0

Posted in stat.ME · 2026-08-30 · Jing Zhou, Dominik Janzing, Sepp Tsang, Patrick Blöbaum, Marco Visentini Scarzanella

A Unified Approach to Interpretable Causal Root Cause Attribution

Understanding why a target metric changes is a fundamental problem in data-driven decision making, beyond anomaly detection alone. We study root cause attribution for metric changes in complex e-commerce systems, focusing on trade-offs between interpretability, efficiency, and causal validity. As a starting point, we extend a...

💬 0 commentsarXiv:2608.29735v1PDF
0

Posted in stat.ML · 2026-08-30 · Zhe Aurore Li, Quentin Clairon, Cécilia Samieri, Rodolphe Thiébaut, Mélanie Prague, Cécile Proust-Lima

Neural ODE enhanced linear mixed effect models for estimating complex association patterns of time-varying covariates with the marker trajectory

Longitudinal cohort studies produce repeated data that enable the assessment of time-varying association patterns between exposures and health outcomes. Classical linear mixed-effects models (LMMs) can accommodate a large variety of association patterns while accounting for the irregularly spaced, partially observed measurement. But...

💬 0 commentsarXiv:2608.29714v1PDF
0

Posted in stat.ME · 2026-08-31 · Jiaye Chen, Rui Qiu, Roulin Wang, Zhou Yu

Marginal Coordinate Test for Fréchet Regression with Random Objects

We develop a marginal coordinate test for regression with Euclidean predictors and a random-object response in a separable metric space. The goal is to test whether a predictor provides additional information about the response conditional on the remaining predictors. In a semi-supervised design, an unlabeled sample is used to...

💬 0 commentsarXiv:2608.30644v1PDF
0

Posted in stat.OT · 2026-08-30 · Anders Gorst-Rasmussen

Statistical Leadership of What? Statistics After AI

Statisticians have spent over a century arguing that we are more than calculators, usually by pointing to what else we know. AI is making that defense harder, since the list of what only statisticians can do grows shorter with each model release. AI makes claims cheap to generate and may eventually make the statistics behind them...

💬 0 commentsarXiv:2608.29629v1PDF
0

Posted in stat.ME · 2026-08-30 · Xinbing Kong, Xiaoying Pan, Long Yu, Tong Zhang

One-step group factor analysis via penalized least squares

In this article, we revisit the problem of group factor analysis and propose a one-step penalized least squares method to estimate the factor loadings and factors in large-dimensional group factor models, offering a distinct alternative to the conventional two-step principal component approach. Our procedure originates from the...

💬 0 commentsarXiv:2608.29625v1PDF
0

Posted in stat.ME · 2026-08-30 · Seojin Lee, Neulpum Jeong, Seonghyun Jeong

Adaptive Functional Clustering with Structured Dependence via Variational Inference

Functional clustering is an important tool for identifying latent heterogeneity in functional data and has been widely applied across various scientific fields. However, many existing methods are not fully adaptive, as they may require the number of clusters to be prespecified and may lack automatic control over the smoothness of the...

💬 0 commentsarXiv:2608.29619v1PDF
0

Posted in stat.ME · 2026-08-30 · Kazuharu Harada, Mitsunori Ogawa

Prognosis-equivalent mapping of clinical measurements via survival analysis

A continuous clinical measurement recorded in fixed physical units may have different prognostic meaning across patients when its effect depends on a patient-level modifier. For example, the same tumor diameter may imply markedly different prognosis in an infant and an adult, because patient size can modify its prognostic effect. We...

💬 0 commentsarXiv:2608.29614v1PDF
0

Posted in stat.ML · 2026-08-29 · Minxing Zheng, Holly Wiberg, Shixiang Zhu

Deciding When to Decide: Testing Operational Suboptimality Under Distributional Shift

Deployed decisions are often optimized once and retained because updates impose operational, regulatory, or switching costs. As operating conditions change, when should such decisions be re-optimized? We study this question for stochastic optimization when the objective's functional form is known but the decision maker's trade-offs...

💬 0 commentsarXiv:2608.29465v1PDF