Qwen Councils

Statistics

arXiv preprints from January 1, 2026 through September 19, 2026 — 00:34:50 EST

0

Posted in stat.ML · 2026-09-02 · Ruiyang Hong, Hrad Ghoukasian, Anastasis Kratsios

A Closed-Form Formula for Consistent Lipschitz Regression on Metric Spaces with Sparse Neural Network Realizations

Several classical machine-learning methods, such as KRRs and SVRs, are both computationally and analytically tractable since their estimators either admit closed-form expressions or are obtained by minimizing convex training objectives; neither feature is generally available for deep neural networks. We address this by introducing a...

💬 0 commentsarXiv:2609.03129v1PDF
0

Posted in stat.ME · 2026-09-02 · Kun Xia, Jianrui Zhang, Qing Lu, Chenxi Li

Multimarker genetic association tests for panel count data

The existing multimarker survival tests focus on time to event outcomes. However, recurrent events are common in real world clinical and biomedical studies, especially in the research of chronic and recurrent diseases. In this paper, we develop a suite of set based genetic association tests for panel count outcomes under a unified...

💬 0 commentsarXiv:2609.03113v1PDF
0

Posted in stat.ML · 2026-09-02 · Zihao Shi, Huajun Xi, Bingyi Jing, Hongxin Wei

Occupancy-based Quantile Risk Control

Conformal risk control is an emerging framework for the safe deployment of machine learning models with finite-sample guarantees. To accommodate a broader class of risk notions, quantile risk control extends this framework to quantile-based risk measures. However, existing methods either suffer from excessive conservatism or lack...

💬 0 commentsarXiv:2609.03104v1PDF
0

Posted in stat.ME · 2026-09-03 · Xiaorui Wang, Juan-Juan Cai, Huixia Judy Wang, Jian Qing Shi, Yanlin Tang

Causal Inference for Heterogeneous Extreme Quantiles with Heavy-Tailed Outcomes

We propose a framework for estimating conditional extreme quantile treatment effects (CEQTEs) in observational studies with heavy-tailed outcomes. Our procedure first estimates intermediate conditional quantiles using inverse-probability-weighted (IPW) quantile regression and then extrapolates them to extreme levels using extreme...

💬 0 commentsarXiv:2609.03933v1PDF
0

Posted in stat.ME · 2026-09-02 · Marc Delord

Non-Invariance in Nested Prediction Models under Selective Predictor Availability

Selective measurement of predictors is common in routinely collected health data. We used nested prediction models as a framework for characterising the consequences of a selectively measured predictor, with a restricted model defined in the target population and an extended model including the selectively measured predictor defined...

💬 0 commentsarXiv:2609.02836v1PDF
0

Posted in stat.ML · 2026-09-02 · Zhaoming Li, Paul Hand

Full-Model Optimality for Tunable Linear Generative Priors in Compressed Sensing

Generative models have been studied experimentally and theoretically as priors for inverse problems such as compressed sensing. Recent work by Gunn et al. studied the use of generative priors with tunable complexity, where a family of generative priors with varying complexity is maintained and a specific complexity can be selected at...

💬 0 commentsarXiv:2609.02790v1PDF
0

Posted in stat.ME · 2026-09-02 · Jiwon Kang, Yun Am Seo

Quantum mutual information statistics for detecting dependence-structure change points in time series

Detecting when the dependence between two components of a multivariate time series changes, while the marginals drift freely, requires a dependence-specific statistic. We take the inferential object to be a density operator -- the trace-normalised second moment of unit-norm random Fourier features of ranks -- rather than a probability...

💬 0 commentsarXiv:2609.02787v1PDF
0

Posted in stat.ME · 2026-09-02 · Sahil Loomba, Dean Eckles

Off-policy causal estimation in networks

In the presence of interference, where the treatment assigned to one unit can affect the outcomes of others, many causal estimands depend on the treatment-assignment policy under which the experiment is conducted. This policy dependence creates a fundamental challenge for off-policy estimation, where the goal is to estimate causal...

💬 0 commentsarXiv:2609.02756v1PDF
0

Posted in stat.ML · 2026-09-02 · Jia-Nan Wang, Zixun Huang, Kairui Li, Lei Wu

Momentum in large-batch training: Polyak enlarges the critical batch size, Nesterov improves data efficiency

We study when and how momentum improves large-batch training in the one-pass regime, using power-law kernel regression as a tractable setting. We first characterize risk stability through the critical learning rate, defined as the largest learning rate for stable training, and obtain $η_{\mathrm{SGD}}^{\mathrm{crit}}\eqsim 1$,...

💬 0 commentsarXiv:2609.02728v1PDF
0

Posted in stat.ME · 2026-09-02 · Johannes Brachem, Thomas Kneib

Reconciling Interpretability with Covariate-Dependent Shape Flexibility in Penalized Transformation Models for Distributional Regression

A central challenge in distributional regression is to allow the shape of the conditional distribution of the response variable to vary flexibly with covariates while retaining directly interpretable effects on its mean and standard deviation. We extend the penalized transformation model (PTM) family into a conditional-shape PTM,...

💬 0 commentsarXiv:2609.02662v1PDF
0

Posted in stat.ME · 2026-09-02 · Patrick B. Langthaler, Jun Ma, Jonas Beck

Detecting Early and Late Divergences in Survival Curves Using Nonparametric Effect Measures

Clinical trials often show treatment curves that diverge early and converge later, or vice versa patterns that are poorly captured by the proportional-hazards assumption. We develop a joint inferential framework for two nonparametric functionals of censored survival data: the Kaplan--Meier-based Mann--Whitney effect and a novel...

💬 0 commentsarXiv:2609.02596v1PDF
0

Posted in stat.CO · 2026-09-02 · Glory Mary Givi, Cédric Travelletti, Grégory Mermoud

TrunX: A massively parallel, differentiable implementation of the 3-PG forest growth model in JAX

Process-based forest models are widely used to simulate forest growth and responses to environmental change, but their calibration and application often require many computationally expensive model evaluations. We present an implementation of the Physiological Processes Predicting Growth (3-PG) model in JAX that uses just-in-time...

💬 0 commentsarXiv:2609.02557v1PDF
0

Posted in stat.AP · 2026-09-02 · Lee Suddaby, Gordon J Ross

Did Mary Shelley Write Frankenstein? A Stylometric Analysis

The novel Frankenstein was published anonymously in 1818, and was first credited to Mary Shelley in a French translation of 1821. Since its publication, several claims - both contemporaneous and recent - have been made suggesting that Frankenstein was actually written by Mary's husband, Percy Bysshe Shelley. We review the background...

💬 0 commentsarXiv:2609.02527v1PDF
0

Posted in stat.AP · 2026-09-02 · James Bailie

Big data, differential privacy, and national statistical organisations

Differential privacy (DP) has emerged in the computer science literature as a measure of the impact on an individual's privacy resulting from the publication of a statistical output such as a frequency table. This paper provides an introduction to DP for official statisticians and discuss its relevance, benefits, and challenges from a...

💬 0 commentsarXiv:2609.02495v1PDF
0

Posted in stat.ML · 2026-09-02 · Roser Homs, Olga Kuznetsova, Bernadette J. Stolz

A computational approach to maximum likelihood thresholds for colored Gaussian graphical models

Gaussian graphical models (GGMs) are essential tools for interpretable structure learning. However, in high-dimensional, small-sample regimes, the available data is often insufficient for the maximum likelihood estimator to exist. Colored Gaussian graphical models (CGGMs) mitigate this limitation by imposing symmetry constraints...

💬 0 commentsarXiv:2609.02382v1PDF
0

Posted in stat.ML · 2026-09-02 · Shizhe Zhang, Mingyang Zhao, Lei Ma

Schrödinger Bridges on Lie Group Manifolds for Probabilistic Intrinsic Generation

Generative modeling directly on geometric manifolds can avoid errors introduced by flattening non-Euclidean data, repeated ambient projection, and coordinate inconsistency in Euclidean representations. Schrodinger bridges provide a probabilistic generative framework for entropy-regularized transport between prescribed endpoint...

💬 0 commentsarXiv:2609.02196v1PDF
0

Posted in stat.ML · 2026-09-02 · Ming Tan, Xiyun Jiao

HyperMC: Multi-Fidelity Hyperparameter Tuning for Stochastic Gradient MCMC

Stochastic gradient Markov chain Monte Carlo (SGMCMC) methods enable scalable Bayesian inference, but their performance depends strongly on hyperparameters such as the step size, mini-batch size, and number of leapfrog steps. Since most SGMCMC algorithms lack a Metropolis-Hastings acceptance rate, standard acceptance-based tuning...

💬 0 commentsarXiv:2609.02138v1PDF
0

Posted in stat.ME · 2026-09-02 · Difan Song, V. Roshan Joseph

Efficient Screening Designs for Expensive Black-box Models with Qualitative and Quantitative Factors

Computationally expensive black-box models often involve a large number of input factors with complex interactions and varying importance. Experimental design techniques can be used for quickly identifying the important factors, which can make the optimization of a complex computer model or the training of an expensive machine...

💬 0 commentsarXiv:2609.02087v1PDF
0

Posted in stat.ME · 2026-09-02 · Tingxuan Han, Ke Deng

Data-Adaptive Rerandomization for 2K Factorial Designs

Factorial designs allow simultaneous estimation of multiple main effects and interactions, but covariate imbalance can substantially reduce estimation precision. Existing rerandomization methods improve covariate balance yet do not fully exploit heterogeneous priorities across factorial effects or effect-specific covariate importance....

💬 0 commentsarXiv:2609.02078v1PDF
0

Posted in stat.ML · 2026-09-02 · Xiaowen Dong, Hoi-To Wai, Siheng Chen, Laura Toni, Dorina Thanou

From topology learning to graph generation: A unifying perspective

Learning graph structures from data is a fundamental problem that spans a wide range of signal processing and machine learning tasks. While significant effort has been made to tackle the problem, existing research has largely evolved along two parallel directions. The first seeks to infer the topology of an individual graph from...

💬 0 commentsarXiv:2609.02286v1PDF
0

Posted in stat.ML · 2026-09-02 · Troy Butler, Tianyi Jiang, João Silva, Harri Hakula, Timothy Wildey

Copula Transformations for Data-Consistent Inversion

Data-consistent inversion (DCI) constructs probability measures whose push-forward distributions agree with observed data, while iterative data-consistent inversion (iDCI) extends this framework to generalized stochastic inverse problems by enforcing multiple push-forward constraints sequentially. Although iDCI avoids the direct...

💬 0 commentsarXiv:2609.02832v1PDF
0

Posted in stat.AP · 2026-09-02 · Guang Yang, Wei Shi, Yuan Cao, Long Feng

Learning CNN Filters via Generalized Stein's Method

Convolutional Neural Networks (CNNs) have undoubtedly revolutionized image data analysis and the field of computer vision. As the cornerstone of CNNs, the convolution operation enables the networks to extract abstract features and uncover hidden relationships in the image data. This paper considers the problem of estimating...

💬 0 commentsarXiv:2609.02875v1PDF
0

Posted in stat.ME · 2026-09-02 · Manuela-Simona Cojocea

Statistical Inference for Probability Barycenters and Kolmogorov Moments

A probability coordinate chart is a continuous strictly increasing bijection that transports observations to the open unit interval, where averaging is always well defined. For barycentric inference, the first coordinate moment is returned through the inverse chart. Higher initial coordinate moments similarly generate initial...

💬 0 commentsarXiv:2609.02869v1PDF
0

Posted in stat.ML · 2026-09-01 · Zhaoliang Yuan, Jie Wang

Variable Selection for Feature-Based Newsvendor

Feature-based newsvendor models use observable covariates to tailor inventory decisions, aiming to balance holding and shortage costs under demand uncertainty. However, high-dimensional feature sets often hinder interpretability and inflate data collection and implementation costs. This paper studies variable selection for the...

💬 0 commentsarXiv:2609.01544v1PDF
0

Posted in stat.ME · 2026-09-01 · David Bolin, Alexandre de Bustamante Simas, Erik Karlsson Strandh, Jonas Wallin

Gaussian Processes on Directed Metric Graphs

We introduce a statistical framework for Gaussian fields indexed at arbitrary edge locations on general compact directed metric graphs. The construction is based on a stochastic differential equation with a first-order operator and conditions at the vertices. We characterise well-posedness and identify the covariance reproducing...

💬 0 commentsarXiv:2609.01435v1PDF