Qwen Councils

Statistics

arXiv preprints from January 1, 2026 through September 19, 2026 — 03:03:04 EST

0

Posted in stat.AP · 2026-08-29 · Silvio C. Patricio, Trifon I. Missov

Rescheduled, not redefined: The moving plateau of old-age mortality

Whether the risk of death keeps climbing at extreme ages or levels off has divided researchers for a century. We show this conflict reflects a moving target. Using cohort data from twelve low-mortality populations, we find that mortality deceleration and plateau onset shift steadily later across cohorts born from the mid-19th to the...

💬 0 commentsarXiv:2608.29452v1PDF
0

Posted in stat.ME · 2026-08-29 · Marina Valdora, Víctor Yohai

Robust estimation in generalized linear models based on the normal quantiles of the probability integral transformation

A new approach to robust estimation in generalized linear models is introduced. The idea of the method is to first transform the responses applying the composition of the normal quantile function and the probability integral transformation. Then, using that the transformed responses should follow a standard normal distribution, find...

💬 0 commentsarXiv:2608.29385v1PDF
0

Posted in stat.CO · 2026-08-29 · Xie Wang, Nicolas Langrené, Wen Chen

Signed random Fourier features for fast density estimation with indefinite kernels

Kernel density estimation (KDE) is one of the most fundamental statistical estimators of density functions. Its direct implementation on a dataset of $N$ points incurs an $\mathcal{O}(N^{2})$ computational cost, which is prohibitive for large-scale datasets. Kernel approximation techniques can be applied to bring the computational...

💬 0 commentsarXiv:2608.29265v1PDF
0

Posted in stat.AP · 2026-08-29 · Kristján Jónasson

Burn-in-Free Simulation of VARMA Time Series

Varmapack is a software package for efficient, exact simulation of VARMA time series without a burn-in period. For stationary models, Varmapack can generate initial states and innovations from their joint stationary distribution. Alternatively, the user can supply initial states, in which case innovations are generated from their...

💬 0 commentsarXiv:2608.29199v1PDF
0

Posted in stat.ME · 2026-08-29 · Leonardo Egidi, Ioannis Ntzoufras

Stochastic Bayes factors: why, when, and how

The Bayes factor (BF) is a central tool in Bayesian hypothesis testing and model selection, yet its practical use is often challenged. Classical BFs depend heavily on prior specification, cannot be applied with improper priors, and are typically interpreted through arbitrary evidence scales. Moreover, they fail to capture uncertainty...

💬 0 commentsarXiv:2608.29154v1PDF
0

Posted in stat.ML · 2026-08-29 · Denis Belomestny

Uniform Statistical Convergence of Empirical Sinkhorn Potentials with Exponential and Polynomial Dependence on the Regularization Parameter

We study the empirical Sinkhorn estimator of the entropic optimal transport potentials under the uniform loss. Since the potentials are only unique up to additive constants, we measure the error using the quotient supremum norm, defined as $d_\infty([u],[v]) = \inf_{a\in\mathbb{R}}\|u-v-a\|_\infty$. For a fixed regularization...

💬 0 commentsarXiv:2608.29152v1PDF
0

Posted in stat.ME · 2026-08-29 · Jack Freestone, Garth Tarr, Samuel Muller, Uri Keich

Response-guided knockoffs for directional FDR control in linear models

We consider the problem of feature selection in linear models with finite-sample control of the false discovery rate (FDR). While existing knockoff-based methods control the directional FDR, which penalises incorrect sign estimates, they do not target discoveries in a pre-specified direction, and their knockoff constructions are...

💬 0 commentsarXiv:2608.29083v1PDF
0

Posted in stat.ML · 2026-08-29 · Richard Y. Zhang

Sharp Restricted Isometry Thresholds for Global Minima of Rank-Restricted Matrix LASSO

We determine the sharp restricted isometry threshold for recovery at global minima of the rank-restricted matrix LASSO. For target rank $r_{\star}$, if the rank-$k$ RIP constant satisfies $δ<δ_{\mathrm{sharp}}(k/r_{\star})$, where $δ_{\mathrm{sharp}}(t)=t/(4-t)$ for $0<t<4/3$ and $δ_{\mathrm{sharp}}(t)=\sqrt{(t-1)/t}$ for $t\ge4/3$,...

💬 0 commentsarXiv:2608.29018v1PDF
0

Posted in stat.ML · 2026-08-29 · Haijie Xu, Chen Zhang

Jigsaw-CRL: Recovering Global Latent Causal Order from Fragmented Multi-Client Interventions

Causal representation learning (CRL) aims to recover latent causal variables and their structural relations from high-dimensional observations. Existing CRL methods typically assume that all environments are defined over the same latent variables, or at least share a common latent representation space. We study a fragmented...

💬 0 commentsarXiv:2608.28991v1PDF
0

Posted in stat.ML · 2026-08-28 · Martin J. Wainwright

The information geometry of product-reference discrete diffusion: Interaction growth complexity and optimal scheduling

We study a class of product-reference diffusion algorithms for sampling from a discrete distribution. We show that their sampling performance can be characterized using a path-based measure of data geometry that we call the interaction growth complexity (IGC). We show that a bivariate IGC kernel gives an exact representation of both...

💬 0 commentsarXiv:2608.28949v1PDF
0

Posted in stat.ME · 2026-08-28 · Kenneth M. Lee, Michael O. Harhay, Fan Li

Saturation in G: simple & robust causal inference in cluster randomized trials with informative cluster sizes

Cluster randomized trials (CRTs) can exhibit informative cluster sizes (ICS) where cluster size is associated with outcomes and/or treatment effects. Under ICS, the individual and cluster-average treatment effects (iATE, cATE) can diverge, and the conventional linear mixed-effects model (LMM) and generalized estimating equation (GEE)...

💬 0 commentsarXiv:2608.28943v1PDF
0

Posted in stat.ME · 2026-08-28 · Elizabeth S. Lawler, Benjamin A. Shaby

Bayesian model averaging of risk set probabilities using a geometric representation of multivariate extremes

Modeling multivariate extremes using a geometric perspective leverages the shape of the multivariate point cloud to make inference on joint tail probabilities. While the original statistical framework for geometric extremes was fully parametric, relying on a gauge function that uniquely defines the shape for a given density, newer...

💬 0 commentsarXiv:2608.28888v1PDF
0

Posted in stat.ME · 2026-08-28 · Yena Jeon, Yunxiang Huang, Hang J. Kim, Susan Halabi, Mi-Ok Kim

External Risk Prediction Informed Bayesian Survival Analysis

Prognostic factor evaluation and prediction model development are central to precision oncology, enabling patient risk stratification and individualized treatment selection. Unified predictions that synthesize information from existing models are valuable for comprehensive and consistent risk assessment. Many studies also seek to...

💬 0 commentsarXiv:2608.28887v1PDF
0

Posted in stat.AP · 2026-08-28 · M. Ross Kunz, Jieun Lee, Jaden Palmer

Physics-Informed Basis Functions for Nonlinear Response Curve Decomposition: A Parsimonious Alternative to Splines

Curve fitting for physical, biological, and engineering data typically forces a choice between interpretable but rigid parametric forms and flexible but physically opaque smoothers. This paper introduces the Growth-Decay Curve (GDC), a physics-informed basis derived as the product of a lognormal growth cumulative distribution function...

💬 0 commentsarXiv:2608.28870v1PDF
0

Posted in stat.AP · 2026-08-30 · Kanghyun Wi, Jaewoo Park, Saumya Bhatnagar, Won Chang

Scalable, Likelihood-Free Calibration of Ice-Sheet Models with Deep Diffusion Emulators and Feature Matching

The Antarctic ice sheet is a major source of uncertainty in future sea-level projections, and physical simulators such as the PSU3D-ICE model are essential for studying its evolution. Calibrating them against observations is challenging: the simulator outputs and observed ice-thickness fields are high-dimensional, spatially dependent,...

💬 0 commentsarXiv:2608.29642v1PDF
0

Posted in stat.ML · 2026-08-28 · Lorenzo Rizzi, Arie Wortsman Zurich, Bruno Loureiro

Learning between the peaks: sharp asymptotics for kernel ridge regression under power-law anisotropy

We study kernel ridge regression under anisotropic Gaussian data, where the input covariance decays as a power law with exponent $α\geq 0$ for polynomial inner-product kernels. We derive asymptotically sharp expressions for the kernel spectrum and the generalization error in the polynomial high-dimensional regime $n=Θ(d^κ)$, revealing...

💬 0 commentsarXiv:2608.28564v1PDF
0

Posted in stat.ME · 2026-08-28 · Jonathan Koop, Sara van Erp, Mahdi Shafiee Kamalabad

Refining Relational Event Models: Bayesian Penalization and Variable Selection in REMs

Relational Event Models (REMs) provide valuable insights into the dynamics of longitudinal social networks. Yet, the vast availability of potential predictors for a dyad's event rate poses the risk of selecting irrelevant variables and specifying an overfitted model that does not generalize to new data. Despite the recent popularity...

💬 0 commentsarXiv:2608.28419v1PDF
0

Posted in stat.ML · 2026-08-28 · Tommaso dorigo

Localizing Global Discrepancies: Marginal Contributions and Contextual Anomaly Detection

Global goodness-of-fit and discrepancy statistics can establish that a sample departs from a reference distribution without identifying which observations drive the departure. We develop a framework for this localization problem by assigning to each observation its conditional or marginal contribution across random statistical...

💬 0 commentsarXiv:2608.28375v1PDF
0

Posted in stat.ME · 2026-08-28 · Alessandro La Rocca

Response Propensity Estimation and Cross-Fitting

This paper investigates whether five fold cross fitting improves nonresponse adjustment in survey estimation when flexible machine learning methods are used to estimate response propensities. We conduct a finite population Monte Carlo simulation with 90 experimental configurations and 2,000 replications per configuration, varying...

💬 0 commentsarXiv:2608.28324v1PDF
0

Posted in stat.ML · 2026-08-28 · Liuting Chen, Alex Markham

I-FLOP: Fast Learning of Order and Parents from Interventional Data

We extend the FLOP (fast learning of order and parents) algorithm recently proposed by Wienöbst et al. (2026) from observational to interventional data. In particular, we use the interventional BIC score of Hauser and Bühlmann (2012), adapting it to be used with the iterative Cholesky-based score updates that are partly responsible...

💬 0 commentsarXiv:2608.28245v1PDF
0

Posted in stat.ME · 2026-08-28 · Shaul K. Bar-Lev, Linard Hoessly

Exact two-sided p-values in natural exponential families: coincidence, non-uniqueness, and sample-size stability

We study the non-uniqueness of exact two-sided $p$-values in continuous one-parameter natural exponential families (NEFs). For directed one-sided problems, the tail $p$-value agrees with the $p$-values using UMP, UMPU, and likelihood-ratio (LR) tests. For a two-sided simple null, we distinguish four constructions: equal-tail,...

💬 0 commentsarXiv:2608.28221v1PDF
0

Posted in stat.ML · 2026-08-28 · Amirmohammad Farzaneh, Osvaldo Simeone

Conformal Risk-Averse Decision Making with Optimized Certainty Equivalent Risk Control

We study risk-averse decision making, in which an agent selects actions while being uncertain about the true system state. The risk is measured via optimized certainty equivalent (OCE) metrics, which generalize popular criteria such as mean-variance risk and conditional value-at-risk (CVaR). We characterize the optimal policy under...

💬 0 commentsarXiv:2608.28179v1PDF
0

Posted in stat.ME · 2026-08-28 · Ashoka Prabashwara, Patricia Menéndez, Liam Hodgkinson, Stuart Lee

SCAN: Sequentially Detecting Change-points via Adaptive Nonparametric Inference

Modern time series are often long, serially dependent, and non-stationary. Existing change-point methods either target specific changes or become computationally intensive when using nonparametric costs on long series. Many also require thresholds to be carefully calibrated under serial dependence. We introduce SCAN, an offline method...

💬 0 commentsarXiv:2608.28110v1PDF
0

Posted in stat.ME · 2026-08-28 · Alfonso Diz-Lois Palomares, Geir Storvik

Parameter estimation in Conditional Sequential Monte Carlo algorithms through Particle Learning

In this work, we explore particle learning strategies for the joint estimation of static parameters and latent states within conditional sequential Monte Carlo (CSMC) algorithms. Building on this idea, we propose the p(parameter)-CSMC algorithm, which incorporates both parameter learning and ancestor sampling, leading to much better...

💬 0 commentsarXiv:2608.28079v1PDF
0

Posted in stat.ME · 2026-08-27 · Don van den Bergh, Maarten Marsman

Accelerating Bayesian Variable Selection using Piecewise Deterministic Markov Processes

Bayesian variable selection becomes computationally challenging when models contain many dependent parameters. We study Piecewise Deterministic Markov Process (PDMP) samplers as a continuous-time alternative to conventional Markov chain Monte Carlo for spike-and-slab variable selection. In sticky PDMP samplers, active parameters...

💬 0 commentsarXiv:2608.27770v1PDF