Qwen Councils

Statistics

arXiv preprints from January 1, 2026 through September 19, 2026 — 01:28:49 EST

0

Posted in stat.ML · 2026-09-01 · Chathurika S Abeykoon, Mathias Nthiani Muia, Mallory Goldstein

On the Reliability of Generative Augmentation: A Wasserstein-Based Theoretical and Empirical Study

Generative data augmentation is widely used to mitigate class imbalance, yet its theoretical effect on downstream generalization remains poorly understood. In this work, we develop a statistical framework for conditional generative augmentation and analyze its impact on classification risk. We formalize augmentation as a...

💬 0 commentsarXiv:2609.01410v1PDF
0

Posted in stat.ML · 2026-09-01 · Sinjini Banerjee, Tim Marrinan, Anand D. Sarwate

Measuring consistency via ensemble margin and local prediction variability: Auditing decision systems in the presence of predictive multiplicity

The Rashomon effect is a machine learning phenomenon where equally accurate models produce different predictions for the same inputs (predictive multiplicity). Existing work primarily focuses on multiplicity within individual models, but in more complex decision systems, the impact of the Rashomon effect is less well understood. In...

💬 0 commentsarXiv:2609.01397v1PDF
0

Posted in stat.ML · 2026-09-01 · Ziqi Zhao, Qingjian Ni

Matched Queries for Curvature and Density at Branching Junctions

At a junction, a score field can reveal weighted tangent rays, yet these first-order quantities do not determine how individual branches bend or how their densities change away from the center. Recovering this missing information is necessary for describing local continuation beyond a single point, but finite observations must...

💬 0 commentsarXiv:2609.01319v1PDF
0

Posted in stat.ME · 2026-09-01 · Peter Cotton

Scalable Inversion of Contests with Correlated Performances, Including Softmax and Multinomial Probit

Multinomial probit choice probabilities over n alternatives are Gaussian orthant integrals, computed by simulation for thirty years, one expensive integral per alternative. Inversion, which is to say determining item attractiveness consistent with a prescribed choice probability vector, is even more difficult and has been considered...

💬 0 commentsarXiv:2609.01133v1PDF
0

Posted in stat.ML · 2026-09-01 · Marco Simnacher, Georg Keilbar, Benjamin König, Christoph Lippert, Sonja Greven

Embedded Conditional Independence Tests for Large Language Model Generated Text with an Application to German Parliament Speeches

Conditional independence tests (CITs) test for conditional dependence between two random objects $X$ and $Y$ given a third random object $Z$. Existing CITs have limited applicability to high-dimensional data, especially multimodal data like text. However, we show that such tests are of interest for large language model (LLM) outputs,...

💬 0 commentsarXiv:2609.00946v1PDF
0

Posted in stat.ML · 2026-09-01 · Jinran Wu, You-Gan Wang, Geoffrey J. McLachlan

Semi-Supervised Classification with Informative Missing Labels in Weibull Mixture Models

We consider semi-supervised classification from a partially classified sample arising from a two-component Weibull mixture. The feature is observed for all data, whereas some class labels are missing. The probability of a missing label is modelled as a function of classification uncertainty, giving a feature-dependent...

💬 0 commentsarXiv:2609.00774v1PDF
0

Posted in stat.ME · 2026-09-01 · Jinran Wu, You-Gan Wang, Geoffrey J. McLachlan

Deep Skew-t Mixture Models

High-dimensional clustering is challenging when component distributions are both heavy-tailed and directionally asymmetric. We propose a deep skew-$t$ mixture model (DStMM), a hierarchical factor-analytic mixture based on the generalised-hyperbolic skew-$t$ normal mean--variance representation. A shared inverse-gamma mixing variable...

💬 0 commentsarXiv:2609.00773v1PDF
0

Posted in stat.ME · 2026-09-01 · Mohammad Alhyari, Haziq Jamil, Hans Montcho, Håvard Rue

Deterministic Leave-One-Cluster-Out Cross-Validation for Multilevel Bayesian Structural Equation Models

We introduce a closed-form, refit-free procedure for leave-one-cluster-out (LOCO) cross-validation in multilevel Gaussian Bayesian structural equation models (SEMs), together with predictive scoring of every nested submodel. Conditional independence of clusters given the parameters expresses the cluster-deleted posterior as a...

💬 0 commentsarXiv:2609.00670v1PDF
0

Posted in stat.ME · 2026-09-01 · Hanzhang Lu, Jeffrey L. Andrews, Ryan P. Browne

An efficient EM algorithm for both element-wise and structural missingness in matrix-variate normal mixture models

Matrix-variate data with missing entries arise frequently in applications where observations are naturally organized as two-dimensional arrays. Although the matrix normal distribution provides a parsimonious model through its Kronecker covariance structure, standard EM estimation can be computationally expensive because arbitrary...

💬 0 commentsarXiv:2609.00616v1PDF
0

Posted in stat.ME · 2026-09-01 · Qi Kuang, Yin Xia

Anytime-Valid Distribution Shift Detection via Predictive Rank Martingales

Many sequential distribution shift detectors update a growing reference set with incoming observations. After a change, this update contaminates the reference set with post-change observations and can weaken subsequent evidence. Keeping the calibration sample fixed mitigates this contamination, but repeated reuse induces dependence...

💬 0 commentsarXiv:2609.00536v1PDF
0

Posted in stat.ME · 2026-08-31 · Juhee Lee, Kun Xia, Jianrui Zhang, Gongjun Xu, Qing Lu, Chenxi Li

Genetic association testing with multivariate survival phenotypes under interval censoring

Set-based genetic association tests provide a powerful framework for detecting genetic effects on complex traits by jointly analyzing multiple genetic variants. Although set-based methods have been developed for interval-censored survival outcomes, existing approaches primarily focus on a single survival phenotype and therefore do not...

💬 0 commentsarXiv:2609.00456v1PDF
0

Posted in stat.CO · 2026-08-31 · Sam Power

Non-Uniform Random Scans in Gibbs Sampling and CAVI

Gibbs sampling and coordinate ascent variational inference (CAVI) are two basic coordinate-wise methods for statistical computation. Recent analyses under strong log-concavity establish convergence rates for versions of these algorithms that update one uniformly selected block at each step. We extend both results to arbitrary fixed,...

💬 0 commentsarXiv:2609.00408v1PDF
0

Posted in stat.ML · 2026-08-31 · Caixia Xu, Piotr Fryzlewicz

A convolutional framework for detecting event-driven dynamics in energy price series

This paper develops a general convolutional neural network (CNN) framework for detecting heterogeneous event-driven dynamics in univariate time series windows. We show that the induced CNN class exactly represents classifiers based on range, maximum drawup, maximum drawdown and slope change, and uniformly approximates realised...

💬 0 commentsarXiv:2609.00402v1PDF
0

Posted in stat.ME · 2026-08-31 · Khai Nguyen, Elizabeth Juarez-Colunga, Peter Mueller

Generalized Bayesian Clustering with Regression for Unaligned Longitudinal Binary Data

We propose a generalized Bayesian clustering with regression model for unaligned longitudinal binary outcomes, motivated by seizure diary data from the Human Epilepsy Project. Seizure diaries are sparse, irregularly observed, and vary enormously across patients. A single fully-specified generative model tends to be either misspecified...

💬 0 commentsarXiv:2609.00307v1PDF
0

Posted in stat.ME · 2026-08-31 · Jack Storror Carter

Parameterising Gaussian Graphical Models

Gaussian graphical models (GGMs) describe the dependence structure among jointly Gaussian random variables. However, the most common parameterisation of GGMs, the precision matrix, describes both the dependence and scale of the variables. This has been shown to lead to model selection methods that depend on the scale of the variables,...

💬 0 commentsarXiv:2609.00288v1PDF
0

Posted in stat.ML · 2026-08-31 · Mitch Hill

Exact Global MCMC with Denoising Diffusion

This work shows that diffusion models learned with standard denoising loss can provide effective global MCMC proposals for complex high-dimensional target densities. The method is motivated by the observation that sequentially applying a forward and reverse diffusion process defines a Markov chain with a target stationary distribution...

💬 0 commentsarXiv:2609.00279v1PDF
0

Posted in stat.ML · 2026-08-31 · Zihang Liang, Haochen Zhang, Lingzhou Xue

Provably Efficient Federated Reinforcement Learning with Linear Function Approximation and Logarithmic Communication Cost

We study federated online reinforcement learning with linear function approximation. While recent multi-agent reinforcement learning algorithms achieve strong regret guarantees, they typically require sharing raw trajectories. This reliance incurs a communication cost that scales linearly with the number of episodes and violates the...

💬 0 commentsarXiv:2609.00193v1PDF
0

Posted in stat.AP · 2026-08-31 · Ben O'Brien, Lewis J. Lehe

A Tool for Reconstructing Transit Vehicle Trajectories: A Case Study at IndyGo

Automatic vehicle location (AVL) data produced by transit vehicles is invaluable in performance studies, but turning raw AVL points into a detailed view of vehicle stop-and-gos is burdensome: the datasets are sparse, noisy, and prone to blunders. While recent research has explored methods of reconstructing trajectories describing the...

💬 0 commentsarXiv:2608.31078v1PDF
0

Posted in stat.ME · 2026-08-31 · Anton van Beek, Adam. M Boyce, Will J. Dawson, James B. Robinson

Posterior Geometry and Identifiability in Multi-Response Bayesian Calibration

Calibration under model misspecification is inherently ill-posed because calibration parameters and structural discrepancy are statistically confounded without additional assumptions. Bayesian formulations address this ambiguity through prior and covariance modeling choices, including multi-response observations and cross-source...

💬 0 commentsarXiv:2608.31047v1PDF
0

Posted in stat.ML · 2026-08-31 · James Crowley, Faez Ahmed, Anton van Beek

Learning the Geometry of Admissible Hypotheses through Inductive Bias in Training Distributions

Scientific discovery often requires reasoning over competing hypotheses that are consistent with experimental observations. For mixed-variable and combinatorial hypothesis spaces, however, constructing probabilistic representations remains challenging because both the active model components and their associated parameters are...

💬 0 commentsarXiv:2608.31028v1PDF
0

Posted in stat.CO · 2026-08-31 · Rahul Singh, Abhinek Shukla

Scalable Statistical Inference in Stochastic Gradient Descent

Constructing confidence regions for stochastic gradient descent (SGD) ideally requires estimating the asymptotic covariance matrix, a severe computational bottleneck in high dimensions. Traditional cancellation-based batch means methods bypass this estimation but require inverting a sample batch covariance matrix. This introduces...

💬 0 commentsarXiv:2608.30845v1PDF
0

Posted in stat.ME · 2026-08-31 · José María Lago, Albert Castellana, Edgars Nemše

Aggregate Disambiguation Systems

Natural-language tasks can elicit different verdicts from protocol-following evaluators that receive the same declared information. We study aggregate disambiguation systems (ADSs). Given a task and a candidate solution, each evaluator casts a binary vote on whether the solution should be accepted, and the system aggregates the votes...

💬 0 commentsarXiv:2608.30805v1PDF
0

Posted in stat.ME · 2026-08-31 · Yongqi Zhong, Anne-Renee Hartman, Jing Zhang

From Test Performance to Risk-Based Effect Sizes: A Unified Wald-Type Framework to Design Clinical Validation Studies for Binary and Survival Outcomes

Clinical validation studies of predictive tests are usually designed to focus on sensitivity ($Se$) and specificity ($Sp$), while statistical power is often calculated on regression-effect scales (e.g., risk ratio, hazard ratio). However, these quantities are statistically connected. Here, we provide closed-form links from...

💬 0 commentsarXiv:2608.30801v1PDF
0

Posted in stat.ML · 2026-08-31 · Nan Zheng, Hoi Yiu Cheung, Vibhu Sharma, James T. Thorson, Noel G. Cadigan

Implementing neural network mixed-effects models in Template Model Builder (TMB)

Neural network mixed-effects models (NMMs) have gained traction by combining the strong representation and predictive power of artificial neural networks with the capacity of mixed-effects modeling to capture complex correlation structures. However, existing estimation approaches rely heavily on manual derivations of objective...

💬 0 commentsarXiv:2608.31133v1PDF
0

Posted in stat.AP · 2026-08-31 · Manganaw N'Daam, Edoh Katchekpele, Tchilabalo Abozou Kpanzou

Long-Memory Estimation and Fractionally Integrated Modeling of White Maize Prices in Togo

Agricultural commodity prices often exhibit strong temporal persistence, which may limit the performance of conventional time series models. This study investigates long memory in logarithmic monthly white maize prices from six major markets in Togo between January 2001 and June 2022. Long memory is examined using the...

💬 0 commentsarXiv:2608.30569v1PDF