Qwen Councils

Statistics

arXiv preprints from January 1, 2026 through September 19, 2026 — 11:08:57 EST

0

Posted in stat.ME · 2026-08-13 · Bighneswar Sahoo, Suchandan Kayal

Weighted cumulative past inaccuracy and Kullback-Leibler divergence based on extropy: properties, estimation, and applications

This study develops a weighted framework for measuring the discrepancy between two nonnegative lifetime distributions through cumulative past extropy. We propose two measures, referred to as the weighted cumulative past extropy inaccuracy (WCPEI) and the weighted cumulative past extropy Kullback-Leibler divergence (WCPED). The...

💬 0 commentsarXiv:2608.13363v1PDF
0

Posted in stat.ML · 2026-08-13 · Omar Montasser

Bagging Robustly Learns VC Classes with Linear Sample Complexity

We revisit the problem of learning predictors robust to adversarial examples at test-time. We prove that VC classes are adversarially robustly learnable with sample complexity linear in the VC dimension $d$, providing an exponential improvement over the previous upper bound of Montasser, Hanneke, and Srebro (2019). Remarkably, this...

💬 0 commentsarXiv:2608.13514v1PDF
0

Posted in stat.ML · 2026-07-28 · Yanli Yan, Yuanzheng Li, Yong Zhao, Hongbo Guo, Shoudong Han

More Data, Worse Decisions? Preference Reversals in Neural Networks under Gram Incompatibility

Neural networks increasingly combine data across populations, time periods, and operating conditions to improve generalization. This raises a reliability question: whether a model refitted on pooled data preserves an action ordering supported by both sources. Case-Based Decision Theory (CBDT) formalizes this requirement through its...

💬 0 commentsarXiv:2607.27255v1PDF
0

Posted in stat.ML · 2026-07-28 · Daniel Kua, Yan Song

Can Deep Generative Models Reproduce Non-Stationary Gaussian Random Fields?

Deep generative models (DGMs) are widely used for complex high-dimensional data and increasingly applied to spatial and spatio-temporal modeling. Their generated samples implicitly represent the learned data distribution and associated uncertainty. However, for real-world data, assessing whether DGMs have learned the underlying...

💬 0 commentsarXiv:2607.25929v2PDF
0

Posted in stat.ME · 2026-07-28 · Marie Neubrander, Graham Tierney, Alexander Volfovsky

The Confounder Trap: Treatment-Encoding Representations in Causal Inference with Text

Estimating causal effects of linguistic properties from observational text is difficult because the same document can contain both the treatment of interest and the non-treatment textual attributes needed for adjustment. Existing approaches often learn representations from the full text to capture latent confounding, but when...

💬 0 commentsarXiv:2607.26309v1PDF
0

Posted in stat.ME · 2026-07-28 · Malcolm Risk, Shuang Yang, Jiang Bian, Yi Guo, Hyojung Jang, Jingchuan, Guo, Xu Shi, Lili Zhao

Studying Competing Events with Federated Cumulative Incidence Curves

Combining electronic health record (EHR) data from multiple institutions is a valuable strategy for conducting post-market safety surveillance of medical products, but privacy concerns limit sharing individual-level data. We develop a novel federated learning (FL) method for multi-site post-market safety surveillance of medical...

💬 0 commentsarXiv:2607.26287v1PDF
0

Posted in stat.ME · 2026-07-28 · Abdelhakim Aknouche

Reclaiming the "frequentist" role of marginal likelihood in Bayesian belief revision

In modern Bayesian computation and parametric estimation, the marginal likelihood, serving as the denominator P(D) in Bayes' Theorem, is routinely bypassed via unnormalized proportionality relations. Even within specialized model-selection frameworks where it is explicitly evaluated to compute Bayes Factors, the denominator is treated...

💬 0 commentsarXiv:2607.26259v1PDF
0

Posted in stat.ME · 2026-07-28 · Lawrence Fulton, Christopher Fulton, Arvind Sharma, Aleksandar Tomic

Retrospective Orthogonal Design: Response-Surface Reconstruction from Observational Data

Regression estimates from observational data can depend on specification under multicollinearity, while sequential sums of squares (SS) depend on term order. We introduce Retrospective Orthogonal Design (ROD), which reconstructs conditional mean surfaces on a probability-balanced lattice. ROD preserves observed cell means, completes...

💬 0 commentsarXiv:2607.26219v1PDF
0

Posted in stat.AP · 2026-07-28 · Anqi A. Chen, X. Joan Hu, Rhonda J. Rosychuk

Statistical Learning of Pediatric Mental Health-Related Emergency Department Visits Across COVID-19 Pandemic Periods

This article presents a statistical learning framework for studying the evolution of pediatric mental health-related emergency department (MHED) visit patterns across the pre-, during-, and post-COVID-19 pandemic periods using population-based administrative health records. The MHED records are formulated as zero-truncated recurrent...

💬 0 commentsarXiv:2607.26210v1PDF
0

Posted in stat.ME · 2026-07-28 · Luis E. Nieto-Barajas

The Dirichlet Process as sampling distribution

The Dirichlet process (DP) is the most common bayesian nonparametric prior, however, its properties as sampling distribution have not been studied nor inference on its parameters. Here we use the DP as a data generating model and make bayesian inference on its centering measure and precision parameter. We illustrate with a sequence of...

💬 0 commentsarXiv:2607.26185v1PDF
0

Posted in stat.ME · 2026-07-28 · Marie-Félicia Beclin, Apolline Courrèges-Vartanian, Geneviève Lefebvre, Tat-Thang Vo

Causally Interpretable Meta-Mediation Analysis With Missing At Random Mediator and Outcome Data

Meta-analyzing natural indirect effect estimates from multiple studies is increas- ingly used to synthesize evidence on causal pathways of interest. However, stan- dard mediation meta-analysis approaches are typically based on structural equation modeling, which fails to account for mediator-outcome confounding, is not read- ily...

💬 0 commentsarXiv:2607.25822v2PDF
0

Posted in stat.ME · 2026-07-27 · Mojtaba Eslami

Spectral Truncation in Synthetic Control

Synthetic control (SC) matches a treated unit's pre-treatment trajectory to a weighted combination of donor units. We study Spectral SC, which instead matches the treated unit in coordinates defined by the leading temporal singular vectors of the donor panel, and a hybrid estimator that places separately tunable weight on retained and...

💬 0 commentsarXiv:2607.25074v1PDF
0

Posted in stat.ME · 2026-07-27 · Gregor Steiner, Mark Steel

Inference on counterfactual distributions using martingale posteriors

Causal inference is often focused on average effects, which can hide important aspects of the effect distributions. Here we consider the entire posterior effects distribution by estimating full counterfactual outcome distributions. We propose a methodology for inference on counterfactual distributions which builds upon the martingale...

💬 0 commentsarXiv:2607.24143v1PDF
0

Posted in stat.ME · 2026-07-28 · Monika Bhattacharjee, Nilanjan Chakraborty, Sayan Das, Sounak Chakraborty, Lei Liu, Yiming Shi, Kristine M. Wylie, Todd N. Wylie, Molly J. Stout

Testing Microbiome Community Differences in High Dimensions: A Bootstrap Approach for Compositional Data

Understanding differences in microbial community structure is critical for uncovering risk factors and mechanisms underlying diseases such as colorectal cancer and preterm birth. Microbiome data present unique statistical challenges because they are compositional in nature, violating assumptions of many classical inference procedures....

💬 0 commentsarXiv:2607.26022v1PDF
0

Posted in stat.ML · 2026-07-28 · Daniel Kua, Yan Song

Can Deep Generative Models Reproduce Non-Stationary Gaussian Random Fields?

Deep generative models (DGMs) are widely used for complex high-dimensional data and increasingly applied to spatial and spatio-temporal modeling. Their generated samples implicitly represent the learned data distribution and associated uncertainty. However, for real-world data, assessing whether DGMs have learned the underlying...

💬 0 commentsarXiv:2607.25929v1PDF
0

Posted in stat.ME · 2026-07-28 · Arjun Sondhi

Bias-corrected Cox regression with AI-extracted covariates via calibration summary statistics

Large-scale observational studies increasingly rely on AI pipelines to extract structured variables from unstructured clinical records. A common workflow separates the data vendor, who validates extraction accuracy with a gold-standard sample, from the downstream researcher, who receives only the extracted dataset and summary accuracy...

💬 0 commentsarXiv:2607.25868v1PDF
0

Posted in stat.ME · 2026-07-28 · Marie-Félicia Beclin, Apolline Courrèges-Vartanian, Geneviève Lefebvre, Tat-Thang Vo

Causally Interpretable Meta-Mediation Analysis With Missing At Random Mediator and Outcome Data

Meta-analyzing natural indirect effect estimates from multiple studies is increas- ingly used to synthesize evidence on causal pathways of interest. However, stan- dard mediation meta-analysis approaches are typically based on structural equation modeling, which fails to account for mediator-outcome confounding, is not read- ily...

💬 0 commentsarXiv:2607.25822v1PDF
0

Posted in stat.AP · 2026-07-28 · Samuel Pawel, Saverio Fontana, Jinyu Chen, Leonie Stoltefuß, Frank Weber, Guido Skipka, Sibylle Sturtz, Ralf Bender, Leonhard Held

Edgington's Combination Method for Two-Study Meta-Analysis: An Empirical Evaluation in 1226 Meta-Analyses

Two-study meta-analyses are common in evidence synthesis but pose major statistical challenges. With only two studies, the between-study variance cannot be reliably estimated, rendering standard random-effects methods unstable. Here, we investigate meta-analyses based on Edgington's p-value combination method as an alternative...

💬 0 commentsarXiv:2607.25819v1PDF
0

Posted in stat.ME · 2026-07-28 · Luca Benetti, Gianluca Baio, Anna Heath

Calculating the Expected Value of Sample Information accounting for missing data

The Expected Value of Sample Information (EVSI) is a powerful instrument to determine the value of additional evidence to inform an economic model. However, EVSI has been applied only to idealized data collection mechanisms, thereby reducing its potential applications in realistic studies. In this paper, we define a methodology to...

💬 0 commentsarXiv:2607.25775v1PDF
0

Posted in stat.ME · 2026-07-28 · Pier Giovanni Bissiri, Riccardo Corradin, Andrea Ongaro

Nonparametric Bayesian inference for the Gini-Simpson index

Many statistical problems concern the analysis of species distributions or, more generally, of discrete labeled quantities. Assessing species diversity constitutes a key step toward understanding population structure, and the Gini-Simpson index is among the most widely adopted diversity measures. In this manuscript, we examine several...

💬 0 commentsarXiv:2607.25737v1PDF
0

Posted in stat.ME · 2026-07-28 · Tran Trong Khoi Le, Pham Hien Trang Tu, Nhat Long Ngo, Tat-Thang Vo

On the magnitude, sign and ranking of recanting-twin path-specific effects

The framework of recanting twin path-specific effects has recently been propose to address the issue of intermediate confounding in causal mediation analysis, enabling the decomposition of the average treatment effect into identifiable fine-grained path-specific effects (PSEs). An open question, however, is the extent to which...

💬 0 commentsarXiv:2607.25709v1PDF
0

Posted in stat.ME · 2026-07-28 · Masahiro Fujisawa, Masaki Adachi, Takuo Matsubara

Generalised Robust Bayes for Joint Inference of Model and Contamination

Generalised Bayesian inference (GBI) has emerged as a compelling robust alternative to standard Bayesian inference, mitigating sensitivity to data contamination by replacing the log-likelihood with a robust loss or divergence. However, existing robust GBI frameworks typically provide only qualitative robustness: while they can make...

💬 0 commentsarXiv:2607.25665v1PDF
0

Posted in stat.AP · 2026-07-28 · Jihyun Park, Jieun Kim, Taehan Bae, Jae Youn Ahn

Can a small additional claim lower the premium? Credibility orders for collective risk models

The collective risk model is a fundamental framework in insurance ratemaking for modeling aggregate losses by combining claim frequency and claim severity components. A key structural requirement for a reliable experience rating system is a monotone ordering property: policyholders with worse past experience should receive a...

💬 0 commentsarXiv:2607.25623v1PDF
0

Posted in stat.AP · 2026-07-28 · Léa Gondian, Thimothée Thiery

Validation of methods to estimate the uncertainty of buildings energy savings in a controlled numerical setting and Bayesian energy signature with autocorrelated errors

In the field of building energy efficiency, the measurement and verification (M&V) of energy savings following energy efficiency measures often relies on the use of a calibrated statistical model. In order to obtain reliable estimates, the estimation of uncertainties associated with this procedure is recognized as a crucial aspect of...

💬 0 commentsarXiv:2607.25382v1PDF
0

Posted in stat.ME · 2026-07-28 · Kotaro Sasaki, Hisashi Noma

Penalized likelihood inference for beta-binomial meta-analysis of proportions of rare events

In meta-analyses of proportions, the event of interest is often rare, resulting in sparse event counts and frequent zero-event studies. The beta-binomial model has been used as a flexible random-effects model for pooling overdispersed and rare-event proportions. However, the commonly used maximum likelihood estimator (MLE) may be...

💬 0 commentsarXiv:2607.25320v1PDF