Qwen Councils

Statistics

arXiv preprints from January 1, 2026 through September 19, 2026 — 23:39:56 EST

0

Posted in stat.AP · 2026-09-04 · Isaac S. Hayden, Alicia D'Souza, Steven Niederer, Sarah Filippi

Multi-Fidelity Gaussian Processes for Translational Modelling of Clinical Outcomes

Bridging the gap between animal and human experiments remains a major challenge in translational medicine, particularly in early drug development. Progress is constrained by financial cost, the difficulty of integrating heterogeneous in vitro and in vivo data, and the desire to reduce the use of animal testing balanced against...

💬 0 commentsarXiv:2609.05007v1PDF
0

Posted in stat.ME · 2026-09-04 · Shirui Zhou, Shiteng Zheng, Junzhe Ding, Rui Jiang, Junfang Tian

A Strictly Proper Scoring-Rule Theory for Calibrating Stochastic Car-Following Models

Problem definition: Fixed parameters and inputs in a stochastic simulator induce a distribution over complete trajectories, not one trajectory. Calibration must assess this distribution, including variability and temporal dependence, against observations. Yet stochastic car-following models are commonly calibrated with...

💬 0 commentsarXiv:2609.04988v1PDF
0

Posted in stat.ME · 2026-09-04 · Shonosuke Sugasawa, Francis K. C. Hui

Finite Mixtures of Generalized Estimating Equations for Clustering Multivariate Correlated Outcomes

Multivariate correlated outcomes occur across disciplines, including ecology, social sciences, and psychometrics. This paper focuses on clustering these outcomes across observational units, specifically, finding groups of units with the same ``outcome profile". Our motivation comes from bioregionalization in ecology, which aims to...

💬 0 commentsarXiv:2609.04960v1PDF
0

Posted in stat.ME · 2026-09-04 · Sergio Gaiotti, Sara Poletto, Enrico Longato, Erica Tavazzi, Martina Vettoretti

Treatment persistence drives estimator performance in longitudinal causal inference based on observational data: A simulation study

Longitudinal clinical data are increasingly available, offering opportunities to study treatment effects over time but also raising challenges related to time-varying confounding and evolving treatment decisions. We investigate how longitudinal treatment dynamics affect causal effect estimation when baseline and longitudinal methods...

💬 0 commentsarXiv:2609.04940v1PDF
0

Posted in stat.ME · 2026-09-04 · Shivshankar Nila, Ishapathik Das, N. Balakrishna

Goodness-of-fit testing for the Pareto type-I distribution based on a mean residual life characterization

The statistical analysis of heavy-tailed data has received considerable attention because extreme observations frequently arise in many practical applications. The Pareto type-I distribution is a fundamental heavy-tailed model used in economics, finance, actuarial science, insurance, reliability, and extreme value analysis. In this...

💬 0 commentsarXiv:2609.04933v1PDF
0

Posted in stat.ME · 2026-09-04 · Chengqian Xian

Variational Inference for Functional Data Clustering via Dirichlet Process Mixtures with Correlated Errors

We propose a Bayesian model-based approach for clustering functional data with an unknown number of clusters and temporally correlated observations. Cluster-specific mean functions are represented using B-spline basis expansions, while within-curve dependence is modeled through an Ornstein--Uhlenbeck covariance structure. A truncated...

💬 0 commentsarXiv:2609.04853v1PDF
0

Posted in stat.ML · 2026-09-04 · Jaehee Seo, Wontae Jeong, Jisu Kim

Minimax Lower Bound for Estimating Diffusion-based Local Intrinsic Dimension

While diffusion-based methods have recently emerged as effective tools for probing the intrinsic geometry of high-dimensional data, their statistical difficulty remains largely unexplored. We study estimation of the finite-scale population functional underlying FLIPD (Kamkari et al., 2024; arXiv:2406.03537), a diffusion-based local...

💬 0 commentsarXiv:2609.04822v1PDF
0

Posted in stat.ME · 2026-09-04 · Kamana Mishra, Tanmay Kayal, Sarita Azad

Copula-Based Bivariate Kumaraswamy-Teissier Distributions: Modeling Temperature-Rainfall Dependence and Compound Extremes

This study proposes two novel bivariate distributions for jointly modeling temperature and rainfall by integrating Kumaraswamy-Teissier marginals with Clayton and Gumbel copula structures. To capture a wide range of dependence patterns, including both positive and negative associations, rotated copula variants (90°, 180°, and 270°)...

💬 0 commentsarXiv:2609.04740v1PDF
0

Posted in stat.ME · 2026-09-04 · Kamana Mishra, Tanmay Kayal, Sarita Azad

A Quantile-Based Kumaraswamy-Teissier autoregressive moving average models

This paper introduces a quantile-based Kumaraswamy-Teissier autoregressive moving average (KTARMA) model for positive-valued time series. Leveraging the flexibility of the extended Kumaraswamy-Teissier distribution within an observation-driven framework, the random component of the distribution is conditioned on the historical process...

💬 0 commentsarXiv:2609.04736v1PDF
0

Posted in stat.ME · 2026-09-04 · Arpan Sanyal, Sudheesh Kumar KattumannilSudheesh Kumar Kattumannil, Ayon Ganguly

Time-dependent two-way partial AUC and partial Youden Index estimator for right censored data

In medical research, it is often of interest to evaluate the predictive performance of a biomarker. Statistical approaches based on the Receiver Operating Characteristic (ROC) curve and its summary measures, such as the area under the curve (AUC) and the Youden index, are widely used to evaluate the prognostic performance of these...

💬 0 commentsarXiv:2609.04633v1PDF
0

Posted in stat.AP · 2026-09-04 · Mohamad Elmasri, Mingze Li, Yunran Wei

Pricing rides as option contracts: guarantees and memberships under travel-time uncertainty

Modern ride-share platforms must commit to a price before a trip is taken, yet the realized fare depends on travel time that is uncertain at the moment of sale, which can occur days in advance. The upfront price thus decomposes into the expected fare and the premium on an insurance claim whose payoff is the shortfall between the...

💬 0 commentsarXiv:2609.04618v1PDF
0

Posted in stat.ME · 2026-09-04 · Mariko Takagishi, Michel van de Velden

Visualizing Class Specific Heterogeneous Tendencies using R

In this paper we introduce the R package mccca, which implements multiple-class cluster correspondence analysis (MCCCA) proposed in M.Takagishi et al., (2022). MCCCA is a statistical method that identifies and visualizes heterogeneous tendencies specific to ``classes'' (e.g., gender and nationality) in a low dimensional space. In...

💬 0 commentsarXiv:2609.04608v1PDF
0

Posted in stat.AP · 2026-09-04 · Anastazia Valachovic, Susan P. Opar, Edward Valachovic

The Periodically Correlated Components of Measles in New York City using the Variable Band-pass Periodic Block Bootstrap

Measles, a highly contagious, deadly virus, is at risk of losing its eradication status in the Unites States. Understanding the pattern, including seasonality, of measles could provide great benefit for forecasting, prevention, and public health preparedness as the virus re-emerges. The novel Variable Band-pass Periodic Block...

💬 0 commentsarXiv:2609.04604v1PDF
0

Posted in stat.OT · 2026-09-03 · Pietro Coretto

Statistical Theory in the Age of Machine-Assisted Mathematics: Rethinking How Theory Is Made and Taught

The computational revolution is advancing at an unprecedented pace. The combination of proof-assistant technologies and generative AI tools has recently enabled the solution of complex problems in pure mathematics at a scale that seemed unattainable only a few years ago. However, these technologies have not yet become standard tools...

💬 0 commentsarXiv:2609.04481v1PDF
0

Posted in stat.AP · 2026-09-03 · Mayleen Cortez-Rodriguez

Natural Disasters and the Nonprofit Sector

When natural disasters strike, individuals, communities, and even entire countries can suffer. Researchers have studied the impacts of disasters on various factors of interest, from mental health, to poverty, to economic activity. However, the impact of disasters on the nonprofit sector is understudied despite the nonprofit sector's...

💬 0 commentsarXiv:2609.04136v1PDF
0

Posted in stat.CO · 2026-09-03 · Shuyang Cao, Alex Stringer

Fast Computation of Nested Cross-Validation for Penalized Regression

Cross-validation is a resampling procedure that provides a point estimate of generalization error for any predictive model. Cross-validation is widely used for model selection and evaluation. Uncertainty in the cross-validation estimate is challenging to quantify, and estimation of its variance is known to require multiple runs of the...

💬 0 commentsarXiv:2609.04126v1PDF
0

Posted in stat.ME · 2026-09-03 · María Eugenia Riaño

Model-assisted estimation with a training subsample: a two-phase sampling approach with design-based variance estimation

When a flexible prediction model is fitted on a training subsample drawn from a probability sample, the model-assisted estimator actually reported arises from one realized partition, yet existing theory quantifies uncertainty only for partition-averaged, cross-fitted, or symmetrized versions of it. We represent the training subsample...

💬 0 commentsarXiv:2609.04082v1PDF
0

Posted in stat.ME · 2026-09-03 · Emanuele Giorgi, Claudio Fronterre, Peter Diggle

Comment on: "The Two Cultures of Prevalence Mapping: Small Area Estimation and Model-Based Geostatistics"

Small Area Estimation (SAE) and Model-Based Geostatistics (MBG) provide complementary approaches to prevalence mapping, with their relative advantages depending on the inferential goals and characteristics of the available data. We argue that a fuller comparison should consider model interpretability, the role of epidemiologically...

💬 0 commentsarXiv:2609.03805v1PDF
0

Posted in stat.ME · 2026-09-03 · Benjamin Poignard, Yoann Potiron

Parametric estimation of Hawkes processes based on ordinary least squares

We develop a parametric estimation framework for self-exciting Hawkes processes whose intensity functions admit a parametric form. The estimation procedure is based on ordinary least squares. To apply the least squares estimation, we restrict to a kernel class that can be expressed as a sum of the product of a parameter and a...

💬 0 commentsarXiv:2609.03696v1PDF
0

Posted in stat.ME · 2026-09-03 · Žikica Lukić, Bojana Milošević

Change-point analysis: a new perspective for unstable financial markets

We introduce two new classes of nonparametric change-point tests for sequences of univariate non-negative random variables. The proposed procedures are based on the empirical modified Hankel transform and the Laplace transform, respectively, and provide new transform-based tools for detecting distributional changes. We derive the...

💬 0 commentsarXiv:2609.03614v1PDF
0

Posted in stat.ME · 2026-09-03 · Giulia Patanè, Sonja Greven, Alessandra Menafoglio

Random mixtures in Bayes Hilbert spaces

We present a framework for the analysis and unmixing of random density mixtures in the Bayes Hilbert space. General identifiability results for mixtures in Hilbert spaces are established and applied to the Bayes Hilbert space setting. Building on these results, we propose a penalised maximum likelihood approach for the unmixing of...

💬 0 commentsarXiv:2609.03523v1PDF
0

Posted in stat.ML · 2026-09-03 · Siyuan He, Bokai Yang, Jie Hu, Ziwen Gao, Yuhong Yang

Towards a Statistical Understanding of Mixture-of-Experts

Mixture-of-experts (MoE) architectures increase model capacity by combining a collection of expert predictors through input-dependent routing, while often activating only a small subset of experts for each input. Despite their growing importance in modern large-scale models, the statistical roles of their design choices, especially...

💬 0 commentsarXiv:2609.03501v1PDF
0

Posted in stat.ML · 2026-09-03 · Quang Hoang Trung, Quang Huu Hieu, Nguyen Van Hoang Phuc, Vo Nguyen Le Duy

ALRA: Adaptive Local Relational Alignment for Logit-Based Pre-training Distillation of Autoregressive Language Models

Logit-based knowledge distillation for autoregressive language models usually aligns teacher and student next-token distributions over the entire vocabulary. However, this global objective overlooks relative preferences among likely token alternatives. Existing local approaches often select candidate tokens from either the teacher or...

💬 0 commentsarXiv:2609.03355v1PDF
0

Posted in stat.ME · 2026-09-03 · Bob Wilson

Randomization Inference for Matched Pairs with Binary Outcomes

We give an exact randomization-based confidence set for the average treatment effect (ATE) in matched-pair studies with a binary outcome, requiring neither monotonicity nor any distributional assumption beyond the within-pair coin flip. At its core is an analytic solution to the worst-case allocation of attributable effects: two...

💬 0 commentsarXiv:2609.03227v1PDF
0

Posted in stat.AP · 2026-09-02 · Yulin Guo, Veera Sundararaghavan, Boris Kramer

Uncertainty quantification of fatigue initiation life for powder bed fusion metal additive manufacturing

Predicting fatigue life with quantified uncertainties is essential for the qualification of critical components produced by laser-based powder bed fusion additive manufacturing. We present a framework that propagates microstructure and defect uncertainties directly to a fatigue initiation life distribution for a specific part. In...

💬 0 commentsarXiv:2609.03163v1PDF