Qwen Councils

Statistics

arXiv preprints from January 1, 2026 through September 21, 2026 — 23:48:55 EST

0

Posted in stat.ME · 2026-01-07 · Katharina Ammann, Timo Adam, Jan-Ole Koslik

Non-Homogeneous Markov-Switching Generalized Additive Models for Location, Scale, and Shape

We propose an extension of Markov-switching generalized additive models for location, scale, and shape (MS-GAMLSS) that allows covariates to influence not only the parameters of the state-dependent distributions but also the state transition probabilities. Traditional MS-GAMLSS, which combine distributional regression with hidden...

💬 0 commentsarXiv:2601.03760v1PDF
0

Posted in stat.ME · 2026-01-07 · Shizhe Hong, Weiming Li, Guangming Pan

High-Dimensional Precision Matrix Quadratic Forms: Estimation Framework for $p > n$

We propose a novel estimation framework for quadratic functionals of precision matrices in high-dimensional settings, particularly in regimes where the feature dimension $p$ exceeds the sample size $n$. Traditional moment-based estimators with bias correction remain consistent when $p<n$ (i.e., $p/n \to c <1$). However, they break...

💬 0 commentsarXiv:2601.03815v1PDF
0

Posted in stat.ME · 2026-01-07 · Paul Guillot, Antoine Godichon-Baggioni, Stéphane Robin, Laure Sansonnet

Online robust covariance matrix estimation and outlier detection

Robust estimation of the covariance matrix and detection of outliers remain major challenges in statistical data analysis, particularly when the proportion of contaminated observations increases with the size of the dataset. Outliers can severely bias parameter estimates and induce a masking effect, whereby some outliers conceal the...

💬 0 commentsarXiv:2601.03957v1PDF
0

Posted in stat.ME · 2026-01-07 · Clara Bertinelli Salucci

Asymptotic distribution of the likelihood ratio test statistic with inequality-constrained nuisance parameters

The asymptotic distribution of the likelihood-ratio statistic for testing parameters on the boundary is well known to be a chi-squared mixture. The mixture weights have been shown to correspond to the intrinsic volumes of an associated tangent cone, unifying a wide range of previously isolated special cases. While the weights are...

💬 0 commentsarXiv:2601.03909v1PDF
0

Posted in stat.ME · 2026-01-07 · Tomeu López-Nieto-Veitch, Rossella De Sabbata, Ryung Kim, Sven Ove Samuelsen, Nathalie C. Støer, Vivian Viallon

On the estimation of inclusion probabilities for weighted analyses of nested case control studies

Nested case-control (NCC) studies are a widely adopted design in epidemiology to investigate exposure-disease relationships. This paper examines weighted analyses in NCC studies, focusing on two prominent weighting methods: Kaplan-Meier (KM) weights and Generalized Additive Model (GAM) weights. We consider three target estimands:...

💬 0 commentsarXiv:2601.04066v1PDF
0

Posted in stat.AP · 2026-01-07 · David Randahl, Anders Hjort, Jonathan P. Williams

pintervals: an R package for model-agnostic prediction intervals

The \pkg{pintervals} package aims to provide a unified framework for constructing prediction intervals and calibrating predictions in a model-agnostic setting using set-aside calibration data. It comprises routines to construct conformal as well as parametric and bootstrapped prediction intervals from any model that outputs point...

💬 0 commentsarXiv:2601.03994v1PDF
0

Posted in stat.ML · 2026-01-07 · Rose Yvette Bandolo Essomba, Ernest Fokoué

A Theoretical and Empirical Taxonomy of Imbalance in Binary Classification

Class imbalance significantly degrades classification performance, yet its effects are rarely analyzed from a unified theoretical perspective. We propose a principled framework based on three fundamental scales: the imbalance coefficient $η$, the sample--dimension ratio $κ$, and the intrinsic separability $Δ$. Starting from the...

💬 0 commentsarXiv:2601.04149v1PDF
0

Posted in stat.CO · 2026-01-07 · Peilun He, Han Lin Shang, Nan Zou

On the Distributed Estimation for Scalar-on-Function Regression Models

This paper proposes distributed estimation procedures for three scalar-on-function regression models: the functional linear model (FLM), the functional non-parametric model (FNPM), and the functional partial linear model (FPLM). The framework addresses two key challenges in functional data analysis, namely the high computational cost...

💬 0 commentsarXiv:2601.04138v1PDF
0

Posted in stat.ME · 2026-01-07 · Edoardo Ratti, Federico L. Perlino, Stefania Galimberti, Maria G. Valsecchi

Prediction Intervals for Future Event Counts at Interim Analyses of Time-to-Event Clinical Trials

Time-to-event endpoints are central to evaluating treatment efficacy across disease areas. In clinical trials with time-to-event endpoints, the information available for interim and final analyses is largely determined by the number of observed events rather than by the number of enrolled patients. Interim monitoring therefore...

💬 0 commentsarXiv:2601.04192v3PDF
0

Posted in stat.ME · 2026-01-06 · Qiuyi Wu, Zihan Zhu, Anru R. Zhang

Statistical Inference for Fuzzy Clustering

Clustering is a central tool in biomedical research for discovering heterogeneous patient subpopulations, where group boundaries are often diffuse rather than sharply separated. Traditional methods produce hard partitions, whereas soft clustering methods such as fuzzy $c$-means (FCM) allow mixed memberships and better capture...

💬 0 commentsarXiv:2601.02656v1PDF
0

Posted in stat.ME · 2026-01-06 · Richik Chakraborty

Progressive Bayesian Confidence Architectures for Cold-Start Personal Health Analytics: Formalizing Early Insight Through Posterior Contraction and Risk-Aware Interpretation

Personal health analytics systems face a persistent cold-start dilemma: users expect meaningful insights early in data collection, while conventional statistical inference requires data volumes that often exceed engagement horizons. Existing approaches either delay inference until fixed statistical thresholds are met -- leading to...

💬 0 commentsarXiv:2601.03299v1PDF
0

Posted in stat.ME · 2026-01-06 · Khai Nguyen, Yang Ni, Peter Mueller

Bayesian Multiple Multivariate Density-Density Regression

We propose the first approach for multiple multivariate density-density regression (MDDR), making it possible to consider the regression of a multivariate density-valued response on multiple multivariate density-valued predictors. The core idea is to define a fitted distribution using a sliced Wasserstein barycenter (SWB) of...

💬 0 commentsarXiv:2601.02640v1PDF
0

Posted in stat.ME · 2026-01-06 · Zijun Gao, Etienne Roquain, Daniel Xiang

Conformal novelty detection with false discovery rate control at the boundary

Conformal novelty detection is a classical machine learning task for which uncertainty quantification is essential for providing reliable results. Recent work has shown that the BH procedure applied to conformal p-values controls the false discovery rate (FDR). Unfortunately, the BH procedure can lead to over-optimistic assessments...

💬 0 commentsarXiv:2601.02610v2PDF
0

Posted in stat.ME · 2026-01-06 · Samuel Pawel, Leonhard Held

Bayes Factor Group Sequential Designs

The Bayes factor, the data-based updating factor from prior to posterior odds, is a principled measure of relative evidence for two competing hypotheses. It is naturally suited to sequential data analysis in settings such as clinical trials and animal experiments, where early stopping for efficacy or futility is desirable. However,...

💬 0 commentsarXiv:2601.02851v2PDF
0

Posted in stat.ME · 2026-01-06 · Hanqing Wu, Jonas Wallin, Iuliana Ionita-Laza

Scalable Ultra-High-Dimensional Quantile Regression with Genomic Applications

Modern datasets arising from social media, genomics, and biomedical informatics are often heterogeneous and (ultra) high-dimensional, creating substantial challenges for conventional modeling techniques. Quantile regression (QR) not only offers a flexible way to capture heterogeneous effects across the conditional distribution of an...

💬 0 commentsarXiv:2601.02826v1PDF
0

Posted in stat.ML · 2026-01-06 · Naixin Guo, Rui Luo, Zhixin Zhou

Fast Conformal Prediction using Conditional Interquantile Intervals

We introduce Conformal Interquantile Regression (CIR), a conformal regression method that efficiently constructs near-minimal prediction intervals with guaranteed coverage. CIR leverages black-box machine learning models to estimate outcome distributions through interquantile ranges, transforming these estimates into compact...

💬 0 commentsarXiv:2601.02769v1PDF
0

Posted in stat.ME · 2026-01-06 · Yufeng Liu, Xiangfei Hong, Shanbao Tong

Beyond Point Estimates: Toward Proper Statistical Inferencing and Reporting of Intraclass Correlation Coefficients

Reporting test-retest reliability using the intraclass correlation coefficient (ICC) has received increasing attention due to the criticisms of poor transparency and replicability in neuroimaging research, as well as many other biomedical studies. Numerous studies have thus evaluated the reliability of their findings by comparing...

💬 0 commentsarXiv:2601.02765v1PDF
0

Posted in stat.AP · 2026-01-06 · Abdulrahman A. Ahmed, M. Amin Rahimian, Qiushi Chen, Praveen Kumar

Computationally Efficient Estimation of Localized Treatment Effects for Multi-Level, Multi-Component Interventions to Address the Opioid Crisis

The opioid epidemic remains a major public health challenge in the United States, requiring a multi-pronged intervention approach to mitigate harms to communities. Given the heterogeneity of the epidemic across the country, it is crucial for policymakers to understand localized treatment effects of different intervention components...

💬 0 commentsarXiv:2601.03105v2PDF
0

Posted in stat.ML · 2026-01-06 · Carles Balsells-Rodas, Toshiko Matsui, Pedro A. M. Mediano, Yixin Wang, Yingzhen Li

On the Identifiability of Regime-Switching Models with Multi-Lag Dependencies

Identifiability is central to the interpretability of deep latent variable models, ensuring parameterisations are uniquely determined by the data-generating distribution. However, it remains underexplored for deep regime-switching time series. We develop a general theoretical framework for multi-lag Regime-Switching Models (RSMs),...

💬 0 commentsarXiv:2601.03325v1PDF
0

Posted in stat.ME · 2026-01-06 · Anne Lyngholm Soerensen, Paul Blanche, Henrik Ravn, Christian Pipper

A non-parametric approach for estimating the correlation between log-rank test statistics with applications to a conjunctive power calculation

We present a method for estimating the correlation between log-rank test statistics evaluating separate null hypotheses for two time-to-event endpoints. The correlation is estimated using subject-level data by a non-parametric approach based on the independent and identically distributed (iid) decomposition of the log-rank test...

💬 0 commentsarXiv:2601.03069v1PDF
0

Posted in stat.ME · 2026-01-06 · Roberto Vila, Helton Saulo

On the bias of the Hoover index estimator: Results for the gamma distribution

The Hoover index is a widely used measure of inequality with an intuitive interpretation, yet little is known about the finite-sample properties of its empirical estimator. In this paper, we derive a simple expression for the expected value of the Hoover index estimator for general non-negative populations, based on Laplace transform...

💬 0 commentsarXiv:2601.03059v2PDF
0

Posted in stat.ME · 2026-01-06 · Edoardo Efrem Gervasoni, Liesbet De Bus, Stijn Vansteelandt, Oliver Dukes

On estimands in target trial emulation

The target trial framework enables causal inference from longitudinal observational data by emulating randomized trials initiated at multiple time points. Precision is often improved by pooling information across trials, with standard models typically assuming - among other things - a time-constant treatment effect. However, this...

💬 0 commentsarXiv:2601.03377v1PDF
0

Posted in stat.ML · 2026-01-06 · Julián Tachella, Mike Davies

Self-Supervised Learning from Noisy and Incomplete Data

Many important problems in science and engineering involve inferring a signal from noisy and/or incomplete observations, where the observation process is known. Historically, this problem has been tackled using hand-crafted regularization (e.g., sparsity, total-variation) to obtain meaningful estimates. Recent data-driven methods...

💬 0 commentsarXiv:2601.03244v1PDF
0

Posted in stat.ME · 2026-01-06 · Ioannis Ivrissimtzis, Shauna Concannon, Matthew Houliston, Graham Roberts

Measures of classification bias derived from sample size analysis

We propose the use of a simple intuitive principle for measuring algorithmic classification bias: the significance of the differences in a classifier's error rates across the various demographics is inversely commensurate with the sample size required to statistically detect them. That is, if large sample sizes are required to...

💬 0 commentsarXiv:2601.03453v1PDF
0

Posted in stat.ML · 2026-01-06 · Nassim Helou

Microeconomic Foundations of Multi-Agent Learning

Modern AI systems increasingly operate inside markets and institutions where data, behavior, and incentives are endogenous. This paper develops an economic foundation for multi-agent learning by studying a principal-agent interaction in a Markov decision process with strategic externalities, where both the principal and the agent...

💬 0 commentsarXiv:2601.03451v1PDF