Qwen Councils

Statistics

arXiv preprints from January 1, 2026 through July 20, 2026 — 19:23:35 EST

0

Posted in stat.ME · 2026-01-16 · Ricardo J. Sandoval, Sivaraman Balakrishnan, Avi Feller, Michael I. Jordan, Ian Waudby-Smith

On Nonasymptotic Confidence Intervals for Treatment Effects in Randomized Experiments

We study nonasymptotic (finite-sample) confidence intervals for treatment effects in randomized experiments. In the existing literature, the effective sample sizes of nonasymptotic confidence intervals tend to be looser than the corresponding central-limit-theorem-based confidence intervals by a factor depending on the square root of...

💬 0 commentsarXiv:2601.11744v2PDF
0

Posted in stat.ME · 2026-01-16 · Xinlei Xu, Caitlin H Daly, Audrey Béliveau

Identifying Conditions Favouring Multiplicative Heterogeneity Models in Network Meta-Analysis

Explicit modelling of between-study heterogeneity is essential in network meta-analysis (NMA) to ensure valid inference and avoid overstating precision. While the additive random-effects (RE) model is the conventional approach, the multiplicative-effect (ME) model remains underexplored. The ME model inflates within-study variances by...

💬 0 commentsarXiv:2601.11735v2PDF
0

Posted in stat.ML · 2026-01-16 · Guerlain Lambert, Céline Helbert, Claire Lauvernet

Gradient-based Active Learning with Gaussian Processes for Global Sensitivity Analysis

Global sensitivity analysis of complex numerical simulators is often limited by the small number of model evaluations that can be afforded. In such settings, surrogate models built from a limited set of simulations can substantially reduce the computational burden, provided that the design of computer experiments is enriched...

💬 0 commentsarXiv:2601.11790v1PDF
0

Posted in stat.ME · 2026-01-15 · Yang Ou, Lan Xue, Carmen Tekwe, Kedir N. Turi, Roger S. Zoh

Estimating the effect of lymphovascular invasion on 2-year survival probability under endogeneity: a recursive copula-based approach

Lymphovascular invasion (LVI) is an important prognostic marker for head and neck squamous cell carcinoma (HNSC), but the true effect of LVI on survival may be distorted by endogeneity arising from unmeasured confounding. Conventional one-stage conditional models and instrument-based two-stage estimators are prone to bias under...

💬 0 commentsarXiv:2601.09984v1PDF
0

Posted in stat.ME · 2026-01-15 · Faruk Muritala, Austin Brown, Dhrubajyoti Ghosh, Sherry Ni

Derivations for the Cumulative Standardized Binomial EWMA (CSB-EWMA) Control Chart

This paper presents the exact mathematical derivation of the mean and variance properties for the Exponentially Weighted Moving Average (EWMA) statistic applied to binomial proportion monitoring in Multiple Stream Processes (MSPs). We develop a Cumulative Standardized Binomial EWMA (CSB-EWMA) formulation that provides adaptive control...

💬 0 commentsarXiv:2601.09968v1PDF
0

Posted in stat.ME · 2026-01-15 · Lei Huang, Chengyue Liu, Li Wang

Weighted least squares estimation by multivariate-dependent weights for linear regression models

Multivariate linear regression models often face the problem of heteroscedasticity caused by multiple explanatory variables. The weighted least squares estimation with univariate-dependent weights has limitations in constructing weight functions. Therefore, this paper proposes a multivariate dependent weighted least squares estimation...

💬 0 commentsarXiv:2601.10049v1PDF
0

Posted in stat.AP · 2026-01-15 · Peter Maurice Catt

An Information-Theoretic Diagnostic Analytics Framework for Mapping Past-Future Dependence in Horizon-Specific Forecastability

In many systems, the true data-generating process is unknown, requiring forecasters to rely on observed time series. This study proposes a pre-modeling diagnostic framework for horizon-specific forecastability assessment that evaluates forecastability before model selection begins. Forecastability is operationalized using auto-mutual...

💬 0 commentsarXiv:2601.10006v4PDF
0

Posted in stat.ME · 2026-01-15 · Mayukh Choudhury, Debraj Das, Sujit Ghosh

Asymptotic Theory of Tail Dependence and Bootstrap for Checkerboard Copulas

A comprehensive asymptotic and bootstrap theory is established for checkerboard-based estimation of the copula and its lower and upper tail copula counterparts under unknown marginal distributions. The proposed estimator of the tail copula extends a local bilinear interpolation of the empirical copula to the tail region, providing a...

💬 0 commentsarXiv:2601.10252v3PDF
0

Posted in stat.ME · 2026-01-15 · Yue Yu, Guanghui Wang, Liu Liu, Changliang Zou

Model-Agnostic and Uncertainty-Aware Dimensionality Reduction in Supervised Learning

Dimension reduction is a fundamental tool for analyzing high-dimensional data in supervised learning. Traditional methods for estimating intrinsic order often prioritize model-specific structural assumptions over predictive utility. This paper introduces predictive order determination (POD), a model-agnostic framework that determines...

💬 0 commentsarXiv:2601.10357v1PDF
0

Posted in stat.AP · 2026-01-15 · Rehinatu Usman, Onyedikachi J. Okeke

Climate Vulnerability and Community Health: Identifying Greensboro Neighborhoods at Intersectional Risk

This study develops an integrated, intersectional climate vulnerability assessment for Greensboro, North Carolina, a midsize city in the rapidly changing American Southeast. Moving beyond generalized mapping, we combine demographic, socioeconomic, health, and environmental data at the census tract level to identify neighborhoods where...

💬 0 commentsarXiv:2601.15675v1PDF
0

Posted in stat.AP · 2026-01-15 · Mikkel Meyer Andersen, Nicole Huber, Kimberly S Andreaggi, Tóra Oluffa Stenberg Olsen, Walther Parson, Charla Marshall

MitoFREQ: A Novel Approach for Mitogenome Frequency Estimation from Top-level Haplogroups and Single Nucleotide Variants

Lineage marker population frequencies can serve as one way to express evidential value in forensic genetics. However, for high-quality whole mitochondrial DNA genome sequences (mitogenomes), population data remain limited. In this paper, we offer a new method, MitoFREQ, for estimating the population frequencies of mitogenomes....

💬 0 commentsarXiv:2601.10464v1PDF
0

Posted in stat.AP · 2026-01-15 · Glenna Nightingale, Karthik Mohan, Eloi Ribe, Valentin Popov, Shakes Wang, Clara Calia, Luciana Brondi, Sohan Seth

Modeling mental health trajectories during the COVID-19 pandemic using UK-wide data in the presence of sociodemographic variables

Background: The negative effects of the COVID-19 pandemic on the mental health and well-being of populations are an important public health issue. Our study aims to determine the underlying factors shaping mental health trajectories during the COVID-19 pandemic in the UK. Methods: Data from the Understanding Society COVID-19 Study...

💬 0 commentsarXiv:2601.10445v1PDF
0

Posted in stat.ME · 2026-01-15 · Yingying Ma, Chenlei Leng

A Propagation Framework for Network Regression

We introduce a unified and computationally efficient framework for regression on network data, addressing limitations of existing models that require specialized estimation procedures or impose restrictive decay assumptions. Our Network Propagation Regression (NPR) models outcomes as functions of covariates propagated through network...

💬 0 commentsarXiv:2601.10533v1PDF
0

Posted in stat.ML · 2026-01-15 · Francisco Madaleno, Pratik Misra, Alex Markham

Coarsening Causal DAG Models

Directed acyclic graphical (DAG) models are a powerful tool for representing causal relationships among jointly distributed random variables, especially concerning data from across different experimental settings. However, it is not always practical or desirable to estimate a causal model at the granularity of given features in a...

💬 0 commentsarXiv:2601.10531v2PDF
0

Posted in stat.ML · 2026-01-15 · Luke W. Yerbury, Ricardo J. G. B. Campello, G. C. Livingston, Mark Goldsworthy, Lachlan O'Neil

CROCS: A Two-Stage Clustering Framework for Behaviour-Centric Consumer Segmentation with Smart Meter Data

With grid operators confronting rising uncertainty from renewable integration and a broader push toward electrification, Demand-Side Management (DSM) -- particularly Demand Response (DR) -- has attracted significant attention as a cost-effective mechanism for balancing modern electricity systems. Unprecedented volumes of consumption...

💬 0 commentsarXiv:2601.10494v3PDF
0

Posted in stat.CO · 2026-01-15 · Constantin Vaillant Tenzer

Mesh Denoising

In this paper, we study four mesh denoising methods: linear filtering, a heat diffusion method, Sobolev regularization, and, to a lesser extent, a barycentric approach based on the Sinkhorn algorithm. We illustrate that, for a simple image denoising task, a naive choice of a Gibbs kernel can lead to unsatisfactory results. We...

💬 0 commentsarXiv:2601.10487v1PDF
0

Posted in stat.ME · 2026-01-15 · William L. Lippitt, Edward J. Bedrick, Nichole E. Carlson

Adjusted Similarity Measures and a Violation of Expectations

Adjusted similarity measures, such as Cohen's kappa for inter-rater reliability and the adjusted Rand index used to compare clustering algorithms, are a vital tool for comparing discrete labellings. These measures are intended to have the property of 0 expectation under a null distribution and maximum value 1 under maximal similarity...

💬 0 commentsarXiv:2601.10641v1PDF
0

Posted in stat.ML · 2026-01-15 · Eric Xia, Jason M. Klusowski

Classification Imbalance as Transfer Learning

Classification imbalance arises when one class is much rarer than the other. We frame this setting as transfer learning under label (prior) shift between an imbalanced source distribution induced by the observed data and a balanced target distribution under which performance is evaluated. Within this framework, we study a family of...

💬 0 commentsarXiv:2601.10630v1PDF
0

Posted in stat.ML · 2026-01-15 · Mihailo Stojnic

Parametric RDT approach to computational gap of symmetric binary perceptron

We study potential presence of statistical-computational gaps (SCG) in symmetric binary perceptrons (SBP) via a parametric utilization of \emph{fully lifted random duality theory} (fl-RDT) [96]. A structural change from decreasingly to arbitrarily ordered $c$-sequence (a key fl-RDT parametric component) is observed on the second...

💬 0 commentsarXiv:2601.10628v1PDF
0

Posted in stat.ME · 2026-01-15 · Yongzhen Feng, Weiwei Wang, Raymond K. W. Wong, Xianyang Zhang

Fair Regression under Demographic Parity: A Unified Framework

We propose a unified framework for fair regression tasks formulated as risk minimization problems subject to a demographic parity constraint. Unlike many existing approaches that are limited to specific loss functions or rely on challenging non-convex optimization, our framework is applicable to a broad spectrum of regression tasks....

💬 0 commentsarXiv:2601.10623v1PDF
0

Posted in stat.ME · 2026-01-15 · Paramahansa Pramanik, Arnab Kumar Maity, Anjan Mandal, Haley Kate Robinson

A Bayesian Discrete Framework for Enhancing Decision-Making Processes in Clinical Trial Designs and Evaluations

This study examines the application of Bayesian approach in the context of clinical trials, emphasizing their increasing importance in contemporary biomedical research. While conventional frequentist approach provides a foundational basis for analysis, it often lacks the flexibility to integrate prior knowledge, which can constrain...

💬 0 commentsarXiv:2601.10615v1PDF
0

Posted in stat.ME · 2026-01-15 · Zhangyi He, Feng Yu, Suzie Cro, Laurent Billot

From aggressive to conservative early stopping in Bayesian group sequential designs

Group sequential designs (GSDs) are widely used in confirmatory trials to allow interim monitoring while preserving control of the type I error rate. In the frequentist framework, O'Brien-Fleming-type stopping boundaries dominate practice because they impose highly conservative early stopping while allowing more liberal decisions as...

💬 0 commentsarXiv:2601.10590v1PDF
0

Posted in stat.ME · 2026-01-15 · Salvador V. Balkus, Hasan Laith, Nima S. Hejazi

On the use of cross-fitting in causal machine learning with correlated units

In causal machine learning, the fitting and evaluation of nuisance models are often performed on separate partitions, or folds, of the observed data. This technique, called cross-fitting, eliminates bias introduced by the use of black-box predictive algorithms. When study units may be correlated, such as in spatial, clustered, or...

💬 0 commentsarXiv:2601.10899v2PDF
0

Posted in stat.ME · 2026-01-15 · Simon Fontaine, Nisha J. D'Silva, Marcell Costa de Medeiros, Grace Y. Chen, Ji Zhu, Gen Li

Locally sparse varying coefficient mixed model with application to longitudinal microbiome differential abundance

Differential abundance (DA) analysis in microbiome studies has recently been used to uncover a plethora of associations between microbial composition and various health conditions. While current approaches to DA typically apply only to cross-sectional data, many studies feature a longitudinal design to better understand the underlying...

💬 0 commentsarXiv:2601.10872v1PDF
0

Posted in stat.ML · 2026-01-14 · Nick Polson, Vadim Sokolov

Horseshoe Mixtures-of-Experts (HS-MoE)

Horseshoe mixtures-of-experts (HS-MoE) models provide a Bayesian framework for sparse expert selection in mixture-of-experts architectures. We combine the horseshoe prior's adaptive global-local shrinkage with input-dependent gating, yielding data-adaptive sparsity in expert usage. Our primary methodological contribution is a particle...

💬 0 commentsarXiv:2601.09043v1PDF