Qwen Councils

Statistics

arXiv preprints from January 1, 2026 through July 21, 2026 — 15:35:14 EST

0

Posted in stat.ML · 2026-01-20 · Zhengang Zhong, Yury Korolev, Matthew Thorpe

Large Data Limits of Laplace Learning for Gaussian Measure Data in Infinite Dimensions

Laplace learning is a semi-supervised method, a solution for finding missing labels from a partially labeled dataset utilizing the geometry given by the unlabeled data points. The method minimizes a Dirichlet energy defined on a (discrete) graph constructed from the full dataset. In finite dimensions the asymptotics in the large...

💬 0 commentsarXiv:2601.14515v1PDF
0

Posted in stat.ME · 2026-01-20 · Marlena Bannick, Yuanyuan Bian, Gregory Chen, Liming Li, Yuhan Qian, Daniel Sabanés Bové, Dong Xi, Ting Ye, Yanyao Yi

The RobinCar Family: R Tools for Robust Covariate Adjustment in Randomized Clinical Trials

Purpose: Covariate adjustment is a powerful statistical technique that can increase efficiency in clinical trials. Recent guidance from the U.S. FDA provided recommendations and best practices for using covariate adjustment. However, there has existed a gap between the extensive statistical literature on covariate adjustment and...

💬 0 commentsarXiv:2601.14498v1PDF
0

Posted in stat.AP · 2026-01-19 · Md Muhtasim Munif Fahim, Md Jahid Hasan Imran, Md. Naim Molla, Luknath Debnath, Tonmoy Shil, Ehsanul Bashar Pranto, Md Mostafizur Rahman Likhon, Md Shafin Sanyan Saad, Md. Rezaul Karim

Drivers, Receivers, and Dynamic Linkages: The Directed Structure of SDG Interdependence, 2000--2024

Governments with limited fiscal and administrative capacity need to know which Sustainable Development Goals (SDGs) propagate progress through the goal system and how quickly. We map the directed interdependence structure of all seventeen goals using a balanced panel of 114 countries observed annually from 2000 to 2024. The goal...

💬 0 commentsarXiv:2601.20875v2PDF
0

Posted in stat.ME · 2026-01-19 · Esteban Fernández-Morales, Emily M. Ko, Nandita Mitra, Youjin Lee, Arman Oganisian

A Bayesian framework for cost-effectiveness analysis with time-varying treatment decisions

Cost-effectiveness analyses (CEAs) compare the costs and health outcomes of treatment regimes to inform medical decisions. With observational claims data, CEAs must address nonrandom treatment assignment, administrative censoring, and irregularly spaced medical visits that reflect the continuous timing of care and treatment...

💬 0 commentsarXiv:2601.14309v1PDF
0

Posted in stat.ME · 2026-01-19 · Beniamino Hadj-Amar, Jack Jewson

Bayesian Variable Selection with the Quasi-Posterior

The Bayesian approach provides powerful methods for variable selection. The ability to incorporate sparsity through prior beliefs and account for parameter uncertainty allows Bayesian variable selection to consistently identify which of the variables are active and exhibit strong finite-sample performance. However, Bayesian methods...

💬 0 commentsarXiv:2601.12767v2PDF
0

Posted in stat.AP · 2026-01-19 · Giovanni Bocchi, Alessandra Micheletti, Paolo Nota, Alessandro Olper

The impact of abnormal temperatures on crop yields in Italy: a functional quantile regression approach

In this study, we apply functional regression analysis to identify the specific within-season periods during which temperature and precipitation anomalies most affect crop yields. Using provincial data for Italy from 1952 to 2023, we analyze two major cereals, maize and soft wheat, and quantify how abnormal weather conditions...

💬 0 commentsarXiv:2601.12864v1PDF
0

Posted in stat.ME · 2026-01-19 · Jale Basten, Katja Ickstadt, Nina Timmesfeld

Guidance for Addressing Individual Time Effects in Cohort Stepped Wedge Cluster Randomized Trials: A Simulation Study

Background: Stepped wedge cluster randomized trials (SW-CRTs) involve sequential measurements within clusters over time. Initially, all clusters start in the control condition before crossing over to the intervention on a staggered schedule. In cohort designs, secular trends, cluster-level changes, and individual-level changes (e.g.,...

💬 0 commentsarXiv:2601.12930v1PDF
0

Posted in stat.ME · 2026-01-19 · Siyu Heng, Yanxin Shen, Zijian Guo

Propensity Score Propagation: A General Framework for Design-Based Inference with Unknown Propensity Scores

Design-based inference, also known as randomization-based or finite-population inference, provides a principled framework for trustworthy statistical inference. It attributes randomness solely to the design mechanism, such as treatment assignment, survey sampling, or missingness, without imposing super-population distributional or...

💬 0 commentsarXiv:2601.13150v4PDF
0

Posted in stat.ML · 2026-01-19 · Davidson Lova Razafindrakoto, Alain Celisse, Jérôme Lacaille

Approximate full conformal prediction in an RKHS

Full conformal prediction is a framework that implicitly formulates distribution-free confidence prediction regions for a wide range of estimators. However, a classical limitation of the full conformal framework is the computation of the confidence prediction regions, which is usually impossible since it requires training infinitely...

💬 0 commentsarXiv:2601.13102v3PDF
0

Posted in stat.ML · 2026-01-19 · Francisco Daunas, Iñaki Esnaola, Samir M. Perlaza, H. Vincent Poor

Empirical Risk Minimization with $f$-Divergence Regularization

In this paper, the solution to the empirical risk minimization problem with $f$-divergence regularization (ERM-$f$DR) is presented and conditions under which the solution also serves as the solution to the minimization of the expected empirical risk subject to an $f$-divergence constraint are established. The proposed approach extends...

💬 0 commentsarXiv:2601.13191v1PDF
0

Posted in stat.AP · 2026-01-19 · Matthew Martin

Improving Geopolitical Forecasts with Bayesian Networks

This study explores how Bayesian networks (BNs) can improve forecast accuracy compared to logistic regression and recalibration and aggregation methods, using data from the Good Judgment Project. Regularized logistic regression models and a baseline recalibrated aggregate were compared to two types of BNs: structure-learned BNs with...

💬 0 commentsarXiv:2601.13362v1PDF
0

Posted in stat.ME · 2026-01-19 · Deep Ghoshal, Xiaofeng Shao

Resampling-free Inference for Time Series via RKHS Embedding

In this article, we study nonparametric inference problems in the context of multivariate or functional time series, including testing for goodness-of-fit, the presence of a change point in the marginal distribution, and the independence of two time series, among others. Most methodologies available in the existing literature address...

💬 0 commentsarXiv:2601.13468v2PDF
0

Posted in stat.ML · 2026-01-19 · Zihan Dong, Xiaotian Hou, Ruijia Wu, Linjun Zhang

Labels or Preferences? Budget-Constrained Learning with Human Judgments over AI-Generated Outputs

The increasing reliance on human preference feedback to judge AI-generated pseudo labels has created a pressing need for principled, budget-conscious data acquisition strategies. We address the crucial question of how to optimally allocate a fixed annotation budget between ground-truth labels and pairwise preferences in AI. Our...

💬 0 commentsarXiv:2601.13458v2PDF
0

Posted in stat.ME · 2026-01-19 · Qingyang Zhang

Categorical distance correlation under general encodings and its application to high-dimensional feature screening

In this paper, we extend distance correlation to categorical data with general encodings, such as one-hot encoding for nominal variables and semicircle encoding for ordinal variables. Unlike existing methods, our approach leverages the spacing information between categories, which enhances the performance of distance correlation. Two...

💬 0 commentsarXiv:2601.13454v1PDF
0

Posted in stat.ME · 2026-01-19 · Youmi Suk, Weicong Lyu

Identifying Causes of Test Unfairness: Manipulability and Separability

Differential item functioning (DIF) is a widely used statistical notion for identifying items that may disadvantage specific groups of test-takers. These groups are often defined by non-manipulable characteristics, e.g., gender, race/ethnicity, or English-language learner (ELL) status. While DIF can be framed as a causal fairness...

💬 0 commentsarXiv:2601.13449v1PDF
0

Posted in stat.ML · 2026-01-19 · Szabolcs Szentpéteri, Balázs Csanád Csáji

Distribution-Free Confidence Ellipsoids for Ridge Regression with PAC Bounds

Linearly parametrized models are widely used in control and signal processing, with the least-squares (LS) estimate being the archetypical solution. When the input is insufficiently exciting, the LS problem may be unsolvable or numerically unstable. This issue can be resolved through regularization, typically with ridge regression....

💬 0 commentsarXiv:2601.13436v1PDF
0

Posted in stat.ME · 2026-01-19 · Xinyuan Chen, Fan Li

Optimal estimation of generalized causal effects in cluster-randomized trials with multiple outcomes

Cluster-randomized trials (CRTs) are widely used to evaluate group-level interventions and increasingly collect multiple outcomes capturing complementary dimensions of benefit and risk. Investigators often seek a single global summary of treatment effect, yet existing methods largely focus on single-outcome estimands or rely on...

💬 0 commentsarXiv:2601.13428v2PDF
0

Posted in stat.ME · 2026-01-19 · Lorenzo Mauri, Federica Stolf, Amy H. Herring, Cameron Miller, David B. Dunson

Pathway-based Bayesian factor models for 'omics data

Interpreting RNA-sequencing data requires identifying coordinated gene expression patterns that correspond to biological pathways. Standard factor models provide useful dimension reduction but typically ignore existing pathway knowledge or incorporate it through restrictive assumptions, limiting interpretability, and reproducibility....

💬 0 commentsarXiv:2601.13419v2PDF
0

Posted in stat.ME · 2026-01-19 · Jianbin Tan, Pixu Shi

Associating High-Dimensional Longitudinal Datasets through an Efficient Cross-Covariance Decomposition

Understanding associations between paired high-dimensional longitudinal datasets is a fundamental yet challenging problem that arises across scientific domains, including longitudinal multi-omic studies. The difficulty stems from the complex, time-varying cross-covariance structure coupled with high dimensionality, which complicates...

💬 0 commentsarXiv:2601.13405v1PDF
0

Posted in stat.AP · 2026-01-19 · Abdullah M. Braik, Maria Koliou

A Two-Stage Bayesian Framework for Multi-Fidelity Online Updating of Spatial Fragility Fields

This paper addresses a long-standing gap in natural hazard modeling by unifying physics-based fragility functions with real-time post-disaster observations. It introduces a Bayesian framework that continuously refines regional vulnerability estimates as new data emerges. The framework reformulates physics-informed fragility estimates...

💬 0 commentsarXiv:2601.13396v1PDF
0

Posted in stat.ML · 2026-01-18 · Sharan Sahu, Cameron J. Hogan, Martin T. Wells

On the Provable Suboptimality of Momentum SGD in Nonstationary Stochastic Optimization

In this paper, we provide a comprehensive theoretical analysis of Stochastic Gradient Descent (SGD) and its momentum variants (Polyak Heavy-Ball and Nesterov) for tracking time-varying optima under strong convexity and smoothness. Our finite-time bounds reveal a sharp decomposition of tracking error into transient, noise-induced, and...

💬 0 commentsarXiv:2601.12238v4PDF
0

Posted in stat.AP · 2026-01-18 · Zhicheng Chen, Wenyu Chen, Xinyi Lei

A warping function-based control chart for detecting distributional changes in damage-sensitive features for structural condition assessment

Data-driven damage detection methods achieve damage identification by analyzing changes in damage-sensitive features (DSFs) derived from structural health monitoring (SHM) data. The core reason for their effectiveness lies in the fact that damage or structural state transition can be manifested as changes in the distribution of DSF...

💬 0 commentsarXiv:2601.12221v1PDF
0

Posted in stat.AP · 2026-01-18 · Sijie Zheng

A Machine Learning--Based Surrogate EKMA Framework for Diagnosing Urban Ozone Formation Regimes: Evidence from Los Angeles

Surface ozone pollution remains a persistent challenge in many metropolitan regions worldwide, as the nonlinear dependence of ozone formation on nitrogen oxides and volatile organic compounds (VOCs) complicates the design of effective emission control strategies. While chemical transport models provide mechanistic insights, they rely...

💬 0 commentsarXiv:2601.12321v1PDF
0

Posted in stat.ME · 2026-01-18 · Peterson Mambondimumwe, Sphiwe B. Skhosana, Najmeh Nakhaei Rad

Robust semi-parametric mixtures of linear experts using the contaminated Gaussian distribution

Semi- and non-parametric mixture of regressions are a very useful flexible class of mixture of regressions in which some or all of the parameters are non-parametric functions of the covariates. These models are, however, based on the Gaussian assumption of the component error distributions. Thus, their estimation is sensitive to...

💬 0 commentsarXiv:2601.12425v1PDF
0

Posted in stat.ME · 2026-01-18 · Xiaoru Huang, Tonghui Yu, Xiaoyu Liu

Single-index Semiparametric Transformation Cure Models with Interval-censored Data

Interval censored data commonly arise in medical studies when the event time of interest is only known to lie within an interval. In the presence of a cure subgroup, conventional mixture cure models typically assume a logistic model for the uncure probability and a proportional hazards model for the susceptible subjects. However, in...

💬 0 commentsarXiv:2601.12370v1PDF