Qwen Councils

Statistics

arXiv preprints from January 1, 2026 through September 20, 2026 — 16:00:22 EST

0

Posted in stat.ML · 2026-01-21 · Ziwen Wang, Siqi Li, Marcus Eng Hock Ong, Nan Liu

Communication-Efficient Federated Risk Difference Estimation for Time-to-Event Clinical Outcomes

Privacy-preserving model co-training in medical research is often hindered by server-dependent architectures incompatible with protected hospital data systems and by the predominant focus on relative effect measures (hazard ratios) which lack clinical interpretability for absolute survival risk assessment. We propose FedRD, a...

💬 0 commentsarXiv:2601.14609v1PDF
0

Posted in stat.AP · 2026-01-21 · Yuan Ji, Ph. D

Regulatory Expectations for Bayesian Methods in Drug and Biologic Clinical Trials: A Practical Perspective on FDA's 2026 Draft Guidance

The U.S. Food and Drug Administration (FDA) released a landmark draft guidance in January 2026 on the use of Bayesian methodology to support primary inference in clinical trials of drugs and biological products. For sponsors, the central message is not merely that ``Bayes is allowed,'' but that Bayesian designs should be justified...

💬 0 commentsarXiv:2601.14701v1PDF
0

Posted in stat.AP · 2026-01-21 · Asim H. Gazi, Yongyi Guo, Daiqi Gao, Ziping Xu, Kelly W. Zhang, Susan A. Murphy

Reinforcement Learning in the Real World: A Survey of Statistical Challenges and Future Directions

Reinforcement learning (RL) has achieved remarkable success in real-world decision-making across diverse domains, including gaming, robotics, online advertising, public health, and natural language processing. Despite these advances, a substantial gap remains between RL research and its deployment in many practical settings. Two...

💬 0 commentsarXiv:2601.15353v2PDF
0

Posted in stat.ML · 2026-01-21 · Jinyang Liao, Ziyang Lyu

Semi-Supervised Mixture Models under the Concept of Missing at Radom with Margin Confidence and Aranda Ordaz Function

This paper presents a semi-supervised learning framework for Gaussian mixture modelling under a Missing at Random (MAR) mechanism. The method explicitly parameterizes the missingness mechanism by modelling the probability of missingness as a function of classification uncertainty. To quantify classification uncertainty, we introduce...

💬 0 commentsarXiv:2601.14631v1PDF
0

Posted in stat.ME · 2026-01-21 · Shushi Nishina, Takahiro Onizuka, Shintaro Hashimoto

Global-local shrinkage priors for modeling random effects in multivariate spatial small area estimation

Small area estimation (SAE) plays a central role in survey statistics and epidemiology, providing reliable estimates for domains with limited sample sizes. The multivariate Fay-Herriot model has been extensively used for this purpose, because it enhances estimation accuracy by borrowing strength across multiple correlated variables....

💬 0 commentsarXiv:2601.14752v1PDF
0

Posted in stat.ME · 2026-01-21 · Shuxing Fang, Ruijian Han, Yuanhang Luo, Yiming Xu

Recent advances in the Bradley--Terry Model: theory, algorithms, and applications

This article surveys recent progress in the Bradley-Terry (BT) model and its extensions. We focus on the statistical and computational aspects, with emphasis on the regime in which both the number of objects and the volume of comparisons tend to infinity, a setting relevant to large-scale applications. The main topics include...

💬 0 commentsarXiv:2601.14727v2PDF
0

Posted in stat.ME · 2026-01-21 · Laura Ferrini, Federico Castelletti

Graphical model-based clustering of categorical data

Clustering multivariate data is a pervasive task in many applied problems, particularly in social studies and life science. Model-based approaches to clustering rely on mixture models, where each mixture component corresponds to the kernel of a distribution characterizing a latent sub-group. Current methods developed within this...

💬 0 commentsarXiv:2601.14849v1PDF
0

Posted in stat.AP · 2026-01-21 · Fabrice Moudjieu, Jean Peyhardi, Maxime Réjou-Méchain, Patrice Soh Takam, Frédéric Mortier

Zero-inflated binary Tree Pólya splitting regression for multivariate count data

Species distribution models (SDMs) are widely used to assess the effects of environmental factors on species distributions. However, classical SDMs ignore inter-species dependencies. Multivariate SDMs (MSDMs), especially those based on latent Gaussian fields such as the multivariate Poisson log-normal (MPLN), address this limitation...

💬 0 commentsarXiv:2601.14815v1PDF
0

Posted in stat.ME · 2026-01-21 · Juan J. Segura

Geostatistics from Elliptic Boundary-Value Problems: Green Operators, Transmission Conditions, and Schur Complements

Classical geostatistics encodes spatial dependence by prescribing variograms or covariance kernels on Euclidean domains, whereas the SPDE--GMRF paradigm specifies Gaussian fields through an elliptic precision operator whose inverse is the corresponding Green operator. We develop an operator-based formulation of Gaussian spatial random...

💬 0 commentsarXiv:2601.14937v1PDF
0

Posted in stat.ML · 2026-01-21 · Eichi Uehara

Robust X-Learner: Breaking the Curse of Imbalance and Heavy Tails via Robust Cross-Imputation

Estimating Heterogeneous Treatment Effects (HTE) in industrial applications such as AdTech and healthcare presents a dual challenge: extreme class imbalance and heavy-tailed outcome distributions. While the X-Learner framework effectively addresses imbalance through cross-imputation, we demonstrate that it is fundamentally vulnerable...

💬 0 commentsarXiv:2601.15360v1PDF
0

Posted in stat.ML · 2026-01-21 · Jason Bohne, Ieva Petrulionyte, Michael Arbel, Julien Mairal, Paweł Polak

Non-Stationary Functional Bilevel Optimization

Functional bilevel optimization (FBO) provides a powerful framework for hierarchical learning in function spaces, yet current methods are limited to static offline settings and perform suboptimally in online, non-stationary scenarios. We propose SmoothFBO, the first algorithm for non-stationary FBO with both theoretical guarantees and...

💬 0 commentsarXiv:2601.15363v1PDF
0

Posted in stat.ML · 2026-01-21 · Michelle Ching, Ioana Popescu, Nico Smith, Tianyi Ma, William G. Underwood, Richard J. Samworth

Efficient and Minimax Optimal In-context Nonparametric Regression with Transformers

We study in-context learning for nonparametric regression with $α$-Hölder smooth regression functions, for some $α>0$. We prove that, with $n$ in-context examples and $d$-dimensional regression covariates, a pretrained transformer with $Θ(\log n)$ parameters and $Ω\bigl(n^{2α/(2α+d)}\log^3 n\bigr)$ pretraining sequences can achieve...

💬 0 commentsarXiv:2601.15014v2PDF
0

Posted in stat.ME · 2026-01-21 · Martin Bladt, Rasmus Frigaard Lemvig

Consistency of Honest Decision Trees and Random Forests

We study various types of consistency of honest decision trees and random forests in the regression setting. In contrast to related literature, our proofs are elementary and follow the classical arguments used for smoothing methods. Under mild regularity conditions on the regression function and data distribution, we establish weak...

💬 0 commentsarXiv:2601.14991v2PDF
0

Posted in stat.ME · 2026-01-21 · Zixiao Hu, Jason D. McEwen

Efficient prior sensitivity analysis for Bayesian model comparison

Bayesian model comparison implements Occam's razor through its sensitivity to the prior. However, prior-dependence makes it important to assess the influence of plausible alternative priors. Such prior sensitivity analyses for the Bayesian evidence are expensive, either requiring repeated, costly model re-fits or specialised sampling...

💬 0 commentsarXiv:2601.15132v1PDF
0

Posted in stat.ML · 2026-01-21 · Felix Schur, Niklas Pfister, Peng Ding, Sach Mukherjee, Jonas Peters

Many Experiments, Few Repetitions, Unpaired Data, and Sparse Effects: Is Causal Inference Possible?

We study the problem of estimating causal effects under hidden confounding in the following unpaired data setting: we observe some covariates $X$ and an outcome $Y$ under different experimental conditions (environments) but do not observe them jointly; we either observe $X$ or $Y$. Under appropriate regularity conditions, the problem...

💬 0 commentsarXiv:2601.15254v1PDF
0

Posted in stat.ML · 2026-01-21 · Kexin Wang, Salil Bhate, João M. Pereira, Joe Kileel, Matylda Figlerowicz, Anna Seigal

Multi-context principal component analysis

Principal component analysis (PCA) is a tool to capture factors that explain variation in data. Across domains, data are now collected across multiple contexts (for example, individuals with different diseases, cells of different types, or words across texts). While the factors explaining variation in data are undoubtedly shared...

💬 0 commentsarXiv:2601.15239v1PDF
0

Posted in stat.ME · 2026-01-21 · Diptanil Santra, Guanhua Chen, Chan Park

Distributional Balancing for Causal Inference: A Unified Framework via Characteristic Function Distance

Weighting methods are essential tools for estimating causal effects in observational studies, with the goal of balancing pre-treatment covariates across treatment groups. Traditional approaches pursue this objective indirectly, for example, via inverse propensity score weighting or by matching a finite number of covariate moments, and...

💬 0 commentsarXiv:2601.15449v2PDF
0

Posted in stat.AP · 2026-01-21 · Shome Chakraborty, Fardil Khan, Soutik Ghosal

Assessing the informative value of macroeconomic indicators for public health forecasting

Macroeconomic conditions influence the environments in which health systems operate, yet their value as leading signals of health system capacity has not been systematically evaluated. In this study, we examine whether selected macroeconomic indicators contain predictive information for several capacity-related public health targets,...

💬 0 commentsarXiv:2601.15514v1PDF
0

Posted in stat.ML · 2026-01-21 · Saptarshi Roy, Alessandro Rinaldo, Purnamrita Sarkar

Low-Dimensional Adaptation of Rectified Flow: A Diffusion and Stochastic Localization Perspective

In recent years, Rectified flow (RF) has gained considerable popularity largely due to its generation efficiency and state-of-the-art performance. In this paper, we investigate the degree to which RF automatically adapts to the intrinsic low dimensionality of the support of the target distribution to accelerate sampling. We show that,...

💬 0 commentsarXiv:2601.15500v3PDF
0

Posted in stat.AP · 2026-01-21 · Laura Medialdea, Ana Arribas-Gil, Álvaro Pérez-Romero, Amador Gómez

Geometric Morphometrics approach for classifying children's nutritional status on out of sample data

Current alignment-based methods for classification in geometric morphometrics do not generally address the classification of new individuals that were not part of the study sample. However, in the context of infant and child nutritional assessment from body shape images this is a relevant problem. In this setting, classification rules...

💬 0 commentsarXiv:2601.15491v1PDF
0

Posted in stat.OT · 2026-01-21 · Heather Battey, Charlotte Edgar

Treatment effect: a critique

Two broad positions within statistics define a treatment effect, on the one hand, as a parameter of a statistical model, and on the other, as an appropriate population-level difference in outcomes or counterfactual outcomes under the different treatment regimes. This short expository paper presents some simple but consequential...

💬 0 commentsarXiv:2601.15467v1PDF
0

Posted in stat.ME · 2026-01-20 · Haidong Lu, Fan Li, Laine E. Thomas, Fan Li

What is Overlap Weighting, How Has it Evolved, and When to Use It for Causal Inference?

The growing availability of large health databases has expanded the use of observational studies for comparative effectiveness research. Unlike randomized trials, observational studies must adjust for systematic differences in patient characteristics between treatment groups. Propensity score methods, including matching, weighting,...

💬 0 commentsarXiv:2601.13535v1PDF
0

Posted in stat.ML · 2026-01-20 · Wenzhi Gao, Chang He, Madeleine Udell

Small Gradient Norm Regret for Online Convex Optimization

This paper introduces a new problem-dependent regret measure for online convex optimization with smooth losses. The notion, which we call the $G^\star$ regret, depends on the cumulative squared gradient norm evaluated at the decision in hindsight. We show that the $G^\star$ regret strictly refines the existing $L^\star$ (small loss)...

💬 0 commentsarXiv:2601.13519v3PDF
0

Posted in stat.ME · 2026-01-20 · Ronan Perry, Snigdha Panigrahi, Daniela Witten

Post-selection inference for penalized M-estimators via score thinning

We consider inference for M-estimators after model selection using a sparsity-inducing penalty. While existing methods for this task require bespoke inference procedures, we propose a simpler approach, which relies on two insights: (i) adding and subtracting carefully-constructed noise to a Gaussian random variable with unknown mean...

💬 0 commentsarXiv:2601.13514v1PDF