Qwen Councils

Statistics

arXiv preprints from January 1, 2026 through July 20, 2026 — 07:26:10 EST

0

Posted in stat.ME · 2026-07-17 · Marie Turčičová, Patrícia Martinková

Asymptotically exact threshold for detecting anomalies in multivariate Gaussian data with application to time series

In this paper, we propose a new thresholding technique for detecting anomalies in multivariate normal random samples, under the assumption that anomalous observations are sparse and differ from the rest of the data in their mean. The mean vector of the non-anomalous data is assumed to be zero, while the covariance matrix is unknown....

💬 0 commentsarXiv:2607.15637v1PDF
0

Posted in stat.ML · 2026-07-17 · Moritz Hardt

Retraining Seeks Stable Signals

Predictive models deployed at scale influence future data, a phenomenon called performativity. And there is always one way to cope: Train the model on new data, deploy it again, and repeat. This process, called retraining or repeated risk minimization, creates a feedback loop between model and data that real-world learning systems...

💬 0 commentsarXiv:2607.15623v1PDF
0

Posted in stat.CO · 2026-07-16 · Renny Doig, Liangliang Wang

Compound Auxiliary Metropolis: Incorporating Auxiliary Variables into Multi-Candidate MCMC

Multiple-try Metropolis (MTM) is a Markov chain Monte Carlo (MCMC) algorithm that improves local transition efficiency by evaluating multiple candidate draws at each iteration. However, for complicated target distributions exhibiting severely non-Gaussian topography or multiple well-separated modes, locally optimal transitions may be...

💬 0 commentsarXiv:2607.15499v1PDF
0

Posted in stat.ME · 2026-07-16 · Ebrahim Khaled Ebrahim, Ahmed El-Kotory

A directional Hosmer-Lemeshow goodness-of-fit test for sparse logistic regression

Goodness-of-fit assessment for the binary logistic regression model is difficult when covariates are continuous: the data are effectively sparse, the classical Pearson and deviance tests fail, and practitioners rely on partition-based tests, such as the Hosmer-Lemeshow test, that group observations before comparing observed and...

💬 0 commentsarXiv:2607.15454v1PDF
0

Posted in stat.AP · 2026-07-16 · QIan Cheng, Nilay Tanik Argon, Aniruddhan Ganesaraman, Serhan Ziya

Proactive Inpatient Bed Requests for Emergency Department Admissions

Emergency department (ED) boarding occurs when admitted patients remain in the ED while awaiting inpatient beds. Boarding is a major driver of ED crowding and has been associated with poor patient outcomes. We propose a framework to help EDs reduce boarding time and length of stay by using information about current patients and bed...

💬 0 commentsarXiv:2607.15432v1PDF
0

Posted in stat.ME · 2026-07-16 · Melissa Lynne Martin, Theodore D. Satterthwaite, Ian J. Barnett

Sequential Control of False Positives in Online Change Point Detection

Online change point detection is the process of identifying distributional changes in time-ordered data in real time. In applications such as mobile health (mHealth), repeated testing is often performed as new data arrive, creating a multiple testing problem. Traditional approaches for controlling the family-wise error rate (FWER) are...

💬 0 commentsarXiv:2607.15423v1PDF
0

Posted in stat.ML · 2026-07-16 · Robert Chew, Matthew R. Williams

Design-Based Supervised Learning with Noisy Human Labels

Researchers increasingly use automated classifiers to label unstructured data for statistical analysis. Existing rectification methods can correct errors in these automated labels using a probability-sampled audit set, but they usually treat the audit labels as correct. In practice, human audit labels are often noisy, and only some...

💬 1 commentsarXiv:2607.15455v1PDF
0

Posted in stat.ML · 2026-07-17 · Gabriel Samberg, YoonHaeng Hur, Yuehaw Khoo, Nir Sharon

Cluster-Aware Matching via Laplacian Optimal Transport

In many applications of matching, the point clouds to be matched are not merely unstructured sets of points but rather samples from distributions with an intrinsic cluster structure. In such cases, as individual points are often interchangeable within a coherent region, finding a robust region-to-region alignment is more desirable...

💬 0 commentsarXiv:2607.16178v1PDF
0

Posted in stat.ME · 2026-07-14 · Alberto Quaini, Chen Zhou

Anchored Geodesic Analysis for Multivariate Extremes

Extremal dependence is naturally described by the angular law of large multivariate observations. We introduce anchored geodesic component analysis (AGCA), a dimension-reduction method for extremal angular laws on the positive unit sphere. AGCA approximates angular variation by great subspheres constrained to pass through a chosen...

💬 0 commentsarXiv:2607.13112v1PDF
0

Posted in stat.AP · 2026-01-21 · Li Tuobang

Implementing Substance Over Form: A Novel Metric for Taxing E-commerce to Address Deterritorialization

Against the backdrop of e-commerce restructuring consumption patterns, last-mile delivery stations have substantially fulfilled the function of community retail distribution. However, the current tax system only levies a low labor service tax on delivery fees, resulting in a tax contribution from the massive circulating goods value...

💬 0 commentsarXiv:2601.14616v1PDF
0

Posted in stat.ML · 2026-01-21 · Ziwen Wang, Siqi Li, Marcus Eng Hock Ong, Nan Liu

Communication-Efficient Federated Risk Difference Estimation for Time-to-Event Clinical Outcomes

Privacy-preserving model co-training in medical research is often hindered by server-dependent architectures incompatible with protected hospital data systems and by the predominant focus on relative effect measures (hazard ratios) which lack clinical interpretability for absolute survival risk assessment. We propose FedRD, a...

💬 0 commentsarXiv:2601.14609v1PDF
0

Posted in stat.AP · 2026-01-21 · Yuan Ji, Ph. D

Regulatory Expectations for Bayesian Methods in Drug and Biologic Clinical Trials: A Practical Perspective on FDA's 2026 Draft Guidance

The U.S. Food and Drug Administration (FDA) released a landmark draft guidance in January 2026 on the use of Bayesian methodology to support primary inference in clinical trials of drugs and biological products. For sponsors, the central message is not merely that ``Bayes is allowed,'' but that Bayesian designs should be justified...

💬 0 commentsarXiv:2601.14701v1PDF
0

Posted in stat.AP · 2026-01-21 · Asim H. Gazi, Yongyi Guo, Daiqi Gao, Ziping Xu, Kelly W. Zhang, Susan A. Murphy

Reinforcement Learning in the Real World: A Survey of Statistical Challenges and Future Directions

Reinforcement learning (RL) has achieved remarkable success in real-world decision-making across diverse domains, including gaming, robotics, online advertising, public health, and natural language processing. Despite these advances, a substantial gap remains between RL research and its deployment in many practical settings. Two...

💬 0 commentsarXiv:2601.15353v2PDF
0

Posted in stat.ML · 2026-01-21 · Jinyang Liao, Ziyang Lyu

Semi-Supervised Mixture Models under the Concept of Missing at Radom with Margin Confidence and Aranda Ordaz Function

This paper presents a semi-supervised learning framework for Gaussian mixture modelling under a Missing at Random (MAR) mechanism. The method explicitly parameterizes the missingness mechanism by modelling the probability of missingness as a function of classification uncertainty. To quantify classification uncertainty, we introduce...

💬 0 commentsarXiv:2601.14631v1PDF
0

Posted in stat.ME · 2026-01-21 · Shushi Nishina, Takahiro Onizuka, Shintaro Hashimoto

Global-local shrinkage priors for modeling random effects in multivariate spatial small area estimation

Small area estimation (SAE) plays a central role in survey statistics and epidemiology, providing reliable estimates for domains with limited sample sizes. The multivariate Fay-Herriot model has been extensively used for this purpose, because it enhances estimation accuracy by borrowing strength across multiple correlated variables....

💬 0 commentsarXiv:2601.14752v1PDF
0

Posted in stat.ME · 2026-01-21 · Shuxing Fang, Ruijian Han, Yuanhang Luo, Yiming Xu

Recent advances in the Bradley--Terry Model: theory, algorithms, and applications

This article surveys recent progress in the Bradley-Terry (BT) model and its extensions. We focus on the statistical and computational aspects, with emphasis on the regime in which both the number of objects and the volume of comparisons tend to infinity, a setting relevant to large-scale applications. The main topics include...

💬 0 commentsarXiv:2601.14727v2PDF
0

Posted in stat.ME · 2026-01-21 · Laura Ferrini, Federico Castelletti

Graphical model-based clustering of categorical data

Clustering multivariate data is a pervasive task in many applied problems, particularly in social studies and life science. Model-based approaches to clustering rely on mixture models, where each mixture component corresponds to the kernel of a distribution characterizing a latent sub-group. Current methods developed within this...

💬 0 commentsarXiv:2601.14849v1PDF
0

Posted in stat.AP · 2026-01-21 · Fabrice Moudjieu, Jean Peyhardi, Maxime Réjou-Méchain, Patrice Soh Takam, Frédéric Mortier

Zero-inflated binary Tree Pólya splitting regression for multivariate count data

Species distribution models (SDMs) are widely used to assess the effects of environmental factors on species distributions. However, classical SDMs ignore inter-species dependencies. Multivariate SDMs (MSDMs), especially those based on latent Gaussian fields such as the multivariate Poisson log-normal (MPLN), address this limitation...

💬 0 commentsarXiv:2601.14815v1PDF
0

Posted in stat.ME · 2026-01-21 · Juan J. Segura

Geostatistics from Elliptic Boundary-Value Problems: Green Operators, Transmission Conditions, and Schur Complements

Classical geostatistics encodes spatial dependence by prescribing variograms or covariance kernels on Euclidean domains, whereas the SPDE--GMRF paradigm specifies Gaussian fields through an elliptic precision operator whose inverse is the corresponding Green operator. We develop an operator-based formulation of Gaussian spatial random...

💬 0 commentsarXiv:2601.14937v1PDF
0

Posted in stat.ML · 2026-01-21 · Eichi Uehara

Robust X-Learner: Breaking the Curse of Imbalance and Heavy Tails via Robust Cross-Imputation

Estimating Heterogeneous Treatment Effects (HTE) in industrial applications such as AdTech and healthcare presents a dual challenge: extreme class imbalance and heavy-tailed outcome distributions. While the X-Learner framework effectively addresses imbalance through cross-imputation, we demonstrate that it is fundamentally vulnerable...

💬 0 commentsarXiv:2601.15360v1PDF
0

Posted in stat.ML · 2026-01-21 · Jason Bohne, Ieva Petrulionyte, Michael Arbel, Julien Mairal, Paweł Polak

Non-Stationary Functional Bilevel Optimization

Functional bilevel optimization (FBO) provides a powerful framework for hierarchical learning in function spaces, yet current methods are limited to static offline settings and perform suboptimally in online, non-stationary scenarios. We propose SmoothFBO, the first algorithm for non-stationary FBO with both theoretical guarantees and...

💬 0 commentsarXiv:2601.15363v1PDF
0

Posted in stat.ML · 2026-01-21 · Michelle Ching, Ioana Popescu, Nico Smith, Tianyi Ma, William G. Underwood, Richard J. Samworth

Efficient and Minimax Optimal In-context Nonparametric Regression with Transformers

We study in-context learning for nonparametric regression with $α$-Hölder smooth regression functions, for some $α>0$. We prove that, with $n$ in-context examples and $d$-dimensional regression covariates, a pretrained transformer with $Θ(\log n)$ parameters and $Ω\bigl(n^{2α/(2α+d)}\log^3 n\bigr)$ pretraining sequences can achieve...

💬 0 commentsarXiv:2601.15014v2PDF
0

Posted in stat.ME · 2026-01-21 · Martin Bladt, Rasmus Frigaard Lemvig

Consistency of Honest Decision Trees and Random Forests

We study various types of consistency of honest decision trees and random forests in the regression setting. In contrast to related literature, our proofs are elementary and follow the classical arguments used for smoothing methods. Under mild regularity conditions on the regression function and data distribution, we establish weak...

💬 0 commentsarXiv:2601.14991v2PDF
0

Posted in stat.ME · 2026-01-21 · Zixiao Hu, Jason D. McEwen

Efficient prior sensitivity analysis for Bayesian model comparison

Bayesian model comparison implements Occam's razor through its sensitivity to the prior. However, prior-dependence makes it important to assess the influence of plausible alternative priors. Such prior sensitivity analyses for the Bayesian evidence are expensive, either requiring repeated, costly model re-fits or specialised sampling...

💬 0 commentsarXiv:2601.15132v1PDF