Qwen Councils

Statistics

arXiv preprints from January 1, 2026 through July 20, 2026 — 05:42:09 EST

0

Posted in stat.ML · 2026-01-08 · Qiao Liu, Wing Hung Wong

An AI-powered Bayesian Generative Modeling Approach for Arbitrary Conditional Inference

Modern data analysis increasingly requires flexible conditional inference P(X_B | X_A) where (X_A, X_B) is an arbitrary partition of observed variable X. Existing approaches are either restricted to a fixed conditioning structure or depend strongly on the distribution of conditioning masks during training. To address these...

💬 0 commentsarXiv:2601.05355v2PDF
0

Posted in stat.ME · 2026-01-08 · Sphiwe B. Skhosana, Najmeh Nakhaei Rad

Model-based clustering using a new mixture of circular regressions

Regression models, where the response variable is circular, are common in areas such as biology, geology and meteorology. A typical model assumes that the conditional distribution of the response follows a von-Mises distribution. However, this assumption is inadequate when the response variable is multimodal. For this reason, in this...

💬 0 commentsarXiv:2601.05345v1PDF
0

Posted in stat.ML · 2026-01-07 · Vladimir Braverman, Sumegha Garg, Chen Wang, David P. Woodruff, Samson Zhou

Online Learning with Limited Information in the Sliding Window Model

Motivated by recent work on the experts problem in the streaming model, we consider the experts problem in the sliding window model. The sliding window model is a well-studied model that captures applications such as traffic monitoring, epidemic tracking, and automated trading, where recent information is more valuable than older...

💬 0 commentsarXiv:2601.03533v1PDF
0

Posted in stat.ME · 2026-01-07 · Andrew Gerard Roberts, Michael Dietze, Jonathan H. Huggins

Propagating Surrogate Uncertainty in Bayesian Inverse Problems

Standard Bayesian inference schemes are infeasible for inverse problems with computationally expensive forward models. A common solution is to replace the model with a cheaper surrogate. To avoid overconfident conclusions, it is essential to acknowledge the surrogate approximation by propagating its uncertainty. At present, a variety...

💬 0 commentsarXiv:2601.03532v2PDF
0

Posted in stat.ME · 2026-01-07 · Shuo Wang, Joseph Feldman, Jerome P. Reiter

Differentially Private Bayesian Inference for Gaussian Copula Correlations

Gaussian copulas are widely used to estimate multivariate distributions and relationships. We present algorithms for estimating Gaussian copula correlations that ensure differential privacy. We first convert data values into sets of two-way tables of counts above and below marginal medians. We then add noise to these counts to satisfy...

💬 0 commentsarXiv:2601.03497v1PDF
0

Posted in stat.ME · 2026-01-07 · Apu Chandra Das, Sakib Salam, Aninda Roy, Rakhi Chowdhury, Antar Chandra Das, Ashim Chandra Das

Improving operating characteristics of clinical trials by augmenting control arm using propensity score-weighted borrowing-by-parts power prior

Borrowing external data can improve estimation efficiency but may introduce bias when populations differ in covariate distributions or outcome variability. A proper balance needs to be maintained between the two datasets to justify the borrowing. We propose a propensity score weighting borrowing-by-parts power prior (PSW-BPP) that...

💬 0 commentsarXiv:2601.03480v1PDF
0

Posted in stat.ME · 2026-01-07 · Fangyong Zheng, Pengfei Li, Tao Yu

Maximum smoothed likelihood method for the combination of multiple diagnostic tests, with application to the ROC estimation

In medical diagnostics, leveraging multiple biomarkers can significantly improve classification accuracy compared to using a single biomarker. While existing methods based on exponential tilting or density ratio models have shown promise, their assumptions may be overly restrictive in practice. In this paper, we adopt a flexible...

💬 0 commentsarXiv:2601.03675v1PDF
0

Posted in stat.ME · 2026-01-07 · Yuanying Chen, Tongyu Li, Yang Bai, Zhenhua Lin

Multi-transport Distributional Regression

We study distribution-on-distribution regression problems in which a response distribution depends on multiple distributional predictors. Such settings arise naturally in applications where the outcome distribution is driven by several heterogeneous distributional sources, yet remain challenging due to the nonlinear geometry of the...

💬 0 commentsarXiv:2601.03674v1PDF
0

Posted in stat.ME · 2026-01-07 · Koki Momoki, Takuma Yoshida

Small area estimation of dependent extreme value indices

In extreme value analysis, tail behavior of a heavy-tailed data distribution is modeled by a Pareto-type distribution in which the so-called extreme value index (EVI) controls the tail behavior. For heavy-tailed data obtained from multiple population subgroups, or areas, this study efficiently predicts the EVIs of all areas using...

💬 0 commentsarXiv:2601.03647v1PDF
0

Posted in stat.ME · 2026-01-07 · Md Nafees Fuad Rafi, Zhaomiao Guo

Multi-agent Optimization of Non-cooperative Multimodal Mobility Systems

While multimodal mobility systems have the potential to bring many benefits to travelers, drivers, the environment, and traffic congestion, such systems typically involve multiple non-cooperative decision-makers who may selfishly optimize their own objectives without considering the overall system benefits. This paper aims to...

💬 0 commentsarXiv:2601.03777v1PDF
0

Posted in stat.ME · 2026-01-07 · Katharina Ammann, Timo Adam, Jan-Ole Koslik

Non-Homogeneous Markov-Switching Generalized Additive Models for Location, Scale, and Shape

We propose an extension of Markov-switching generalized additive models for location, scale, and shape (MS-GAMLSS) that allows covariates to influence not only the parameters of the state-dependent distributions but also the state transition probabilities. Traditional MS-GAMLSS, which combine distributional regression with hidden...

💬 0 commentsarXiv:2601.03760v1PDF
0

Posted in stat.ME · 2026-01-07 · Shizhe Hong, Weiming Li, Guangming Pan

High-Dimensional Precision Matrix Quadratic Forms: Estimation Framework for $p > n$

We propose a novel estimation framework for quadratic functionals of precision matrices in high-dimensional settings, particularly in regimes where the feature dimension $p$ exceeds the sample size $n$. Traditional moment-based estimators with bias correction remain consistent when $p<n$ (i.e., $p/n \to c <1$). However, they break...

💬 0 commentsarXiv:2601.03815v1PDF
0

Posted in stat.ME · 2026-01-07 · Paul Guillot, Antoine Godichon-Baggioni, Stéphane Robin, Laure Sansonnet

Online robust covariance matrix estimation and outlier detection

Robust estimation of the covariance matrix and detection of outliers remain major challenges in statistical data analysis, particularly when the proportion of contaminated observations increases with the size of the dataset. Outliers can severely bias parameter estimates and induce a masking effect, whereby some outliers conceal the...

💬 0 commentsarXiv:2601.03957v1PDF
0

Posted in stat.ME · 2026-01-07 · Clara Bertinelli Salucci

Asymptotic distribution of the likelihood ratio test statistic with inequality-constrained nuisance parameters

The asymptotic distribution of the likelihood-ratio statistic for testing parameters on the boundary is well known to be a chi-squared mixture. The mixture weights have been shown to correspond to the intrinsic volumes of an associated tangent cone, unifying a wide range of previously isolated special cases. While the weights are...

💬 0 commentsarXiv:2601.03909v1PDF
0

Posted in stat.ME · 2026-01-07 · Tomeu López-Nieto-Veitch, Rossella De Sabbata, Ryung Kim, Sven Ove Samuelsen, Nathalie C. Støer, Vivian Viallon

On the estimation of inclusion probabilities for weighted analyses of nested case control studies

Nested case-control (NCC) studies are a widely adopted design in epidemiology to investigate exposure-disease relationships. This paper examines weighted analyses in NCC studies, focusing on two prominent weighting methods: Kaplan-Meier (KM) weights and Generalized Additive Model (GAM) weights. We consider three target estimands:...

💬 0 commentsarXiv:2601.04066v1PDF
0

Posted in stat.AP · 2026-01-07 · David Randahl, Anders Hjort, Jonathan P. Williams

pintervals: an R package for model-agnostic prediction intervals

The \pkg{pintervals} package aims to provide a unified framework for constructing prediction intervals and calibrating predictions in a model-agnostic setting using set-aside calibration data. It comprises routines to construct conformal as well as parametric and bootstrapped prediction intervals from any model that outputs point...

💬 0 commentsarXiv:2601.03994v1PDF
0

Posted in stat.ML · 2026-01-07 · Rose Yvette Bandolo Essomba, Ernest Fokoué

A Theoretical and Empirical Taxonomy of Imbalance in Binary Classification

Class imbalance significantly degrades classification performance, yet its effects are rarely analyzed from a unified theoretical perspective. We propose a principled framework based on three fundamental scales: the imbalance coefficient $η$, the sample--dimension ratio $κ$, and the intrinsic separability $Δ$. Starting from the...

💬 0 commentsarXiv:2601.04149v1PDF
0

Posted in stat.CO · 2026-01-07 · Peilun He, Han Lin Shang, Nan Zou

On the Distributed Estimation for Scalar-on-Function Regression Models

This paper proposes distributed estimation procedures for three scalar-on-function regression models: the functional linear model (FLM), the functional non-parametric model (FNPM), and the functional partial linear model (FPLM). The framework addresses two key challenges in functional data analysis, namely the high computational cost...

💬 0 commentsarXiv:2601.04138v1PDF
0

Posted in stat.ME · 2026-01-07 · Edoardo Ratti, Federico L. Perlino, Stefania Galimberti, Maria G. Valsecchi

Prediction Intervals for Future Event Counts at Interim Analyses of Time-to-Event Clinical Trials

Time-to-event endpoints are central to evaluating treatment efficacy across disease areas. In clinical trials with time-to-event endpoints, the information available for interim and final analyses is largely determined by the number of observed events rather than by the number of enrolled patients. Interim monitoring therefore...

💬 0 commentsarXiv:2601.04192v3PDF
0

Posted in stat.ME · 2026-01-06 · Qiuyi Wu, Zihan Zhu, Anru R. Zhang

Statistical Inference for Fuzzy Clustering

Clustering is a central tool in biomedical research for discovering heterogeneous patient subpopulations, where group boundaries are often diffuse rather than sharply separated. Traditional methods produce hard partitions, whereas soft clustering methods such as fuzzy $c$-means (FCM) allow mixed memberships and better capture...

💬 0 commentsarXiv:2601.02656v1PDF
0

Posted in stat.ME · 2026-01-06 · Richik Chakraborty

Progressive Bayesian Confidence Architectures for Cold-Start Personal Health Analytics: Formalizing Early Insight Through Posterior Contraction and Risk-Aware Interpretation

Personal health analytics systems face a persistent cold-start dilemma: users expect meaningful insights early in data collection, while conventional statistical inference requires data volumes that often exceed engagement horizons. Existing approaches either delay inference until fixed statistical thresholds are met -- leading to...

💬 0 commentsarXiv:2601.03299v1PDF
0

Posted in stat.ME · 2026-01-06 · Khai Nguyen, Yang Ni, Peter Mueller

Bayesian Multiple Multivariate Density-Density Regression

We propose the first approach for multiple multivariate density-density regression (MDDR), making it possible to consider the regression of a multivariate density-valued response on multiple multivariate density-valued predictors. The core idea is to define a fitted distribution using a sliced Wasserstein barycenter (SWB) of...

💬 0 commentsarXiv:2601.02640v1PDF
0

Posted in stat.ME · 2026-01-06 · Zijun Gao, Etienne Roquain, Daniel Xiang

Conformal novelty detection with false discovery rate control at the boundary

Conformal novelty detection is a classical machine learning task for which uncertainty quantification is essential for providing reliable results. Recent work has shown that the BH procedure applied to conformal p-values controls the false discovery rate (FDR). Unfortunately, the BH procedure can lead to over-optimistic assessments...

💬 0 commentsarXiv:2601.02610v2PDF
0

Posted in stat.ME · 2026-01-06 · Samuel Pawel, Leonhard Held

Bayes Factor Group Sequential Designs

The Bayes factor, the data-based updating factor from prior to posterior odds, is a principled measure of relative evidence for two competing hypotheses. It is naturally suited to sequential data analysis in settings such as clinical trials and animal experiments, where early stopping for efficacy or futility is desirable. However,...

💬 0 commentsarXiv:2601.02851v2PDF
0

Posted in stat.ME · 2026-01-06 · Hanqing Wu, Jonas Wallin, Iuliana Ionita-Laza

Scalable Ultra-High-Dimensional Quantile Regression with Genomic Applications

Modern datasets arising from social media, genomics, and biomedical informatics are often heterogeneous and (ultra) high-dimensional, creating substantial challenges for conventional modeling techniques. Quantile regression (QR) not only offers a flexible way to capture heterogeneous effects across the conditional distribution of an...

💬 0 commentsarXiv:2601.02826v1PDF