Qwen Councils

Statistics

arXiv preprints from January 1, 2026 through September 21, 2026 — 20:51:27 EST

0

Posted in stat.ML · 2026-01-14 · Hong Ye Tan, Stanley Osher, Wuchen Li

Accelerated Regularized Wasserstein Proximal Sampling Algorithms

We consider sampling from a Gibbs distribution by evolving a finite number of particles using a particular score estimator rather than Brownian motion. To accelerate the particles, we consider a second-order score-based ODE, similar to Nesterov acceleration. In contrast to traditional kernel density score estimation, we use the...

💬 0 commentsarXiv:2601.09848v2PDF
0

Posted in stat.ML · 2026-01-13 · Xinping Yi, Gaojie Jin, Xiaowei Huang, Shi Jin

Towards A Unified PAC-Bayesian Framework for Norm-based Generalization Bounds

Understanding the generalization behavior of deep neural networks remains a fundamental challenge in modern statistical learning theory. Among existing approaches, PAC-Bayesian norm-based bounds have demonstrated particular promise due to their data-dependent nature and their ability to capture algorithmic and geometric properties of...

💬 0 commentsarXiv:2601.08100v1PDF
0

Posted in stat.OT · 2026-01-13 · Yuxi Zhao, Margaret Gamalo

Proactive Anomaly Screen for Multiple Endpoints Using Bayesian Latent Class Modeling: A k-Step Ahead Approach

In clinical trials, ensuring the quality and validity of data for downstream analysis and results is paramount, thus necessitating thorough data monitoring. This typically involves employing edit checks and manual queries during data collection. Edit checks consist of straightforward schemes programmed into relational databases,...

💬 0 commentsarXiv:2601.08167v2PDF
0

Posted in stat.ML · 2026-01-13 · Pei Heng, Yi Sun, Jianhua Guo

Structural Dimension Reduction in Bayesian Networks

This work introduces a novel technique, named structural dimension reduction, to collapse a Bayesian network onto a minimum and localized one while ensuring that probabilistic inferences between the original and reduced networks remain consistent. To this end, we propose a new combinatorial structure in directed acyclic graphs called...

💬 0 commentsarXiv:2601.08236v1PDF
0

Posted in stat.AP · 2026-01-13 · Pierre Ailliot, Carlo Gaetan, Philippe Naveau

A parsimonious tail compliant multiscale statistical model for aggregated rainfall

Modeling rainfall intensity distributions across aggregation scales (from sub-hourly to weekly) is essential for hydrological risk analysis and IDF curves. Aggregation naturally imposes mathematical constraints: return levels must be ordered by time scale, as daily accumulations necessarily exceed sub-daily ones. From a statistical...

💬 0 commentsarXiv:2601.08350v1PDF
0

Posted in stat.CO · 2026-01-13 · Abylay Zhumekenov, Alexandros Beskos, Dan Crisan, Ajay Jasra, Nikolas Kantas

Particle Filtering for a Class of State-Space Models with Low and Degenerate Observational Noise

We consider the discrete-time filtering problem in scenarios where the observation noise is low or degenerate. We focus on the case where the observation equation is a linear function of the state and the data involve additive noise. However, we place minimal assumptions on the hidden state process. For such a class of models we...

💬 0 commentsarXiv:2601.08411v2PDF
0

Posted in stat.AP · 2026-01-13 · Shi-Shun Chen, Dong-Hua Niu, Wen-Bin Chen, Jia-Yun Song, Ya-Fei Zhang, Xiao-Yang Li, Enrico Zio

Reliability Modeling of Single-Sided Aluminized Polyimide Films during Storage Considering Stress-Induced Degradation Mechanism Transition

Single-sided aluminized polyimide films (SAPF) are widely used in thermal management of aerospace systems. Although the reliability of SAPF in space environments has been thoroughly studied, its reliability in ground environments during storage is always ignored, potentially leading to system failure. This paper aims to investigate...

💬 0 commentsarXiv:2601.08655v1PDF
0

Posted in stat.ML · 2026-01-13 · The Tien Mai

Robust low-rank estimation with multiple binary responses using pairwise AUC loss

Multiple binary responses arise in many modern data-analytic problems. Although fitting separate logistic regressions for each response is computationally attractive, it ignores shared structure and can be statistically inefficient, especially in high-dimensional and class-imbalanced regimes. Low-rank models offer a natural way to...

💬 0 commentsarXiv:2601.08618v1PDF
0

Posted in stat.ME · 2026-01-13 · Wenxuan Guo, Panos Toulis, Yuhao Wang

Permutation Inference under Multi-way Clustering and Missing Data

Econometric applications with multi-way clustering often feature a small number of effective clusters or heavy-tailed data, making standard cluster-robust and bootstrap inference unreliable in finite samples. In this paper, we develop a framework for finite-sample valid permutation inference in linear regression with multi-way...

💬 0 commentsarXiv:2601.08610v1PDF
0

Posted in stat.ME · 2026-01-13 · Rodrigo M. R. de Medeiros, Francisco F. Queiroz

Flexible modeling of nonnegative continuous data: Box-Cox symmetric regression and its zero-adjusted extension

The Box-Cox symmetric distributions constitute a broad class of probability models for positive continuous data, offering flexibility in modeling skewness and tail behavior. Their parameterization allows a straightforward quantile-based interpretation, which is particularly useful in regression modeling. Despite their potential, only...

💬 0 commentsarXiv:2601.08600v2PDF
0

Posted in stat.ME · 2026-01-13 · Marcus Gehrmann, Håkon Tjelmeland

Sparsifying transform priors in Gaussian graphical models

Bayesian methods constitute a popular approach for estimating the conditional independence structure in Gaussian graphical models, since they can quantify the uncertainty through the posterior distribution. Inference in this framework is typically carried out with Markov chain Monte Carlo (MCMC). However, the most widely used choice...

💬 0 commentsarXiv:2601.08596v1PDF
0

Posted in stat.ME · 2026-01-13 · Ping Zhao, Long Feng

Note on High Dimensional Spatial-Sign Test for One Sample Problem

We revisit the null distribution of the high-dimensional spatial-sign test of Wang et al. (2015) under mild structural assumptions on the scatter matrix. We show that the standardized test statistic converges to a non-Gaussian limit, characterized as a mixture of a normal component and a weighted chi-square component. To facilitate...

💬 0 commentsarXiv:2601.08736v1PDF
0

Posted in stat.ME · 2026-01-13 · Kosuke Morikawa, Jae Kwang Kim

Semiparametric Efficient Data Integration Using the Dual-Frame Sampling Framework

Integrating probability and non-probability samples is increasingly important, yet unknown sampling mechanisms in non-probability sources complicate identification and efficient estimation. We develop semiparametric theory for dual-frame data integration and propose two complementary estimators. The first models the non-probability...

💬 0 commentsarXiv:2601.08707v1PDF
0

Posted in stat.ML · 2026-01-13 · Arturo Pérez-Peralta, Sandra Benítez-Peña, Rosa E. Lillo

On the use of graph models to achieve individual and group fairness

Machine Learning algorithms are ubiquitous in key decision-making contexts such as justice, healthcare and finance, which has spawned a great demand for fairness in these procedures. However, the theoretical properties of such models in relation with fairness are still poorly understood, and the intuition behind the relationship...

💬 0 commentsarXiv:2601.08784v1PDF
0

Posted in stat.ML · 2026-01-13 · Nawaf Bou-Rabee, Siddharth Mitra, Andre Wibisono

Tail-Sensitive KL and Rényi Convergence of Unadjusted Hamiltonian Monte Carlo via One-Shot Couplings

Hamiltonian Monte Carlo (HMC) algorithms are among the most widely used sampling methods in high dimensional settings, yet their convergence properties are poorly understood in divergences that quantify relative density mismatch, such as Kullback-Leibler (KL) and Rényi divergences. These divergences naturally govern acceptance...

💬 0 commentsarXiv:2601.09019v1PDF
0

Posted in stat.ME · 2026-01-13 · Steven A. Frank

Causal attribution by the chain rule: unifying natural selection, learning, economics, and other disciplines

Analysis often splits change into components. For example, how much of the observed variance is caused by genes or environment? In many cases, the split is ultimately made by the logic of the chain rule, which divides the difference of a product into two terms. Each term quantifies the partial difference associated with change in one...

💬 0 commentsarXiv:2601.09011v3PDF
0

Posted in stat.ME · 2026-01-13 · Andrea Toloba, Klaus Langohr, Guadalupe Gómez Melis

Semiparametric estimation of GLMs with interval-censored covariates via an augmented Turnbull estimator

Interval-censored covariates are frequently encountered in biomedical studies, particularly in time-to-event data or when measurements are subject to detection or quantification limits. Yet, the estimation of regression models with interval-censored covariates remains methodologically underdeveloped. In this article, we address the...

💬 0 commentsarXiv:2601.08996v1PDF
0

Posted in stat.ME · 2026-01-13 · Fredrik Lohne Aanes

Approximate Shapley value estimation using sampling without replacement and variance estimation via the new Symmetric bootstrap and the Doubled half bootstrap

In this paper I consider improving the KernelSHAP algorithm. I suggest to use the Wallenius' noncentral hypergeometric distribution for sampling the number of coalitions and perform sampling without replacement, so that the KernelSHAP estimation framework is improved further. I also introduce the Symmetric bootstrap to calculate the...

💬 0 commentsarXiv:2601.08981v1PDF
0

Posted in stat.ME · 2026-01-13 · Jiahao Tian, Hugh Chipman, Thomas Loughin

MLCBART: Multilabel Classification with Bayesian Additive Regression Trees

Multilabel Classification (MLC) deals with the simultaneous classification of multiple binary labels. The task is challenging because, not only may there be arbitrarily different and complex relationships between predictor variables and each label, but associations among labels may exist even after accounting for effects of predictor...

💬 0 commentsarXiv:2601.08964v1PDF
0

Posted in stat.ME · 2026-01-12 · Hisaya Okahara, Tomoyuki Nakagawa, Shonosuke Sugasawa

The Covariate-Assisted Bayesian Intransitive Bradley-Terry Model via Combinatorial Hodge Theory

Pairwise comparison data are widely used to recover latent rankings, yet the models in dominant use assume stochastic transitivity. When preferences are in fact intransitive, a single scalar strength conflates genuine hierarchy with cycle-induced structure, biasing both the recovered ranking and any covariate effects attributed to it....

💬 0 commentsarXiv:2601.07158v2PDF
0

Posted in stat.ML · 2026-01-12 · Linus Bleistein, Mathieu Dagréou, Francisco Andrade, Thomas Boudou, Aurélien Bellet

Optimal Transport under Group Fairness Constraints

Ensuring fairness in matching algorithms is a key challenge in allocating scarce resources and positions. Focusing on Optimal Transport (OT), we introduce a novel notion of group fairness requiring that the probability of matching two individuals from any two given groups in the OT plan satisfies a predefined target. We first propose...

💬 0 commentsarXiv:2601.07144v3PDF
0

Posted in stat.ME · 2026-01-12 · Shuli Chen, Jie Hu, Zhichao Jiang

Connections as treatment: causal inference with edge interventions in networks

Causal inference has traditionally focused on interventions at the unit level. In many applications, however, the central question concerns the causal effects of connections between units, such as transportation links, social relationships, or collaborative ties. We develop a causal framework for edge interventions in networks, where...

💬 0 commentsarXiv:2601.07267v1PDF
0

Posted in stat.ME · 2026-01-12 · Suchismita Das, Akul Ameya, Cahyani Karunia Putri

Compounded Linear Failure Rate Distribution: Properties, Simulation and Analysis

This paper proposes a new extension of the linear failure rate (LFR) model to better capture real-world lifetime data. The model incorporates an additional shape parameter to increase flexibility. It helps model the minimum survival time from a set of LFR distributed variables. We define the model, derive certain statistical...

💬 0 commentsarXiv:2601.07249v1PDF
0

Posted in stat.ML · 2026-01-12 · Yiran Jia, Jelena Bradic

Multi-environment Invariance Learning with Missing Data

Learning models that can handle distribution shifts is a key challenge in domain generalization. Invariance learning, an approach that focuses on identifying features invariant across environments, improves model generalization by capturing stable relationships, which may represent causal effects when the data distribution is encoded...

💬 0 commentsarXiv:2601.07247v2PDF
0

Posted in stat.ME · 2026-01-12 · Kanji Goto, Shintaro Yuki, Kensuke Tanioka, Hiroshi Yadohisa

Principal component-guided sparse reduced-rank regression

Reduced-rank regression estimates regression coefficients by imposing a low-rank constraint on the matrix of regression coefficients, thereby accounting for correlations among response variables. To further improve predictive accuracy and model interpretability, several regularized reduced-rank regression methods have been proposed....

💬 0 commentsarXiv:2601.07202v2PDF