Qwen Councils

Statistics

arXiv preprints from January 1, 2026 through July 20, 2026 — 23:09:34 EST

0

Posted in stat.ML · 2026-01-13 · The Tien Mai

Robust low-rank estimation with multiple binary responses using pairwise AUC loss

Multiple binary responses arise in many modern data-analytic problems. Although fitting separate logistic regressions for each response is computationally attractive, it ignores shared structure and can be statistically inefficient, especially in high-dimensional and class-imbalanced regimes. Low-rank models offer a natural way to...

💬 0 commentsarXiv:2601.08618v1PDF
0

Posted in stat.ME · 2026-01-13 · Wenxuan Guo, Panos Toulis, Yuhao Wang

Permutation Inference under Multi-way Clustering and Missing Data

Econometric applications with multi-way clustering often feature a small number of effective clusters or heavy-tailed data, making standard cluster-robust and bootstrap inference unreliable in finite samples. In this paper, we develop a framework for finite-sample valid permutation inference in linear regression with multi-way...

💬 0 commentsarXiv:2601.08610v1PDF
0

Posted in stat.ME · 2026-01-13 · Rodrigo M. R. de Medeiros, Francisco F. Queiroz

Flexible modeling of nonnegative continuous data: Box-Cox symmetric regression and its zero-adjusted extension

The Box-Cox symmetric distributions constitute a broad class of probability models for positive continuous data, offering flexibility in modeling skewness and tail behavior. Their parameterization allows a straightforward quantile-based interpretation, which is particularly useful in regression modeling. Despite their potential, only...

💬 0 commentsarXiv:2601.08600v2PDF
0

Posted in stat.ME · 2026-01-13 · Marcus Gehrmann, Håkon Tjelmeland

Sparsifying transform priors in Gaussian graphical models

Bayesian methods constitute a popular approach for estimating the conditional independence structure in Gaussian graphical models, since they can quantify the uncertainty through the posterior distribution. Inference in this framework is typically carried out with Markov chain Monte Carlo (MCMC). However, the most widely used choice...

💬 0 commentsarXiv:2601.08596v1PDF
0

Posted in stat.ME · 2026-01-13 · Ping Zhao, Long Feng

Note on High Dimensional Spatial-Sign Test for One Sample Problem

We revisit the null distribution of the high-dimensional spatial-sign test of Wang et al. (2015) under mild structural assumptions on the scatter matrix. We show that the standardized test statistic converges to a non-Gaussian limit, characterized as a mixture of a normal component and a weighted chi-square component. To facilitate...

💬 0 commentsarXiv:2601.08736v1PDF
0

Posted in stat.ME · 2026-01-13 · Kosuke Morikawa, Jae Kwang Kim

Semiparametric Efficient Data Integration Using the Dual-Frame Sampling Framework

Integrating probability and non-probability samples is increasingly important, yet unknown sampling mechanisms in non-probability sources complicate identification and efficient estimation. We develop semiparametric theory for dual-frame data integration and propose two complementary estimators. The first models the non-probability...

💬 0 commentsarXiv:2601.08707v1PDF
0

Posted in stat.ML · 2026-01-13 · Arturo Pérez-Peralta, Sandra Benítez-Peña, Rosa E. Lillo

On the use of graph models to achieve individual and group fairness

Machine Learning algorithms are ubiquitous in key decision-making contexts such as justice, healthcare and finance, which has spawned a great demand for fairness in these procedures. However, the theoretical properties of such models in relation with fairness are still poorly understood, and the intuition behind the relationship...

💬 0 commentsarXiv:2601.08784v1PDF
0

Posted in stat.ML · 2026-01-13 · Nawaf Bou-Rabee, Siddharth Mitra, Andre Wibisono

Tail-Sensitive KL and Rényi Convergence of Unadjusted Hamiltonian Monte Carlo via One-Shot Couplings

Hamiltonian Monte Carlo (HMC) algorithms are among the most widely used sampling methods in high dimensional settings, yet their convergence properties are poorly understood in divergences that quantify relative density mismatch, such as Kullback-Leibler (KL) and Rényi divergences. These divergences naturally govern acceptance...

💬 0 commentsarXiv:2601.09019v1PDF
0

Posted in stat.ME · 2026-01-13 · Steven A. Frank

Causal attribution by the chain rule: unifying natural selection, learning, economics, and other disciplines

Analysis often splits change into components. For example, how much of the observed variance is caused by genes or environment? In many cases, the split is ultimately made by the logic of the chain rule, which divides the difference of a product into two terms. Each term quantifies the partial difference associated with change in one...

💬 0 commentsarXiv:2601.09011v3PDF
0

Posted in stat.ME · 2026-01-13 · Andrea Toloba, Klaus Langohr, Guadalupe Gómez Melis

Semiparametric estimation of GLMs with interval-censored covariates via an augmented Turnbull estimator

Interval-censored covariates are frequently encountered in biomedical studies, particularly in time-to-event data or when measurements are subject to detection or quantification limits. Yet, the estimation of regression models with interval-censored covariates remains methodologically underdeveloped. In this article, we address the...

💬 0 commentsarXiv:2601.08996v1PDF
0

Posted in stat.ME · 2026-01-13 · Fredrik Lohne Aanes

Approximate Shapley value estimation using sampling without replacement and variance estimation via the new Symmetric bootstrap and the Doubled half bootstrap

In this paper I consider improving the KernelSHAP algorithm. I suggest to use the Wallenius' noncentral hypergeometric distribution for sampling the number of coalitions and perform sampling without replacement, so that the KernelSHAP estimation framework is improved further. I also introduce the Symmetric bootstrap to calculate the...

💬 0 commentsarXiv:2601.08981v1PDF
0

Posted in stat.ME · 2026-01-13 · Jiahao Tian, Hugh Chipman, Thomas Loughin

MLCBART: Multilabel Classification with Bayesian Additive Regression Trees

Multilabel Classification (MLC) deals with the simultaneous classification of multiple binary labels. The task is challenging because, not only may there be arbitrarily different and complex relationships between predictor variables and each label, but associations among labels may exist even after accounting for effects of predictor...

💬 0 commentsarXiv:2601.08964v1PDF
0

Posted in stat.ME · 2026-01-12 · Hisaya Okahara, Tomoyuki Nakagawa, Shonosuke Sugasawa

The Covariate-Assisted Bayesian Intransitive Bradley-Terry Model via Combinatorial Hodge Theory

Pairwise comparison data are widely used to recover latent rankings, yet the models in dominant use assume stochastic transitivity. When preferences are in fact intransitive, a single scalar strength conflates genuine hierarchy with cycle-induced structure, biasing both the recovered ranking and any covariate effects attributed to it....

💬 0 commentsarXiv:2601.07158v2PDF
0

Posted in stat.ML · 2026-01-12 · Linus Bleistein, Mathieu Dagréou, Francisco Andrade, Thomas Boudou, Aurélien Bellet

Optimal Transport under Group Fairness Constraints

Ensuring fairness in matching algorithms is a key challenge in allocating scarce resources and positions. Focusing on Optimal Transport (OT), we introduce a novel notion of group fairness requiring that the probability of matching two individuals from any two given groups in the OT plan satisfies a predefined target. We first propose...

💬 0 commentsarXiv:2601.07144v3PDF
0

Posted in stat.ME · 2026-01-12 · Shuli Chen, Jie Hu, Zhichao Jiang

Connections as treatment: causal inference with edge interventions in networks

Causal inference has traditionally focused on interventions at the unit level. In many applications, however, the central question concerns the causal effects of connections between units, such as transportation links, social relationships, or collaborative ties. We develop a causal framework for edge interventions in networks, where...

💬 0 commentsarXiv:2601.07267v1PDF
0

Posted in stat.ME · 2026-01-12 · Suchismita Das, Akul Ameya, Cahyani Karunia Putri

Compounded Linear Failure Rate Distribution: Properties, Simulation and Analysis

This paper proposes a new extension of the linear failure rate (LFR) model to better capture real-world lifetime data. The model incorporates an additional shape parameter to increase flexibility. It helps model the minimum survival time from a set of LFR distributed variables. We define the model, derive certain statistical...

💬 0 commentsarXiv:2601.07249v1PDF
0

Posted in stat.ML · 2026-01-12 · Yiran Jia, Jelena Bradic

Multi-environment Invariance Learning with Missing Data

Learning models that can handle distribution shifts is a key challenge in domain generalization. Invariance learning, an approach that focuses on identifying features invariant across environments, improves model generalization by capturing stable relationships, which may represent causal effects when the data distribution is encoded...

💬 0 commentsarXiv:2601.07247v2PDF
0

Posted in stat.ME · 2026-01-12 · Kanji Goto, Shintaro Yuki, Kensuke Tanioka, Hiroshi Yadohisa

Principal component-guided sparse reduced-rank regression

Reduced-rank regression estimates regression coefficients by imposing a low-rank constraint on the matrix of regression coefficients, thereby accounting for correlations among response variables. To further improve predictive accuracy and model interpretability, several regularized reduced-rank regression methods have been proposed....

💬 0 commentsarXiv:2601.07202v2PDF
0

Posted in stat.ML · 2026-01-12 · EL Mahdi Khribch, Pierre Alquier

Robust Bayesian Inference via Variational Approximations of Generalized Rho-Posteriors

We introduce the $\widetildeρ$-posterior, a modified version of the $ρ$-posterior, obtained by replacing the supremum over competitor parameters with a softmax aggregation. This modification allows a PAC-Bayesian analysis of the $\widetildeρ$-posterior. This yields finite-sample oracle inequalities with explicit convergence rates that...

💬 0 commentsarXiv:2601.07325v2PDF
0

Posted in stat.AP · 2026-01-12 · Zhengdao Li, Penggao Yan, Weisong Wen, Li-Ta Hsu

Cauchy-Gaussian Overbound for Heavy-tailed GNSS Measurement Errors

Overbounds of heavy-tailed measurement errors are essential to meet stringent navigation requirements in integrity monitoring applications. This paper proposes to leverage the bounding sharpness of the Cauchy distribution in the core and the Overbounds of heavy-tailed measurement errors are essential for meeting stringent navigation...

💬 0 commentsarXiv:2601.07299v2PDF
0

Posted in stat.ME · 2026-01-12 · Junjun Lang, Qiong Zhang, Yukun Liu

Minimum Wasserstein distance estimator under covariate shift: closed-form, super-efficiency and irregularity

Covariate shift arises when covariate distributions differ between source and target populations while the conditional distribution of the response remains invariant, and it underlies problems in missing data and causal inference. We propose a minimum Wasserstein distance estimation framework for inference under covariate shift that...

💬 0 commentsarXiv:2601.07282v1PDF
0

Posted in stat.ML · 2026-01-12 · Likun Zhang, Wei Ma

Covariance-Driven Regression Trees: Reducing Overfitting in CART

Decision trees are powerful machine learning algorithms, widely used in fields such as economics and medicine for their simplicity and interpretability. However, decision trees such as CART are prone to overfitting, especially when grown deep or the sample size is small. Conventional methods to reduce overfitting include pre-pruning...

💬 0 commentsarXiv:2601.07281v1PDF
0

Posted in stat.CO · 2026-01-12 · Alessandra Ragni, Lara Cavinato, Francesca Ieva

Penalized Likelihood Optimization for Adaptive Neighborhood Clustering in Time-to-Event Data with Group-Level Heterogeneity

The identification of patient subgroups with comparable event-risk dynamics plays a key role in supporting informed decision-making in clinical research. In such settings, it is important to account for the inherent dependence that arises when individuals are nested within higher-level units, such as hospitals. Existing survival...

💬 0 commentsarXiv:2601.07446v1PDF
0

Posted in stat.ML · 2026-01-12 · Niklas Kormann, Benjamin Doerr, Johannes F. Lutzeyer

Position: Don't be Afraid of Over-Smoothing And Over-Squashing

Over-smoothing and over-squashing have been extensively studied in the literature on Graph Neural Networks (GNNs) over the past years. We challenge this prevailing focus in GNN research, arguing that these phenomena are less critical for practical applications than assumed. We suggest that performance decreases often stem from...

💬 0 commentsarXiv:2601.07419v1PDF
0

Posted in stat.ME · 2026-01-12 · Wai Leong Ng, Xinyi Tang, Mun Lau Cheung, Jiacheng Gao, Chun Yip Yau, Holger Dette

Inference for Multiple Change-points in Piecewise Locally Stationary Time Series

Change-point detection and locally stationary time series modeling are two major approaches for the analysis of non-stationary data. The former aims to identify stationary phases by detecting abrupt changes in the dynamics of a time series model, while the latter employs (locally) time-varying models to describe smooth changes in...

💬 0 commentsarXiv:2601.07400v2PDF