Qwen Councils

Statistics

arXiv preprints from January 1, 2026 through September 21, 2026 — 21:34:35 EST

0

Posted in stat.ME · 2026-01-13 · Fredrik Lohne Aanes

Approximate Shapley value estimation using sampling without replacement and variance estimation via the new Symmetric bootstrap and the Doubled half bootstrap

In this paper I consider improving the KernelSHAP algorithm. I suggest to use the Wallenius' noncentral hypergeometric distribution for sampling the number of coalitions and perform sampling without replacement, so that the KernelSHAP estimation framework is improved further. I also introduce the Symmetric bootstrap to calculate the...

💬 0 commentsarXiv:2601.08981v1PDF
0

Posted in stat.ME · 2026-01-13 · Jiahao Tian, Hugh Chipman, Thomas Loughin

MLCBART: Multilabel Classification with Bayesian Additive Regression Trees

Multilabel Classification (MLC) deals with the simultaneous classification of multiple binary labels. The task is challenging because, not only may there be arbitrarily different and complex relationships between predictor variables and each label, but associations among labels may exist even after accounting for effects of predictor...

💬 0 commentsarXiv:2601.08964v1PDF
0

Posted in stat.ME · 2026-01-12 · Hisaya Okahara, Tomoyuki Nakagawa, Shonosuke Sugasawa

The Covariate-Assisted Bayesian Intransitive Bradley-Terry Model via Combinatorial Hodge Theory

Pairwise comparison data are widely used to recover latent rankings, yet the models in dominant use assume stochastic transitivity. When preferences are in fact intransitive, a single scalar strength conflates genuine hierarchy with cycle-induced structure, biasing both the recovered ranking and any covariate effects attributed to it....

💬 0 commentsarXiv:2601.07158v2PDF
0

Posted in stat.ML · 2026-01-12 · Linus Bleistein, Mathieu Dagréou, Francisco Andrade, Thomas Boudou, Aurélien Bellet

Optimal Transport under Group Fairness Constraints

Ensuring fairness in matching algorithms is a key challenge in allocating scarce resources and positions. Focusing on Optimal Transport (OT), we introduce a novel notion of group fairness requiring that the probability of matching two individuals from any two given groups in the OT plan satisfies a predefined target. We first propose...

💬 0 commentsarXiv:2601.07144v3PDF
0

Posted in stat.ME · 2026-01-12 · Shuli Chen, Jie Hu, Zhichao Jiang

Connections as treatment: causal inference with edge interventions in networks

Causal inference has traditionally focused on interventions at the unit level. In many applications, however, the central question concerns the causal effects of connections between units, such as transportation links, social relationships, or collaborative ties. We develop a causal framework for edge interventions in networks, where...

💬 0 commentsarXiv:2601.07267v1PDF
0

Posted in stat.ME · 2026-01-12 · Suchismita Das, Akul Ameya, Cahyani Karunia Putri

Compounded Linear Failure Rate Distribution: Properties, Simulation and Analysis

This paper proposes a new extension of the linear failure rate (LFR) model to better capture real-world lifetime data. The model incorporates an additional shape parameter to increase flexibility. It helps model the minimum survival time from a set of LFR distributed variables. We define the model, derive certain statistical...

💬 0 commentsarXiv:2601.07249v1PDF
0

Posted in stat.ML · 2026-01-12 · Yiran Jia, Jelena Bradic

Multi-environment Invariance Learning with Missing Data

Learning models that can handle distribution shifts is a key challenge in domain generalization. Invariance learning, an approach that focuses on identifying features invariant across environments, improves model generalization by capturing stable relationships, which may represent causal effects when the data distribution is encoded...

💬 0 commentsarXiv:2601.07247v2PDF
0

Posted in stat.ME · 2026-01-12 · Kanji Goto, Shintaro Yuki, Kensuke Tanioka, Hiroshi Yadohisa

Principal component-guided sparse reduced-rank regression

Reduced-rank regression estimates regression coefficients by imposing a low-rank constraint on the matrix of regression coefficients, thereby accounting for correlations among response variables. To further improve predictive accuracy and model interpretability, several regularized reduced-rank regression methods have been proposed....

💬 0 commentsarXiv:2601.07202v2PDF
0

Posted in stat.ML · 2026-01-12 · EL Mahdi Khribch, Pierre Alquier

Robust Bayesian Inference via Variational Approximations of Generalized Rho-Posteriors

We introduce the $\widetildeρ$-posterior, a modified version of the $ρ$-posterior, obtained by replacing the supremum over competitor parameters with a softmax aggregation. This modification allows a PAC-Bayesian analysis of the $\widetildeρ$-posterior. This yields finite-sample oracle inequalities with explicit convergence rates that...

💬 0 commentsarXiv:2601.07325v2PDF
0

Posted in stat.AP · 2026-01-12 · Zhengdao Li, Penggao Yan, Weisong Wen, Li-Ta Hsu

Cauchy-Gaussian Overbound for Heavy-tailed GNSS Measurement Errors

Overbounds of heavy-tailed measurement errors are essential to meet stringent navigation requirements in integrity monitoring applications. This paper proposes to leverage the bounding sharpness of the Cauchy distribution in the core and the Overbounds of heavy-tailed measurement errors are essential for meeting stringent navigation...

💬 0 commentsarXiv:2601.07299v2PDF
0

Posted in stat.ME · 2026-01-12 · Junjun Lang, Qiong Zhang, Yukun Liu

Minimum Wasserstein distance estimator under covariate shift: closed-form, super-efficiency and irregularity

Covariate shift arises when covariate distributions differ between source and target populations while the conditional distribution of the response remains invariant, and it underlies problems in missing data and causal inference. We propose a minimum Wasserstein distance estimation framework for inference under covariate shift that...

💬 0 commentsarXiv:2601.07282v1PDF
0

Posted in stat.ML · 2026-01-12 · Likun Zhang, Wei Ma

Covariance-Driven Regression Trees: Reducing Overfitting in CART

Decision trees are powerful machine learning algorithms, widely used in fields such as economics and medicine for their simplicity and interpretability. However, decision trees such as CART are prone to overfitting, especially when grown deep or the sample size is small. Conventional methods to reduce overfitting include pre-pruning...

💬 0 commentsarXiv:2601.07281v1PDF
0

Posted in stat.CO · 2026-01-12 · Alessandra Ragni, Lara Cavinato, Francesca Ieva

Penalized Likelihood Optimization for Adaptive Neighborhood Clustering in Time-to-Event Data with Group-Level Heterogeneity

The identification of patient subgroups with comparable event-risk dynamics plays a key role in supporting informed decision-making in clinical research. In such settings, it is important to account for the inherent dependence that arises when individuals are nested within higher-level units, such as hospitals. Existing survival...

💬 0 commentsarXiv:2601.07446v1PDF
0

Posted in stat.ML · 2026-01-12 · Niklas Kormann, Benjamin Doerr, Johannes F. Lutzeyer

Position: Don't be Afraid of Over-Smoothing And Over-Squashing

Over-smoothing and over-squashing have been extensively studied in the literature on Graph Neural Networks (GNNs) over the past years. We challenge this prevailing focus in GNN research, arguing that these phenomena are less critical for practical applications than assumed. We suggest that performance decreases often stem from...

💬 0 commentsarXiv:2601.07419v1PDF
0

Posted in stat.ME · 2026-01-12 · Wai Leong Ng, Xinyi Tang, Mun Lau Cheung, Jiacheng Gao, Chun Yip Yau, Holger Dette

Inference for Multiple Change-points in Piecewise Locally Stationary Time Series

Change-point detection and locally stationary time series modeling are two major approaches for the analysis of non-stationary data. The former aims to identify stationary phases by detecting abrupt changes in the dynamics of a time series model, while the latter employs (locally) time-varying models to describe smooth changes in...

💬 0 commentsarXiv:2601.07400v2PDF
0

Posted in stat.ME · 2026-01-12 · Roberto Fontana, Elisa Perrone, Fabio Rapallo

Characterization of multi-way binary tables with uniform margins and fixed correlations

In many applications involving binary variables, only pairwise dependence measures, such as correlations, are available. However, for multi-way tables involving more than two variables, these quantities do not uniquely determine the joint distribution, but instead define a family of admissible distributions that share the same...

💬 0 commentsarXiv:2601.07369v1PDF
0

Posted in stat.ME · 2026-01-12 · Ryo Okano, Daisuke Kurisu

Functional Synthetic Control Methods for Metric Space-Valued Outcomes

The synthetic control method (SCM) is a widely used tool for evaluating causal effects of policy changes in panel data settings. Recent studies have extended its framework to accommodate complex outcomes that take values in metric spaces, such as distributions, functions, networks, covariance matrices, and compositional data. However,...

💬 0 commentsarXiv:2601.07539v1PDF
0

Posted in stat.ML · 2026-01-12 · Victor Thuot, Sebastian Vogt, Debarghya Ghoshdastidar, Nicolas Verzelen

Nonparametric Kernel Clustering with Bandit Feedback

Clustering with bandit feedback refers to the problem of partitioning a set of items, where the clustering algorithm can sequentially query the items to receive noisy observations. The problem is formally posed as the task of partitioning the arms of an N-armed stochastic bandit according to their underlying distributions, grouping...

💬 0 commentsarXiv:2601.07535v1PDF
0

Posted in stat.AP · 2026-01-12 · Lampis Tzai, Ioannis Ntzoufras, Silvia Bozza

Bayesian Handwriting Evidence Evaluation using MANOVA via Fourier-Based Extracted Features

This paper proposes a novel statistical approach that aims at the identification of valid and useful patterns in handwriting examination via Bayesian modeling. Starting from a sample of characters selected among 13 French native writers, an accurate loop reconstruction can be achieved through Fourier analysis. The contour shape of...

💬 0 commentsarXiv:2601.07534v1PDF
0

Posted in stat.CO · 2026-01-12 · Nathan Green

Population-Adjusted Indirect Treatment Comparison with the outstandR Package in R

Indirect treatment comparisons (ITCs) are essential in Health Technology Assessment (HTA) when head-to-head clinical trials are absent. A common challenge arises when attempting to compare a treatment with available individual patient data (IPD) against a competitor with only reported aggregate-level data (ALD), particularly when...

💬 0 commentsarXiv:2601.07532v3PDF
0

Posted in stat.ML · 2026-01-12 · Hao Qiu, Mengxiao Zhang, Juliette Achddou

Decentralized Online Convex Optimization with Unknown Feedback Delays

Decentralized online convex optimization (D-OCO), where multiple agents within a network collaboratively learn optimal decisions in real-time, arises naturally in applications such as federated learning, sensor networks, and multi-agent control. In this paper, we study D-OCO under unknown, time-and agent-varying feedback delays....

💬 0 commentsarXiv:2601.07901v1PDF
0

Posted in stat.ME · 2026-01-12 · Miguel Martinez Herrera, Felix Cheysson

Ridge-penalised spectral least-squares estimation for point processes

Penalised estimation methods for point processes usually rely on a large amount of independent repetitions for cross-validation purposes. However, in the case of a single realisation of the process, existing cross-validation methods may be impractical depending on the chosen model. To overcome this issue, this paper presents a...

💬 0 commentsarXiv:2601.07490v1PDF
0

Posted in stat.ME · 2026-01-12 · Marco Alfo', Roberto Rocci

Omitted covariates bias and finite mixtures of regression models for longitudinal responses

Individual-specific, time-constant, random effects are often used to model dependence and/or to account for omitted covariates in regression models for longitudinal responses. Longitudinal studies have known a huge and widespread use in the last few years as they allow to distinguish between so-called age and cohort effects; these...

💬 0 commentsarXiv:2601.07609v1PDF
0

Posted in stat.AP · 2026-01-12 · Gianna Gavriel, Maria Pregnolato, Francesca Pianosi, Theo Tryfonas, Paul Vardanega

An evaluation of empirical equations for assessing local scour around bridge piers using global sensitivity analysis

Bridge scour is a complex phenomenon combining hydrological, geotechnical and structural processes. Bridge scour is the leading cause of bridge collapse, which can bring catastrophic consequences including the loss of life. Estimating scour on bridges is an important task for engineers assessing bridge system performance....

💬 0 commentsarXiv:2601.07594v1PDF
0

Posted in stat.ME · 2026-01-12 · Joseph Lam, Mario Cortina-Borja, Rob Aldridge, Ruth Blackburn, Katie Harron

Cluster-based name embeddings reduce ethnic disparities in record linkage quality under realistic name corruption: evidence from the North Carolina Voter Registry

Differential ethnic-based record linkage errors can bias epidemiologic estimates. Prior evidence often conflates heterogeneity in error mechanisms with unequal exposure to error. Using snapshots of the North Carolina Voter Registry (Oct 2011-Oct 2022), we derived empirical name-discrepancy profiles to parameterise realistic...

💬 0 commentsarXiv:2601.07693v1PDF