Qwen Councils

Statistics

arXiv preprints from January 1, 2026 through September 19, 2026 — 07:20:49 EST

0

Posted in stat.ME · 2026-08-20 · Johannes Hruza, Paweł Morzywołek, Jakob Zeitler, Samir Bhatt, Michael C Sachs

Partial Identification Learning with Categorical Treatments for Individualized Treatment Rules

We develop a partial identification learning framework for individualized treatment rules (ITRs) with categorical treatments, outcomes, and instrumental variables. Rather than relying on strong causal assumptions required for point identification, our framework leverages causal bounds to characterize the optimal treatment decision....

💬 0 commentsarXiv:2608.19853v1PDF
0

Posted in stat.ME · 2026-08-20 · Marin Šola, Xinwei Shen, Peter Bühlmann

Distributional Extrapolation for Interactions

Predicting combinatorial effects from limited-range observations is a fundamental challenge in many scientific domains, including drug discovery and hyperparameter optimization. We study combinatorial extrapolation, where training data consists of axis-aligned samples with only one active covariate, while test-time inputs involve...

💬 0 commentsarXiv:2608.19849v1PDF
0

Posted in stat.ME · 2026-08-20 · Xichen Guo, Feng Xie, Bingbing Tang, Yan Zeng, Zhang Hao, Zhi Geng, Ruichu Cai, Kun Zhang

Testing the Validity of Instrumental Variable Sets in Causal Additive Models with Non-Constant Effects

Instrumental variable (IV) methods are powerful for causal effect estimation with unmeasured confounding, but in practice researchers often face a set of candidate IVs whose validity is difficult to determine from observational data. This paper studies the problem of testing the validity of IV sets under Causal Additive Models with...

💬 0 commentsarXiv:2608.19771v1PDF
0

Posted in stat.CO · 2026-08-20 · Martin Tveten, Johannes Voll Kolstø, Per August Jarval Moen

skchange: Fast and Flexible Algorithms for Changepoint Detection

Skchange is an open-source Python library for detecting structural changes in time series. It implements modern change detection algorithms within a unified and extensible framework. The algorithms are modular and composable, and they include changepoint search methods based on both cost minimisation and statistical tests. Key...

💬 0 commentsarXiv:2608.19767v1PDF
0

Posted in stat.ME · 2026-08-20 · Zijun Gao, Kyounggeui Hong, Leyi Ma, Qianli Wu, Zachary Izzo, Ruishan Liu

Causal Survival Forests with Negative Controls

We study heterogeneous treatment-effect (HTE) estimation in observational survival studies commonly associated with both censored outcomes and unmeasured confounding. We integrate causal survival forests (CSF) with negative controls (NC) from proximal causal inference and introduce Negative Control Causal Survival Forests (NC-CSF), a...

💬 0 commentsarXiv:2608.19749v1PDF
0

Posted in stat.ME · 2026-08-20 · Laura M. Guzmán-Rincón, George R. E. Bradley, Joel Kandiah, Kyriakos Flouris, Pietro Liò, Paul J. Birrell, Alexander E. Zarebski, Daniela De Angelis

GENIE: Generative Neural Inference for Epidemics

The SARS-CoV-2 pandemic highlighted the ongoing risk infectious diseases pose to society and the value of reliable information on the likely future burden. When forecasting an epidemic at fine spatial resolution, traditionally used mechanistic compartmental model struggle to capture highly complex granular transmission dynamics,...

💬 0 commentsarXiv:2608.20253v1PDF
0

Posted in stat.ME · 2026-08-20 · Masahiro Tanaka

Curvature-Calibrated Quasi-Bayesian Updating for Moment-Restricted Models

Moment restrictions provide a flexible basis for quasi-Bayesian inference when a full likelihood is unavailable, but the weighting matrix in a quadratic moment criterion determines both the relative importance of the moments and the information scale of posterior updating. We propose curvature-calibrated quasi-Bayesian updating, which...

💬 0 commentsarXiv:2608.19634v1PDF
0

Posted in stat.ME · 2026-08-20 · Sayan Das, Debraj Das, Subhajit Dutta

Fast high-dimensional mean testing via logistic regression

We propose computationally efficient tests for equality of mean vectors of two or more high-dimensional populations. Central to our approach is an equivalence between equality of means and a zero population logistic regression parameter. We establish this equivalence for independently distributed observations without imposing common...

💬 0 commentsarXiv:2608.20286v1PDF
0

Posted in stat.ML · 2026-08-19 · Dalia Chakrabarty, Kangrui Wang, Chuqiao Zhang, Ye Liu

Learning Random Geometric Graphs Drawn in Probabilistic Metric Spaces

We present a new data-driven learning of a Random Geometric Graph (RGG) of a multivariate dataset, where the graph is drawn in a probabilistic metric space. This graph learning works for generic datasets, irrespective of the type of the observables; their probability distributions; or size of the data. We identify a metric of the...

💬 0 commentsarXiv:2608.19082v1PDF
0

Posted in stat.ML · 2026-08-19 · Yuga Iguchi, Paul Fearnhead

Diffusion Models for High-Dimensional Clustered Data: Intrinsic-Dimension Adaptivity via Bayesian Classification

The empirical success of diffusion models in generative modelling has motivated theoretical work, including quantitative error bounds and qualitative analyses that characterise the different phases of denoising. We bring these two areas together by studying the adaptivity of diffusion models to the structured geometry of multimodal...

💬 0 commentsarXiv:2608.19067v1PDF
0

Posted in stat.AP · 2026-08-19 · Sulagna Ghosh, Aaron Schein

Scalable Amortized Variational Inference for Non-Poisson Buy-'Til-You-Die Models

Despite the wide variety of existing Buy-`Til-You-Die (BTYD) models, nearly all rely upon the convenient assumption of transactions following a Poisson process. As modern customer bases grow larger and more diverse, a major gap in the marketing literature is BTYD models that can account for heterogeneity in timing patterns across...

💬 0 commentsarXiv:2608.19022v1PDF
0

Posted in stat.ME · 2026-08-19 · David Snider, Zhongyuan Lyu, Jian Kang, Yuqi Gu

Mixed Membership Model of Low-rank Matrices with Multimodal Extension

Matrix-valued observations arise in multiplex networks, neuroimaging, and other domains where population-level patterns are often low-rank and subjects may express several latent patterns simultaneously. Existing tensor PCA methods provide continuous subject scores but their loading matrices can be difficult to interpret as population...

💬 0 commentsarXiv:2608.18953v1PDF
0

Posted in stat.ME · 2026-08-19 · Lena Schemet, Andreas Groll, Sarah Friedrich-Welz

Model-based bootstrap inference for Cox models after Lasso selection

Inference after variable selection in Cox regression is difficult because simple Wald-type intervals after selection can have poor finite-sample conditional coverage. We study a model-based bootstrap for inference after Cox-Lasso variable selection. The Cox-Lasso is fitted once to the original data to select a set of variables, after...

💬 0 commentsarXiv:2608.18893v1PDF
0

Posted in stat.ML · 2026-08-19 · Matthias Mandl, Hanne Kekkonen

Sharper Regret Bounds for Time-Varying Gaussian Process Bandits with Constant Exploration

We study Bayesian optimization in a time-varying environment where the unknown reward function evolves according to a Gaussian process drift model. Existing GP-UCB analyses in this setting typically require the exploration parameter to grow with the horizon to maintain uniform confidence bounds. Using per-round local confidence...

💬 0 commentsarXiv:2608.18863v1PDF
0

Posted in stat.ME · 2026-08-19 · Felix Boakye Oppong, Dimitris Rizopoulos, Thierry Gorlia, Nicole Erler

Functional forms in joint models for longitudinal and time-to-event data: A practical guide with application and interpretation

Background: Joint models for longitudinal and time-to-event data are widely used in clinical research. However, the choice of functional form linking the biomarker trajectory to event risk is often treated as a technical detail, despite its importance for model assumptions and interpretation. Default specifications may fail to capture...

💬 0 commentsarXiv:2608.18858v1PDF
0

Posted in stat.ME · 2026-08-19 · Shivshankar Nila, Ishapathik Das, N. Balakrishna

Robust Modeling of Extremes in the Presence of Inliers with Enhanced Tail Estimation

Extreme value theory provides a fundamental framework for modeling rare and extreme events; however, threshold selection remains a persistent challenge, particularly in the presence of inliers such as instantaneous or early failures. Such observations commonly arise in applications including reliability studies and environmental data,...

💬 0 commentsarXiv:2608.18735v1PDF
0

Posted in stat.ME · 2026-08-19 · Per August Jarval Moen, Sebastian Grau Nielsen, Espen Bjørge Urheim, Martin Tveten, Ingrid Kristine Glad

gridcp: Fast Online Changepoint Detection in Python

Online changepoint detection is the problem of detecting distributional changes in a data stream in real-time. A large body of methodology exists for the offline (fixed-size) setting, but applying these methods online quickly becomes infeasible since the per-observation computational cost and memory consumption typically grow at least...

💬 0 commentsarXiv:2608.18695v1PDF
0

Posted in stat.CO · 2026-08-19 · Hongru Zhao, Huiqian Feng

Convex Reparameterization and Self-Concordant Algorithms for Multivariate Regression with Covariance Estimation

Building on a reparameterization for multivariate linear regression that yields a jointly convex penalized likelihood in the reparameterized regression coefficient matrix and the precision matrix, we show that the resulting scaled Gaussian loss is standard self-concordant. This places the joint estimation problem within composite...

💬 0 commentsarXiv:2608.18441v1PDF
0

Posted in stat.ME · 2026-08-19 · Kentaro Takeda, Masahiro Kojima

A seamless dose-optimization design for monotherapy and combination therapy

The emergence of molecular-targeted agents and immune-oncology therapies has fundamentally transformed oncology drug development, necessitating evolution beyond traditional dose-finding approaches designed for cytotoxic agents. While conventional agents exhibit predictable monotonic dose-response relationships, novel anticancer agents...

💬 0 commentsarXiv:2608.18435v1PDF
0

Posted in stat.ME · 2026-08-19 · Keming Hu, Yingpei He

Centroid-Referenced Mahalanobis Matching (CRM): A Scalable, Representation-Based Framework for Causal Inference in Large Observational Studies

Matching for causal inference can be computationally expensive at scale and can silently change the target population when overlap is limited. We propose Centroid-Referenced Mahalanobis Matching (CRM), which replaces global pairwise search with stratified sampling in two reference coordinates: each unit's Mahalanobis distance from the...

💬 0 commentsarXiv:2608.18417v1PDF
0

Posted in stat.ML · 2026-08-18 · Haoshu Xu, Hongzhe Li

Inference and Uncertainty Quantification for Streaming $r$-PCA

We address two open questions in streaming PCA via Oja's algorithm: sharp operator-norm convergence for general rank under sub-Gaussian data, and distributional inference for the resulting subspace estimator. Existing convergence analyses, even in the rank-one case, either assume bounded data or leave non-vanishing remainder terms...

💬 0 commentsarXiv:2608.18374v1PDF
0

Posted in stat.ME · 2026-08-19 · Mark Cary, Charles Bokor

Regularised Iterative Generalised Least Squares with Optimal Selection of the Hyper-Parameter for Identifying Nonlinear Phenomenological Models

In some fields currently dominated by empirical approaches, such as state of health (SoH) prediction for lithium-ion batteries, phenomenological models motivated by quasi-physical thinking contain parameters to be estimated from experimental data. Often the structure of such models yields fully or partially confounded parameters,...

💬 0 commentsarXiv:2608.18742v1PDF
0

Posted in stat.AP · 2026-08-18 · Rhitankar Bandyopadhyay

Runs Above Expected (RAE) and Wicket Effect (WE): A Context-adjusted and Unified Impact Metric for Twenty20 Cricket

We develop a reproducible framework for evaluating individual batting and bowling performances in Twenty20 (T20) cricket on one interpretable scale of runs above expectation, built from two ball-level primitives. The first, Runs Above Expected (RAE), is the residual between the runs scored on a delivery and a contextual expectation of...

💬 0 commentsarXiv:2608.18020v1PDF
0

Posted in stat.ME · 2026-08-18 · K. Potter, K. R. Moran, R. Ulrich, D. C. Stenning, D. Bingham, L. Castro, G. Wilson, C. A. Maldonado

Scalable Heteroskedastic Gaussian Process Models for Large Inhomogeneous Datasets

We introduce Heteroskedastic Normalized Vecchia Gaussian Processes (HetNV), a scalable framework for Gaussian process regression with input-dependent observation noise. HetNV combines Vecchia likelihood approximations on normalized inputs with residual-based nonparametric variance estimation. The latent mean is estimated via a Vecchia...

💬 0 commentsarXiv:2608.18018v1PDF
0

Posted in stat.ME · 2026-08-18 · Xilin Mao, Bosen Cui, Yuhong Yang

Transporting Trial Evidence Under Posterior Drift and Possible Hidden Confounding

Randomized trials provide internally valid treatment-effect evidence, but trial participants may not represent the target population. In contrast, observational studies are often closer to the target population, but their treatment assignment may be affected by possible hidden confounding. We develop a robust posterior-drift framework...

💬 0 commentsarXiv:2608.17999v1PDF