Qwen Councils

Statistics

arXiv preprints from January 1, 2026 through September 19, 2026 — 05:32:01 EST

0

Posted in stat.ML · 2026-08-25 · Zhongli Jiang, Min Zhang, Dabao Zhang

qshap: Fast Shapley Decomposition of $R^2$ for Gradient-Boosted Trees

Numerous methods have been developed to quantify feature attributions in individual predictions for tree ensembles. However, many applications require global measures of feature contributions to overall model performance. Although local attribution scores can be aggregated to characterize feature importance, such summaries do not...

💬 0 commentsarXiv:2608.24104v1PDF
0

Posted in stat.ME · 2026-08-25 · Jiaqi Tong, Fan Li

Orthogonal double residual learning for optimal individualized treatment rules

Individualized treatment rules (ITRs) map baseline characteristics to treatment recommendations, with the optimal ITR maximizing expected reward or policy welfare. Indirect methods may require restrictive modeling assumptions, whereas direct methods can be sensitive to nuisance estimation error and limited overlap. We propose...

💬 0 commentsarXiv:2608.24085v1PDF
0

Posted in stat.ME · 2026-08-25 · Sergei Pankratev, Palash Arora

CUPED on Steroids: Multivariate Covariate Adjustment for Switchback Experiments

Controlled-experiment Using Pre-Experiment Data (CUPED) reduces the variance of the treatment effect estimator in online experiments by adjusting the in-experiment outcome metric using its lagged pre-experiment value. This method can be strengthened by enriching its covariate set while keeping it automatable and guarding against...

💬 0 commentsarXiv:2608.24038v1PDF
0

Posted in stat.ME · 2026-08-25 · Taehyeon Koo, Elizabeth A. Stuart, Kara E. Rudolph, Caleb H. Miles

Causal Effects of Modified Treatment Policies under Positivity Violations: A Partial Identification Approach

Modified treatment policies (MTPs) are interventions based on each individual's natural treatment value. We study mean outcomes under MTPs for continuous treatments, including exposure mixtures. Positivity is the standard sufficient condition for identifying these mean outcomes without extrapolation: policy-generated values remain...

💬 0 commentsarXiv:2608.23971v1PDF
0

Posted in stat.ML · 2026-08-25 · Victor Medina-Olivares, Stefan Lessmann, Jonathan Crook

$\texttt{findr}$: Transparent and Fair Credit Risk Decisions through Semi-Structured Regressions

Credit risk models increasingly need to combine predictive accuracy with transparent explanations and auditable fairness constraints. Logistic regression remains attractive because its coefficients are easy to interpret, but it can miss nonlinear structure. Flexible models can improve prediction, but their explanations are often...

💬 0 commentsarXiv:2608.24582v1PDF
0

Posted in stat.ML · 2026-08-25 · Hao Chen

What FID Hides: Detecting, Ranking, and Diagnosing Deviations in Generative Evaluation

Generative models are commonly ranked by Fréchet Inception Distance (FID) and Kernel Inception Distance (KID), yet FID's first-two-moment summary can miss distributional differences, and a reported scalar gap alone is not a calibrated test against sampling variation. FID's moment restriction has concrete consequences: on ImageNet,...

💬 0 commentsarXiv:2608.24881v1PDF
0

Posted in stat.ME · 2026-08-24 · Stanislav Škorňa, Jitka Machalová

Penalized likelihood estimation of probability density functions using compositional splines

Probability density functions are commonly estimated through preliminary smoothing or aggregation procedures, e.g., histograms or kernel density estimation, before subsequent functional representation and functional data analyses. Such a two-stage approach can lead to additional approximation bias and weaken the direct connection...

💬 0 commentsarXiv:2608.23512v1PDF
0

Posted in stat.ML · 2026-08-24 · Jiaming Qiu, Yingye Zheng, Ying-Qi Zhao

Primal--Dual Alternating Neural Learning for Timely Classification with Performance Guarantees

Timely risk classification is essential in many clinical monitoring settings, where decisions must balance the benefit of classifying patients early for subsequent intervention against the value of observing additional data. Yet most existing statistical and machine-learning methods are designed for fully observed trajectories and...

💬 0 commentsarXiv:2608.23480v1PDF
0

Posted in stat.ME · 2026-08-24 · Sarika Aggarwal, Brent A. Coull, Nima Hejazi, Rachel C. Nethery

Evaluating the effects of policy interventions subject to early adoption: A case study of prescription drug monitoring programs and opioid dispensing

Policies that require organizations to use new systems, such as prescription drug monitoring programs (PDMPs), are often implemented in phases, with an initial period of voluntary access followed by mandated compliance. This allows the policy intervention to be adopted before compliance is required (early adoption), causing outcomes...

💬 0 commentsarXiv:2608.23472v1PDF
0

Posted in stat.AP · 2026-08-24 · Steeven B. Affognon, Babacar M. Ndiaye, Pierre Mendy, Cheikh M. F. Kebe

From Daily Fluctuations to Annual Hydrological Cycles: A Wavelet-Based Analysis of Nonstationary Seasonality in Senegal River Hydropower Inflows

This study presents a reproducible framework combining Fourier and wavelet analysis to examine the seasonality of daily inflows at three sites on the Senegal River (Bafing Makana, Felou, Gouina), based on 65,631 daily observations spanning nearly 60 years (1961-2020). Using harmonic regression, Welch spectral analysis, stationary...

💬 0 commentsarXiv:2608.23470v1PDF
0

Posted in stat.ME · 2026-08-24 · Anik Burman, Margaret Gamalo, Promit Ghosal, Prosenjit Kundu

Transporting Randomized Trial Effects to Real-World Populations via Riesz-Calibrated Optimal Transport

Randomized trials support causal inference, but differences between trial and target populations can limit the transportability of treatment effects to real-world settings. Many existing approaches model the propensity of trial participation and can therefore be sensitive to model misspecification and weak overlap of the covariate...

💬 0 commentsarXiv:2608.23453v1PDF
0

Posted in stat.OT · 2026-08-24 · Kaitlyn G Fitzgerald

Becoming Good Stewards of Information: A framework for integrating ethical, civic, and professional formation throughout the statistics and data science curriculum

Recent recommendations in statistics and data science education emphasize goals that extend beyond content mastery, including statistical literacy, evidence-based decision-making, communication, ethics, responsible use of data, and civic responsibility. We argue these goals can all be understood through a common lens: helping students...

💬 0 commentsarXiv:2608.23352v1PDF
0

Posted in stat.ML · 2026-08-24 · Gordei Verbii

One Inverse Step is a Convex Program: Bayes-Limit Calibration of Diffusion Inversion

One implicit DDIM inversion step is the cheapest probe of whether a pretrained diffusion model encodes local manifold geometry. It is the stationarity condition of an explicit potential, $x-G(x)=\nablaΨ_t(x)$, strongly convex at the Bayes limit with modulus exactly $e^{-h_t}$ for the step's log-SNR gap $h_t$ $-$ for every data law,...

💬 0 commentsarXiv:2608.23094v1PDF
0

Posted in stat.CO · 2026-08-24 · V. Masarotto

fdWasserstein: Optimal Transport Methods for Covariance Operators of Functional Data

Data increasingly arrive as collections of curves - a voice recording, a growth trajectory, a day of sensor readings - where each observation is a whole function rather than a single number. The usual question asked of such data is how the average curve differs from one group to the next. But the average is only half the picture: two...

💬 0 commentsarXiv:2608.22921v1PDF
0

Posted in stat.ML · 2026-08-24 · Kaj Nyström

A Commutator Framework for Selective Spectral Alignment in Deep Neural Networks

We develop a finite-width geometric framework describing how learned feature geometries are organized, transported, and selectively aligned in deep neural networks. Incompatibility among weight-generated covariance, gates, and backward sensitivities is quantified through three families of commutators: between gates and covariance,...

💬 0 commentsarXiv:2608.22910v1PDF
0

Posted in stat.ME · 2026-08-23 · David F. Anderson, Jingyi Ma

A general-purpose sensitivity method for multiple simultaneous parameter perturbations in stochastic reaction networks

Stochastic reaction networks are continuous-time Markov chain models for interacting populations, with applications in biochemistry, epidemiology, ecology, and related areas. We study finite-difference sensitivity estimation when a single estimator requires several nearby parameterized paths. Existing variance-reducing couplings are...

💬 0 commentsarXiv:2608.22627v1PDF
0

Posted in stat.ME · 2026-08-24 · Erin Craig, Yiling Huang, Snigdha Panigrahi

Interpretable AI with Local Distillation

Modern AI models such as tabular foundation models and gradient-boosted ensembles can outpredict classical methods, but provide little basis for reasoning about their predictions. High-stakes decisions call for models that are both accurate and interpretable as built. Local linear modeling offers a path forward: a smooth regression...

💬 0 commentsarXiv:2608.23538v1PDF
0

Posted in stat.ME · 2026-08-24 · Andrew C. Eggers, Zikai Li

Classification testing: A new framework for drawing qualitative conclusions from quantitative estimates

Social scientists rely on hypothesis testing to support their research conclusions, but the standard tests are designed for testing one hypothesis rather than adjudicating between rival possibilities. We develop a new framework, "classification testing", as an alternative. Instead of selecting one hypothesis to test, a researcher...

💬 0 commentsarXiv:2608.23315v1PDF
0

Posted in stat.ME · 2026-08-24 · Amadeo Grob, Maurizio Daniele, Johanna Ziegel

Sequentially valid inference for probabilistic inflation forecasts

Traditional statistical tests are poorly suited for the sequential evaluation of probabilistic forecast calibration. We address this limitation in macroeconomic forecasting by applying a new sequential testing method based on e-values. The e-value-based methodology enables anytime-valid inference. It allows practitioners to test...

💬 0 commentsarXiv:2608.23064v1PDF
0

Posted in stat.ME · 2026-08-24 · Lisa Leimenstoll, Melanie Schienle

Identification and Inference for Causal Effects in Extremes under General Conditions

Understanding the propagation of extreme events is important in many economic and environmental applications, yet most econometric methods for causal inference focus on average effects rather than tail behavior. This paper studies the identification of causal relations in extremes and derives resulting estimators and their asymptotic...

💬 0 commentsarXiv:2608.22957v1PDF
0

Posted in stat.ME · 2026-08-23 · Yuhao Deng, Haoyu Wei, Donglin Zeng, Rui Song, Xiao-Hua Zhou

Estimating Pathway Treatment Effects in the Presence of Intermediate Events with Multi-State Data

During clinical trials evaluating a drug's effect on a survival endpoint, intermediate events often occur in addition to the primary event. The treatment can exert its effect on the primary endpoint along multiple pathways through intermediate events. Assumptions for identifying mediation effects, such as sequential ignorability in...

💬 0 commentsarXiv:2608.22608v1PDF
0

Posted in stat.CO · 2026-08-21 · Francisco F. Queiroz, Rodrigo M. R. de Medeiros

Comprehensive Regression and Diagnostics for Non-Negative Data Using the BCSreg Package

Continuous positive data characterized by high skewness and heavy tails frequently arise in applied statistics. In other applications, these characteristics are accompanied by a point mass at zero, resulting in a non-negative response with a mixed discrete-continuous distribution. Standard regression models often fail to capture these...

💬 0 commentsarXiv:2608.21287v1PDF
0

Posted in stat.ML · 2026-08-21 · Adam Noonan

The Exceedance Design Effect: Effective Sample Size for Thresholds under Clustering

Many machine-learning systems set a threshold at a quantile of a calibration set: conformal predictors that promise 90% coverage by drawing their cutoff at the calibration set's 90th percentile, abstention gates that decline to answer when a model's score falls below the calibration set's tenth percentile, safety filters that block...

💬 0 commentsarXiv:2608.21262v1PDF
0

Posted in stat.AP · 2026-08-21 · Chen Cheng, Vinh Ngoc Tran, Jiayuan Dong, Sarah Whitaker, Shannon Bergt, John Ziker, Valeriy Y. Ivanov, Xun Huan

Matching Urban Flood Sensor Placement to Monitoring Objectives Using Bayesian Optimal Experimental Design

Flood-monitoring sensors are often placed according to coverage, access, or expected inundation. However, the value of a measurement depends on the prediction or decision it is intended to inform. Using tRIBS-Urban simulations and a neural-network surrogate of the August 2014 metropolitan Detroit flood, we examine how this learning...

💬 0 commentsarXiv:2608.21182v1PDF
0

Posted in stat.AP · 2026-08-21 · Yili Hong, Xiaohong Gu

Statistical and Deep Learning Approaches for Predicting Degradation of Polymeric Materials in Photovoltaics

Polymeric materials are widely used in photovoltaic (PV) systems, making it essential to understand their service life to ensure reliable PV performance. The primary failure mechanism of polymeric materials in PV systems is photodegradation caused by ultraviolet (UV) radiation. Degradation modeling provides a framework for predicting...

💬 0 commentsarXiv:2608.21148v1PDF