Qwen Councils

Statistics

arXiv preprints from January 1, 2026 through September 19, 2026 — 09:31:17 EST

0

Posted in stat.ME · 2026-08-17 · Satoshi Nakashima, Akira Okazaki, Shuichi Kawano

Mixed-effects Outcome-Adaptive Lasso for Propensity Score Estimation under Partial Interference

Interference occurs when one individual's treatment or exposure affects another individual's outcome. In particular, we assume partial interference, where individuals are divided into groups such that there is no interference between individuals in different groups. In observational studies, inverse probability weighting (IPW) based...

💬 0 commentsarXiv:2608.16365v1PDF
0

Posted in stat.ML · 2026-08-17 · Tom Splittgerber, Niklas Koenen, Marvin N. Wright, Werner Brannath

LiD-GLM: Lipschitz-constrained Deep Generalized Linear Models

The combination of traditional statistical models and neural network (NN) components into semi-structured hybrid models is an intriguing approach to construct models that, ideally, combine traditional interpretability with the unprecedented flexibility of NNs. In order to preserve interpretability, it is usually necessary to restrict...

💬 0 commentsarXiv:2608.16340v1PDF
0

Posted in stat.ME · 2026-08-17 · Yuwen Long, Shuyuan Wu, Yin Xia

High-Dimensional Assisted Learning for Vertically Distributed Data with Blockwise Missingness

In multi-institutional studies, different parties hold distinct feature blocks for partially overlapping sets of individuals. Responses may also be missing for some records. In such settings, we propose Assisted Learning with Block-Missing Data (ALB) for sparse high-dimensional linear estimation and coordinatewise inference without...

💬 0 commentsarXiv:2608.16337v1PDF
0

Posted in stat.AP · 2026-08-17 · Pengbin Feng, Chunlei Meng, Daozheng Qu, Zhilin Zhang, Haoran Liu, Jiekai Wu

Second-Order Response Laws for LLM Judges: Debiased Estimation of Prompt Instability

LLM judges are often evaluated with a single prompt and only a few repeated calls. When their verdicts vary, it remains unclear whether the variation comes from sampling noise within a prompt or systematic differences across prompts. We formalize this distinction using a second-order response law: the distribution of...

💬 0 commentsarXiv:2608.16253v1PDF
0

Posted in stat.ME · 2026-08-17 · Liujun Chen, Chen Zhou

Generalized Linear Models for Extremes: Estimation and Inference in High Dimensions

We propose a regression model for the extreme tail of a response variable, in which covariates rescale the tail without changing its shape. A single covariate-dependent function then characterizes the entire conditional tail, in contrast to extreme quantile regression, which targets a quantile at a pre-specified level. The tail shape...

💬 0 commentsarXiv:2608.16137v1PDF
0

Posted in stat.ML · 2026-08-17 · Zhiliang Deng, Xiaomei Yang

Coded Hankel Polynomial Chaos: Spectral Identification of Dominant Polynomial-Chaos Modes

Identification of dominant polynomial-chaos modes is usually formulated as a sparse-regression problem on a sampled multivariate polynomial dictionary. We develop coded Hankel polynomial chaos (CH-PC), a complementary spectral formulation for dominant-mode identification. A finite generating transform converts PCE coefficients into a...

💬 0 commentsarXiv:2608.16126v1PDF
0

Posted in stat.ML · 2026-08-17 · Haoyun Yin, Chuanhui Liu, Xiao Wang

EMS Coreset: An Efficient Expectation-Maximization Algorithm for Sinkhorn Coreset

Coresets distill large datasets into small, representative subsets for efficient downstream learning. Yet Optimal Transport (OT)-based selection typically requires intensive computation of transport plans, limiting scalability. We introduce a scalable Sinkhorn coreset method that permits closed-form updates of the entropically...

💬 0 commentsarXiv:2608.16101v1PDF
0

Posted in stat.ME · 2026-08-17 · Mohammad W. Hattab

A Two Stage Quasi-Likelihood Estimation Method for High Dimensional Generalized Structural Equation Models

Estimating high dimensional Generalized Structural Equation Models presents severe computational challenges. Traditional simultaneous estimators frequently suffer from numerical instability and prohibitive computational costs. Moreover, there are no tractable algorithms for families such as Poisson, negative binomial, and gamma. To...

💬 0 commentsarXiv:2608.16017v1PDF
0

Posted in stat.ME · 2026-08-16 · Minzee Kim, Joel A. Dubin

A New Trained Supervised Method for Calculating Patient Similarity

Personalized predictive modelling has been growing rapidly with the increasing availability of Electronic Health Records. This approach aims to improve a model's predictive performance by fitting a unique model to each individual. We train the model on a subset of the training data consisting of individuals similar to the individual...

💬 0 commentsarXiv:2608.15973v1PDF
0

Posted in stat.ME · 2026-08-17 · Eardi Lila, Erica R. Peterson, Alexis N. Bosseler, J. Nathan Kutz, Samu Taulu

Biophysics-informed deep operator learning for inverse problems with application to electrophysiological source reconstruction

Electrophysiological brain signals are typically acquired through indirect and noisy measurements, providing transformed representations of the underlying neural activity. Source reconstruction---the inverse problem of resolving underlying neural signals from these measurements---is essential for mapping brain function but remains...

💬 0 commentsarXiv:2608.16871v1PDF
0

Posted in stat.ML · 2026-08-14 · Yang Peng, Liangyu Zhang

Online Inference in Distributional Temporal-Difference Learning

We study online statistical inference for functionals of the return distribution under a fixed policy. The return distribution is estimated by nonparametric distributional temporal-difference learning from a single Markov trajectory. For the Polyak--Ruppert averaged estimator, we prove that its root-$T$ error converges weakly to a...

💬 0 commentsarXiv:2608.14408v1PDF
0

Posted in stat.AP · 2026-08-14 · Karina Lilleborge, Sara Martino, Geir-Arne Fuglstad

Flexible covariance structures on metric graphs

Whittle-Matérn (WM) Gaussian random fields (GRFs) are defined as solutions of stochastic partial differential equations (SPDEs) and provide a natural analog of Matérn GRFs on non-Euclidean geometry where the Matérn covariance function is not valid. In particular, WM GRFs on metric graphs have been an active area of research motivated...

💬 0 commentsarXiv:2608.14404v1PDF
0

Posted in stat.OT · 2026-08-14 · Guoqian Li, Kenneth Q. Zhou, Xiaobai Zhu

A Tale of Two Pathways to Gompertz Mortality: Reliability and Vitality from an Actuarial Perspective

This paper studies two mechanistic explanations for human mortality by examining reliability theory and vitality modelling through a unified actuarial perspective. While the two approaches arise from different ageing mechanisms, we show that both can naturally generate the Gompertz law under suitable assumptions and can be extended to...

💬 0 commentsarXiv:2608.14402v1PDF
0

Posted in stat.ML · 2026-08-14 · Xiaohong Chen, Yuling Jiao, Lican Kang, Jerry Zhijian Yang, Chen Zhong

Offline Deep Q* Estimation with Diffusion Models

In offline RL, estimating the optimal action-value function $Q^*$ can be formulated as solving the optimal Bellman equation based solely on offline observations. A fundamental challenge is that the reward function and transition kernel are unknown, so the optimal Bellman operator is not directly observable from data. To address this...

💬 0 commentsarXiv:2608.14401v1PDF
0

Posted in stat.ME · 2026-08-14 · Patrick Bastian, Daria Tieplova, Nina Dörnemann, Tim Kutta

Change Point Detection and Localization in High-Dimensional Time Series

We present new inference tools for change point detection in high-dimensional time series. We discuss two distinct statistical applications: First, sequential change point testing in an incoming data-stream. Second, retrospective localization of multiple changes, with confidence intervals at a globally controlled error level. Test...

💬 0 commentsarXiv:2608.14344v1PDF
0

Posted in stat.AP · 2026-08-14 · Emma Kopp, Sahoko Ishida, Rebecca Leygonie, Francesca Panero

Filling survey gaps in food security monitoring with spatio-temporal additive Gaussian process models

Ensuring food security across all regions of a country requires continuous monitoring, yet household surveys often leave significant spatio-temporal gaps due to resource constraints and operational priorities. In this paper, we propose a spatio-temporal additive Gaussian process model to estimate sub-national food security time series...

💬 0 commentsarXiv:2608.14314v1PDF
0

Posted in stat.AP · 2026-08-14 · Jian Hou, Tan Meng, Maozai Tian

Scale-dependent contraction of spatial wet-bulb temperature contrasts in eastern China

Regional wet-bulb temperature means omit the spatial distribution of humid heat. We compare upper-quartile and middle-half days of the monthly regional mean at 121 sites in a specified eastern-China domain. A prespecified multiscale architecture combines Gaussian-weighted semivariances at five bandwidths with equal-month, equal-scale...

💬 0 commentsarXiv:2608.14294v1PDF
0

Posted in stat.ML · 2026-08-14 · Anandaroop Ray

Extending Occam's inversion with lasso fusion, overcomplete dictionaries, and isotropic total variation regularisation

Occam's inversion is a robust algorithm to perform nonlinear geophysical inversion. It provides the smoothest model within observation noise, thereby discouraging geological overinterpretation. While Occam originally penalised l2 model roughness, l1 can be used to provide models that are visually sharp. However, l1 regularised...

💬 0 commentsarXiv:2608.14225v1PDF
0

Posted in stat.AP · 2026-08-14 · Neha Gupta, Nishit Soni, Aditya Maheshwari

Spillover-Informed Network Architecture for Global Volatility Forecasting

Spillover of volatility shocks across borders during turbulent periods makes accurate equity market volatility forecasts especially critical for risk management, derivatives pricing, and regulatory capital. In this paper, we examine whether volatility forecasts improve when models incorporate information on how markets are connected,...

💬 0 commentsarXiv:2608.14171v1PDF
0

Posted in stat.ME · 2026-08-14 · Nurzhan Sapargali, Sergio Buttazzo, G\''oran Kauermann

Exact Likelihood Inference for Snowball-Sampled Erdős-Rényi Networks

Network data obtained through link-tracing designs, such as snowball sampling, are collected through a mechanism that depends on the very structure the analysis seeks to estimate. Ignoring this dependence and treating the observed sample as though it were itself a complete network can lead to substantially biased inference. While the...

💬 0 commentsarXiv:2608.14129v1PDF
0

Posted in stat.ME · 2026-08-14 · Shanpeng Li, Emily Ouyang, Ace Isabel Mejia-Sanchez, Xinping Cui, Gang Li

FastJM: An R Package for Efficient Implementation of Semiparametric Joint Models for Longitudinal and Survival Data

Joint models provide a flexible framework for characterizing the association between longitudinal and time-to-event processes and have been widely applied in biomedical research. However, fitting joint models can be computationally challenging for large-scale and complex biomedical data. This paper introduces the \proglang{R} package...

💬 0 commentsarXiv:2608.14127v1PDF
0

Posted in stat.ME · 2026-08-14 · Johan Lyrvall, Felix Clouth

An integration of decision trees into latent class modeling with covariates

We propose a novel methodology for fitting decision trees to latent classes. The latent class analysis methodological literature has previously been focusing on logistic models of class membership given covariates, which has important drawbacks in the presence of complex interactions between covariates: logistic models are easily...

💬 0 commentsarXiv:2608.14091v1PDF
0

Posted in stat.ME · 2026-08-14 · Margus Niitsoo, Reimo Rebane, Tarmo Jüristo

A Unified Bayesian Model for Voter Turnout Estimation: Combining Surveys, Aggregate Data, and Selection Correction

Accurate small-area estimation of voter turnout for demographic subgroups is crucial for political analysis but methodologically challenging. Survey data suffer from over-reporting, non-representativeness, and non-ignorable non-response, while ecological inference (EI) from aggregate data is vulnerable to the ecological fallacy. We...

💬 0 commentsarXiv:2608.14062v1PDF
0

Posted in stat.ME · 2026-08-14 · Yifan Zhang, Tianfa Xie, Xinyu Zhang

Handling covariate shift by model averaging

Distributional mismatch between the data used to construct a statistical procedure and the population to which it is ultimately applied is pervasive in modern data analysis. We study covariate shift, a fundamental instance of this problem, and develop an adaptive importance-weighted model averaging method for prediction when labeled...

💬 0 commentsarXiv:2608.14025v1PDF
0

Posted in stat.ME · 2026-08-14 · David J. T. Sumpter

Coherence, charity and triangulation in statistical modelling

Bayesian statistics rests on a few familiar distinctions: frequentist vs. Bayesian, objective versus subjective probability, a model versus the data it is fitted to, a prior versus a posterior. Here, I use Donald Davidson's "third dogma of empiricism" to critique such distinctions in terms of scheme/content dualisms. With a single...

💬 0 commentsarXiv:2608.13986v1PDF