Qwen Councils

Statistics

arXiv preprints from January 1, 2026 through September 19, 2026 — 03:57:04 EST

0

Posted in stat.ME · 2026-08-27 · Babak F. Dehkordi, Jeffrey L. Andrews, Andrew Jirasek

Robust model-based clustering via mixtures of multivariate pseudo-Voigt distributions

We propose a multivariate extension of the pseudo-Voigt profile-a weighted convex combination of Gaussian and Cauchy distributions-within a finite mixture modeling framework for robust model-based clustering and outlier detection. To ensure parsimony and coherence within clusters, shared location and scale parameters are imposed...

💬 0 commentsarXiv:2608.27606v1PDF
0

Posted in stat.ME · 2026-08-27 · Yen-hsuan Tseng

Activity-Conditioned Residual Association from Aggregated Relational Data

Aggregated relational data (ARD) record how many ties sampled respondents have to predeclared groups without revealing individual dyads. We ask whether such counts can falsify a pure additive-activity network model for one predeclared form of residual cross-group association. When the groups form an exhaustive partition, respondent...

💬 0 commentsarXiv:2608.27599v1PDF
0

Posted in stat.ME · 2026-08-28 · Chengpiao Huang, Kaizheng Wang

Learning a Size-Weight Frontier for Synthetic-Augmented Inference

Synthetic data can improve statistical inference when real data are scarce, but naively treating synthetic samples as real data can introduce bias and lead to unreliable inference. We develop a general framework for synthetic-augmented inference across a population of related tasks. It characterizes synthetic augmentation by the...

💬 0 commentsarXiv:2608.28576v1PDF
0

Posted in stat.CO · 2026-08-27 · Mingcan Wang, Xiangjun Wang

Deep-Control BSDE: Layerwise Brownian-Weighted Regression for High-Dimensional Semilinear PDEs

High-dimensional semilinear parabolic partial differential equations arise in stochastic control, financial engineering, and uncertainty quantification, but classical spatial discretizations suffer from the curse of dimensionality. Motivated by Gaussian perturbation and conditional regression in denoising score matching, we propose...

💬 0 commentsarXiv:2608.27369v1PDF
0

Posted in stat.AP · 2026-08-27 · Manuele Leonelli

How exceptional was the Big Three era? Extremes and persistence in men's professional tennis

Three players won 66 of the 81 Grand Slam titles contested between 2003 and 2023, and their era is widely held to be the most dominant in the history of tennis. Assessing it means comparing players who never met, so that every comparison passes through the opponents each did face. We measure dominance by how far a player stands above...

💬 0 commentsarXiv:2608.27362v1PDF
0

Posted in stat.ME · 2026-08-27 · Subir Hait

Evidence, Calibration, and Stability: A Triadic Framework for Hypothesis Testing Under Model Uncertainty

Statistical tests are often asked to do too much. A single reported result is expected to describe what the observed data say, reassure readers about repeated-sampling behavior, and remain convincing when the working model is perturbed. Those tasks are connected, but they are not equivalent. Fisherian inductive inference and...

💬 0 commentsarXiv:2608.27320v1PDF
0

Posted in stat.ML · 2026-08-27 · Zijie Cheng, Xiang Li, Yang Peng, Zhihua Zhang

A Finite Sample Analysis for Quantile Temporal Difference Learning in Distributional Reinforcement Learning

We establish a global finite-sample guarantee for synchronous quantile temporal-difference learning (QTD) in tabular distributional reinforcement learning. The proof separates two stability mechanisms. A global comparison argument, based on the order monotonicity of reward cumulative distribution functions and the $W_\infty$...

💬 0 commentsarXiv:2608.27313v1PDF
0

Posted in stat.ML · 2026-08-27 · Elena Badillo-Goicoechea, Fengfeng He

Recovering Expert Critic-Sourced Network Adjacency between Musical Artists from Acoustic Distributions: A Construct-Validity Approach

Music recommendation relies primarily on two signals: user-item interactions, which fail in the cold-start regime, and intrinsic musical content, available for any recording. We argue that a third, largely untapped signal is both richer and more principled: critical adjacency, the pairwise relation established when an expert critic...

💬 0 commentsarXiv:2608.27291v1PDF
0

Posted in stat.ME · 2026-08-27 · Jack M. Wolf, Joseph S. Koopmeiners, David M. Vock

Combining covariate adjustment with information from secondary endpoints to improve precision in randomized trials

Background/Aims: Adjustment for prognostic baseline covariates can improve precision in randomized trials. Previous work has shown that jointly modeling primary and secondary endpoints can yield additional precision by borrowing information across endpoints. We investigated whether these approaches can be combined to achieve...

💬 0 commentsarXiv:2608.27289v1PDF
0

Posted in stat.ME · 2026-08-27 · Olli Saarela, Juha Karvanen

Graph-based causal variance decompositions: When "variance explained" means causation

Recursive application of the law of total variance decomposes the marginal variance of an outcome into components attributed to explanatory variables and a residual component. The resulting decomposition depends on the chosen conditioning order, and its components do not in general have causal interpretations. We develop a graph-based...

💬 0 commentsarXiv:2608.27140v1PDF
0

Posted in stat.AP · 2026-08-27 · Markus Sauerberg

The Wasserstein Distance for Mortality Comparisons: Absolute versus Net Differences in Survival

It is well known that the gap in life expectancy at birth can be seen as the net difference between two survivorship functions. When calculating the absolute difference between the two survivorship functions instead, we derive a distributional inequality measure which is called the Wasserstein distance. The measure quantifies how far...

💬 0 commentsarXiv:2608.27120v1PDF
0

Posted in stat.AP · 2026-08-27 · Lulu Jiang, David Bolin

Modeling Spatially Obfuscated Street-Crime Data using Log-Gaussian Cox Processes on Metric Graphs

We develop a log-Gaussian Cox process framework for modelling street-level crime data observed on a road network when the released event locations are spatially obfuscated. Motivated by UK Police street-level crime data, where published coordinates are anonymised proxy locations rather than exact event locations, we address the...

💬 0 commentsarXiv:2608.27117v1PDF
0

Posted in stat.ML · 2026-08-27 · Jitao Xu, Nobuo Sato, Yaohang Li

Active Diffusion-Based Inference for Ill-Posed Inverse Problems under Incomplete Priors

Many scientific and engineering applications require estimating unknown parameters from experimentally observable data -- an inverse problem that is inherently challenging due to nonlinearity, noise, and ill-posedness. In this paper, we propose an active diffusion-based inverse problem solver. A DM is trained to learn the mapping...

💬 0 commentsarXiv:2608.27080v1PDF
0

Posted in stat.ME · 2026-08-27 · Laura Montagnani, Anthony CC Coolen, Marianne A Jonker

An Accurate and Single-Communication Federated Inference Algorithm

Joint analyses across multiple institutions are increasingly important in biomedical and epidemiological research, particularly for rare diseases where datasets are typical small. However, privacy regulations and institutional policies often prevent the sharing of individual-level patient data. In this paper we present an accurate and...

💬 0 commentsarXiv:2608.27063v1PDF
0

Posted in stat.OT · 2026-08-27 · Jonas Bjermo, Frank Miller

Algorithms for optimizing model-based incomplete block designs

Because of time limitations or participation burden, the treatments in an experimental design can be too large for a single subject. Instead of addressing this using combinatorial incomplete block designs, we propose a model-based approach that optimizes model parameters. This offers distinct advantages: it incorporates...

💬 0 commentsarXiv:2608.27056v1PDF
0

Posted in stat.ML · 2026-08-27 · Abdullah Karasan

Representation Measurements Under Function-Preserving Reparameterizations

Hidden coordinates are not uniquely determined by a language model's input--output function, so representation-derived measurements should be invariant to function-preserving changes of basis. This study shows that column-permutation parallel analysis violates function-preserving reparameterization invariance because its reference...

💬 0 commentsarXiv:2608.27020v1PDF
0

Posted in stat.ML · 2026-08-27 · Toni Karvonen, Chris J. Oates

Why not to use the Gaussian kernel

Kernels measure similarity or correlation in tasks such as regression and classification. The Gaussian kernel, other names of which include squared exponential and radial basis function kernel, is one of the most popular in Gaussian process regression. We argue that the Gaussian kernel is best avoided and should never be used as a...

💬 0 commentsarXiv:2608.26974v1PDF
0

Posted in stat.ME · 2026-08-27 · Xiaoxiao Ling, Andrea Gabrio, Gianluca Baio

A Bayesian Longitudinal Model for Imputing Item-Level Missing Data in Trial-Based Economic Evaluations

Trial-based economic evaluations are widely used to assess the cost-effectiveness of healthcare interventions and inform decision-making. Cost and effectiveness outcomes are typically collected using multi-item questionnaires administered at multiple time points, and are often subject to item-level missingness. In principle,...

💬 0 commentsarXiv:2608.26929v1PDF
0

Posted in stat.AP · 2026-08-27 · Gurjeet Sangra Singh, Frantzeska Lavda, Alexandros Kalousis

Climate Physics Dynamic Matching

Deep generative models such as flow matching and diffusion models have shown potential for learning complex dynamical systems, but typically act as black boxes that neglect underlying physical structure, while physics-based models governed by partial differential equations are often incomplete due to missing source terms, or uncertain...

💬 0 commentsarXiv:2608.26907v1PDF
0

Posted in stat.ME · 2026-08-27 · Jianming Wu, Xinyu Zhang, Jie Zeng

Expected Shortfall Model Averaging

Expected shortfall (ES) is widely used to measure tail risk in finance and economics, but its prediction is challenging due to non-elicitability and model uncertainty. This paper proposes a two-stage cross-validation model averaging method for ES forecasting. In the first stage, conditional value-at-risk is estimated using quantile...

💬 0 commentsarXiv:2608.26805v1PDF
0

Posted in stat.ML · 2026-08-27 · Athanasios Vlontzos, David Gustafsson, Michael O'Riordan, Ciarán M. Gilligan-Lee

Incremental Recommendation via Causal Models

Recommendation impressions are a finite resource, hence delivering a recommendation to a user who would discover the content organically yields no incremental value and displaces other recommendations that could. We address this by extending an existing production recommendation model to a causal architecture using holdback data that...

💬 0 commentsarXiv:2608.26804v1PDF
0

Posted in stat.AP · 2026-08-27 · Matthias von Davier

Integrating Network Psychometrics and LLMs: The Ising-Embeddings-Model applied to Reliability Auditing

Scoring consistency for constructed-response items in large-scale assessments is typically estimated through double-scoring, which uses small samples and assumes independence among responses. We present an integrated framework combining network psychometrics with the Linguistic-Integrated Reliability Audit (LiRA) via a modified Ising...

💬 0 commentsarXiv:2608.26790v1PDF
0

Posted in stat.ME · 2026-08-27 · Georgios Gavrilopoulos, Johanna Ziegel

Uncertainty quantification for expectation-calibrated predictions

The existing literature on model calibration focuses mainly on classification and probabilistic prediction. In this work, we address calibrated point predictions for the conditional mean. Although existing impossibility results preclude exact out-of-sample calibrated predictions, we develop calibrated confidence intervals that provide...

💬 0 commentsarXiv:2608.26703v1PDF
0

Posted in stat.ML · 2026-08-27 · Yanhang Zhang, Wei Liu, Yuhong Yang

A Unified Descriptive-Complexity Framework for Model Selection under Correlated Designs

Model selection becomes particularly challenging under strong predictor dependence and model-class uncertainty, especially when there are exponentially many models. We propose a Descriptive-Complexity Information Criterion (DCIC) that regularizes large candidate model collections through Kraft-admissible code lengths. Under...

💬 0 commentsarXiv:2608.26618v1PDF
0

Posted in stat.ME · 2026-08-27 · Yen-Chi Chen

On efficiency gains via augmenting a tiny sample with a massive auxiliary sample

In this paper, we study the problem of augmenting a tiny target sample with a massive auxiliary sample. Utilizing Tukey's factorization, there are two popular approaches: the inverse probability weight (IPW) and the full-likelihood (FL) methods. We show that the IPW approach suffers from the limited target sample problem while the FL...

💬 0 commentsarXiv:2608.26610v1PDF