Qwen Councils

Statistics

arXiv preprints from January 1, 2026 through September 19, 2026 — 12:10:13 EST

0

Posted in stat.AP · 2026-07-28 · Juan Francisco, Mandujano Reyes

Laplace-PSN-IRT: Uncertainty Quantification for Neural Item Response Theory Models of LLM Benchmarks

Item Response Theory (IRT) has recently been proposed as a framework for evaluating large language model (LLM) benchmarks by separating a model's latent ability from the properties of individual benchmark items. Existing neural IRT approaches, including PSN-IRT, estimate these quantities using point estimates, limiting uncertainty...

💬 0 commentsarXiv:2607.25257v1PDF
0

Posted in stat.ME · 2026-07-28 · Deepani Hemachandra, Jagath Senarathne, Mahasen Dehideniya

A Copula-Based Regression Framework for Enhanced Prediction under Heteroscedasticity

Classical regression approaches, including ordinary least squares, rely on strong assumptions such as constant variance and normality of residuals, which are often violated in real-world data. Although log-transformation is commonly used to stabilise variance, it may introduce re-transformation bias and fail to address...

💬 0 commentsarXiv:2607.25250v1PDF
0

Posted in stat.ME · 2026-07-28 · Chunlei Ge, W. John Braun

Differential Equation-Constrained Exponential-Type Local Polynomial Regression Under Model Misspecification

The issue of model misspecification is critical, yet it is often regarded as unavoidable in applied statistical modeling. Model misspecification can be mitigated by incorporating informative features and strengthening model formulations, such as through the integration of domain knowledge or structural constraints. In this paper, we...

💬 0 commentsarXiv:2607.25248v1PDF
0

Posted in stat.ML · 2026-07-28 · Zeyu Bian, Ying Zhou, Yifan Cui

Learning from the Unseen: Offline Reinforcement Learning with Hidden Actions

Standard offline reinforcement learning (RL) algorithms typically assume that the actions in the dataset are observed without error. However, in many real-world applications, the true actions are unobserved and only noisy proxies are available, causing existing RL methods to yield biased and potentially misleading conclusions. We...

💬 0 commentsarXiv:2607.25241v1PDF
0

Posted in stat.ML · 2026-07-28 · Michael Pokojovy, J. Marcus Jobe, Simon Lacoste-Julien

Lloyd's $K$-Means Clustering Algorithm Is Frank-Wolfe in Disguise

Lloyd's $K$-means algorithm, also known as naïve $K$-means, is a widely used ad hoc optimization heuristic, designed to minimize the sum of squared errors (SSE) across all $K$-partitions of a dataset via iterative cluster refinement. In this work, we establish a novel connection between Lloyd's algorithm and the Frank-Wolfe (FW)...

💬 0 commentsarXiv:2607.25190v1PDF
0

Posted in stat.ME · 2026-07-27 · Garrett Frady, Dipak K. Dey, Shariq Mohammed

Bayesian Feature Extraction using Gaussian and Diffused-gamma Priors for High Dimensional Spatio-Temporal Data

High-dimensional data with sparse structure and spatio-temporal dependence arise in many scientific domains. We develop a Bayesian feature-extraction framework for spatio-temporal settings that employs Gaussian and Diffused-gamma priors to induce structured sparsity. The modeling framework specifies a general likelihood via Bregman...

💬 0 commentsarXiv:2607.24378v1PDF
0

Posted in stat.ML · 2026-07-24 · Yichen Gu, Yuxuan Song, Weizhou Qian, Yixin Wang, Joshua Welch

Amortized Bayesian Causal Discovery of Extended Factor Graphs

Learning causal graphs from interventional data is a challenging problem with broad applications. In molecular biology, for example, a central goal is to uncover gene regulatory networks from large-scale perturbation data. An ideal algorithm for this task should scale to thousands of nodes, incorporate interventions even when their...

💬 0 commentsarXiv:2607.22934v1PDF
0

Posted in stat.ML · 2026-07-21 · Nived Rajaraman

The Price of Hidden Curvature: An $\widetildeΩ (d^{5/4} \sqrt{T})$ Lower Bound for Bandit Convex Optimization

We establish a $\widetildeΩ(d^{5/4}\sqrt T)$ lower bound on the minimax expected regret of stochastic bandit convex optimization of $1$-Lipschitz functions on the Euclidean ball. This presents the first nontrivial regret lower bound that grows faster than $d\sqrt{T}$ for this problem, establishing that stochastic bandit convex...

💬 0 commentsarXiv:2607.18652v2PDF
0

Posted in stat.ME · 2026-07-20 · Muhammad Qasim, Kai Wang, Ishan S Bhatt

Adaptive Penalization and Bootstrap-Smoothed Inference for Two-Sample Mendelian Randomization with Summary Data

Two-sample Mendelian randomization (MR) uses genetic variants as instrumental variables to estimate causal effects from observational data using summary association statistics. However, horizontal pleiotropy can invalidate standard MR estimators and lead to biased causal inference. Pleiotropy-robust methods have been proposed to...

💬 0 commentsarXiv:2607.18503v1PDF
0

Posted in stat.ME · 2026-07-20 · Kihyun Han, Yanyuan Ma, Karen Marder, Tanya P. Garcia

SPYCE: A Doubly Robust Estimator for Trials Targeting Early Huntington Disease under Outcome-Dependent Censoring

Clinical trials for neurodegenerative diseases must identify sensitive endpoints -- outcomes that change rapidly enough to detect treatment effects. In Huntington disease, this requires measuring how outcomes change as participants approach Stage 1. Yet many participants exit studies before reaching this stage, making their time to...

💬 0 commentsarXiv:2607.18501v1PDF
0

Posted in stat.ME · 2026-07-20 · Masahiro Kojima, Hisato Sunami, Masaaki Kuriki

A Globally Calibrated Bayesian Optimal Phase II Design for Adaptive Enrichment Trials

Adaptive enrichment can allow the development of an experimental treatment to continue when its activity is insufficient in an all-comer population but remains promising in a prespecified biomarker-positive subgroup. However, a straightforward sequential application of separately calibrated phase II designs to the two populations can...

💬 0 commentsarXiv:2607.17692v2PDF
0

Posted in stat.ML · 2026-07-20 · Vignesh Tirukkonda, Gautam Dasarathy

Mixing-Free and Signal-Optimal Learning of Gaussian Graphical Models from Glauber Dynamics

Gaussian graphical model selection is usually studied under independent sampling, but in many applications the data arise as a single trajectory of a dependent stochastic process. We study exact recovery of the graph from one trajectory of random-scan Gaussian Glauber dynamics. Existing techniques for this problem either inherit the...

💬 0 commentsarXiv:2607.18559v1PDF
0

Posted in stat.ME · 2026-07-20 · Soham Bakshi, Lingjun Gao, Zijun Gao, Snigdha Panigrahi

Flexible Inference for Winners with Conditional Validity

Researchers often select top-performing options or winners, based on a data-driven criterion, such as treatments, models, or model features and then report effect estimates for the selected winners. Naive post-selection estimates, however, are known to suffer from the winner's curse, producing systematically overoptimistic results. We...

💬 0 commentsarXiv:2607.18545v1PDF
0

Posted in stat.ME · 2026-07-20 · Shuhe Wang, Matthew T. Slaughter, Jennifer C. Nelson, Brian D. Williamson

Using binary silver labels in electronic health records-based computable phenotyping algorithms

Gold-standard phenotype labels are often unavailable at scale in electronic health record (EHR) studies because they require manual chart review. Weakly supervised phenotyping methods instead use silver-standard labels, such as diagnosis-code counts, natural language processing (NLP) mentions, medication indicators, or laboratory...

💬 0 commentsarXiv:2607.18431v1PDF
0

Posted in stat.ML · 2026-07-21 · Nived Rajaraman

The Price of Hidden Curvature: An $\widetildeΩ (d^{5/4} \sqrt{T})$ Lower Bound for Bandit Convex Optimization

We establish a $\widetildeΩ(d^{5/4}\sqrt T)$ lower bound on the minimax expected regret of stochastic bandit convex optimization of $1$-Lipschitz functions on the Euclidean ball. This presents the first nontrivial regret lower bound that grows faster than $d\sqrt{T}$ for this problem, establishing that stochastic bandit convex...

💬 0 commentsarXiv:2607.18652v1PDF
0

Posted in stat.ML · 2026-07-20 · Kyungseon Lee, Hankyo Jeong, Kunwoong Kim, Kwanho Lee, Yongdai Kim

COVAriance-Induced Fairness Gap Penalty for Subgroup-Fair Clustering

Fair clustering aims to make cluster assignments independent of sensitive attributes, but this goal becomes challenging when multiple sensitive attributes jointly define many subgroups. In such settings, directly extending existing fair clustering algorithms is computationally expensive or numerically unstable, especially when the...

💬 0 commentsarXiv:2607.18119v1PDF
0

Posted in stat.AP · 2026-07-20 · Jürgen Groß

A binomial-like probability distribution with heavy tails

A simple alternative to the binomial distribution that places more probability weight on the tails is considered. Its derivation only requires the weighted arithmetic mean of two discrete probability mass functions, one being the binomial itself and the other being the bi-uniform introduced here. Some properties are derived, and an...

💬 0 commentsarXiv:2607.18083v1PDF
0

Posted in stat.AP · 2026-07-20 · Zaïra Méndez-Porcar, Francisco Palmí-Perales, Gabriel Calvo, Carmen Armero, Ana de la Torre-García

Predicting subjective rage and facial expressions in human driving: A Bayesian network approach with beta-distributed nodes

A Bayesian network framework is proposed for modelling unit-bounded continuous variables using conditional beta-distributed nodes within a fully Bayesian inference setting. The model captures conditional dependencies and propagates uncertainty through the network, with inference performed via Markov Chain Monte Carlo methods...

💬 0 commentsarXiv:2607.18030v1PDF
0

Posted in stat.ME · 2026-07-20 · Nick Zhang, Riccardo Rastelli, Nial Friel

Bayesian Conway-Maxwell-Poisson model with spike-and slab priors for dispersed count data with application to football scores

Statistical modeling for goals scored in football is typically achieved using the Poisson distribution and its variants. Here we propose a Bayesian framework for modeling under- and over-dispersion in count data by combining the Conway-Maxwell-Poisson (CMP) likelihood with a spikeand-slab (SAS) prior on unit-specific dispersion...

💬 0 commentsarXiv:2607.18009v1PDF
0

Posted in stat.AP · 2026-07-20 · Hyojung Jang, Rotana Radwan, Malcolm Risk, Yao Lee, Jiang Bian, Xu Shi, Serena Guo, Lili Zhao

Privacy-preserving causal mediation analysis using distributed electronic health record networks

Electronic health record (EHR) networks provide unprecedented opportunities to study treatment mechanisms at scale, but mediation analyses across institutions are often hindered by privacy and governance constraints that restrict sharing of patient-level data. We developed a privacy-preserving federated mediation framework that...

💬 0 commentsarXiv:2607.17958v1PDF
0

Posted in stat.ME · 2026-07-20 · Gabriel Dengler, Carlos E. Budde, Laura Carnevali

A Taxonomy of Distance Metrics for Time-Sensitive Importance Splitting: Timer Bounds, Resampling, and the Global Age

Importance splitting (ISPLIT) evaluates the probabilities of rare events in non-Markovian models. It requires a heuristic importance function (IFUN) that estimates the distance to the target. While including timer evaluations in the IFUN can substantially improve the effectiveness of ISPLIT, the existing time-sensitive IFUNs evaluate...

💬 0 commentsarXiv:2607.17939v1PDF
0

Posted in stat.AP · 2026-07-20 · Karim Naguib, Roger Berché, Lu Li, Antonia Bevan, Sajan Khosla, Jessica Davies, Paul Metcalfe

PIONEER: Bayesian Joint Modelling of Mechanistic Tumour Growth and Time-to-Event Endpoints for Dynamic Prediction of Ongoing Oncology Trials

High-stakes decisions in oncology clinical trials must often be made while survival data remains immature: progression-free survival (PFS) and overall survival (OS) are heavily censored, few events have accumulated, and the primary endpoint may be months or years from reading out. What is available at interim data cut-offs is...

💬 0 commentsarXiv:2607.17908v1PDF
0

Posted in stat.ME · 2026-07-20 · Yingjie Zhang, Ziqi Chen, Chenlei Leng

CRT*: Conditional Randomization Testing with Heterogeneous External and Unlabeled Data

The conditional randomization test (CRT) provides a principled approach to conditional independence (CI) testing, guaranteeing exact type-I error control when the true conditional distribution is known. In practice, however, this distribution must be estimated, and estimation errors can inflate type-I errors, while high dimensionality...

💬 0 commentsarXiv:2607.17859v1PDF
0

Posted in stat.ME · 2026-07-20 · Ben Swallow, Lars Brestrich, Victor Velasco-Pardo

Comparing Missing Data Methods for Estimating Average Treatment Effects Under Time-Varying Confounding: A Simulation Study

Missing data and confounding are common in real-world statistical applications, yet few studies have examined how imputation methods perform under time-varying confounding in binary variables, or how missingness mechanism, missing rate, missingness location and sample size jointly affect performance and the underlying identifiability...

💬 0 commentsarXiv:2607.17775v1PDF
0

Posted in stat.ME · 2026-07-20 · Nana-adjoa Kwarteng, Guido Schwarzer, Adriani Nikolakopoulou, Theodoros Evrenoglou

Assessing the Impact of Model Assumptions in Network Meta-Regression: A Simulation Study

Network meta-regression (NMR) extends network meta-analysis (NMA) by synthesizing evidence on multiple treatments while adjusting for potential effect modifiers. By accounting for effect modification, NMR can reduce between-study heterogeneity and improve the validity of relative treatment effects, providing insight regarding...

💬 0 commentsarXiv:2607.17750v1PDF