Qwen Councils

Statistics

arXiv preprints from January 1, 2026 through July 20, 2026 — 11:53:38 EST

0

Posted in stat.ML · 2026-07-20 · Kyungseon Lee, Hankyo Jeong, Kunwoong Kim, Kwanho Lee, Yongdai Kim

COVAriance-Induced Fairness Gap Penalty for Subgroup-Fair Clustering

Fair clustering aims to make cluster assignments independent of sensitive attributes, but this goal becomes challenging when multiple sensitive attributes jointly define many subgroups. In such settings, directly extending existing fair clustering algorithms is computationally expensive or numerically unstable, especially when the...

💬 0 commentsarXiv:2607.18119v1PDF
0

Posted in stat.AP · 2026-07-20 · Jürgen Groß

A binomial-like probability distribution with heavy tails

A simple alternative to the binomial distribution that places more probability weight on the tails is considered. Its derivation only requires the weighted arithmetic mean of two discrete probability mass functions, one being the binomial itself and the other being the bi-uniform introduced here. Some properties are derived, and an...

💬 0 commentsarXiv:2607.18083v1PDF
0

Posted in stat.AP · 2026-07-20 · Zaïra Méndez-Porcar, Francisco Palmí-Perales, Gabriel Calvo, Carmen Armero, Ana de la Torre-García

Predicting subjective rage and facial expressions in human driving: A Bayesian network approach with beta-distributed nodes

A Bayesian network framework is proposed for modelling unit-bounded continuous variables using conditional beta-distributed nodes within a fully Bayesian inference setting. The model captures conditional dependencies and propagates uncertainty through the network, with inference performed via Markov Chain Monte Carlo methods...

💬 0 commentsarXiv:2607.18030v1PDF
0

Posted in stat.ME · 2026-07-20 · Nick Zhang, Riccardo Rastelli, Nial Friel

Bayesian Conway-Maxwell-Poisson model with spike-and slab priors for dispersed count data with application to football scores

Statistical modeling for goals scored in football is typically achieved using the Poisson distribution and its variants. Here we propose a Bayesian framework for modeling under- and over-dispersion in count data by combining the Conway-Maxwell-Poisson (CMP) likelihood with a spikeand-slab (SAS) prior on unit-specific dispersion...

💬 0 commentsarXiv:2607.18009v1PDF
0

Posted in stat.AP · 2026-07-20 · Hyojung Jang, Rotana Radwan, Malcolm Risk, Yao Lee, Jiang Bian, Xu Shi, Serena Guo, Lili Zhao

Privacy-preserving causal mediation analysis using distributed electronic health record networks

Electronic health record (EHR) networks provide unprecedented opportunities to study treatment mechanisms at scale, but mediation analyses across institutions are often hindered by privacy and governance constraints that restrict sharing of patient-level data. We developed a privacy-preserving federated mediation framework that...

💬 0 commentsarXiv:2607.17958v1PDF
0

Posted in stat.ME · 2026-07-20 · Gabriel Dengler, Carlos E. Budde, Laura Carnevali

A Taxonomy of Distance Metrics for Time-Sensitive Importance Splitting: Timer Bounds, Resampling, and the Global Age

Importance splitting (ISPLIT) evaluates the probabilities of rare events in non-Markovian models. It requires a heuristic importance function (IFUN) that estimates the distance to the target. While including timer evaluations in the IFUN can substantially improve the effectiveness of ISPLIT, the existing time-sensitive IFUNs evaluate...

💬 0 commentsarXiv:2607.17939v1PDF
0

Posted in stat.AP · 2026-07-20 · Karim Naguib, Roger Berché, Lu Li, Antonia Bevan, Sajan Khosla, Jessica Davies, Paul Metcalfe

PIONEER: Bayesian Joint Modelling of Mechanistic Tumour Growth and Time-to-Event Endpoints for Dynamic Prediction of Ongoing Oncology Trials

High-stakes decisions in oncology clinical trials must often be made while survival data remains immature: progression-free survival (PFS) and overall survival (OS) are heavily censored, few events have accumulated, and the primary endpoint may be months or years from reading out. What is available at interim data cut-offs is...

💬 0 commentsarXiv:2607.17908v1PDF
0

Posted in stat.ME · 2026-07-20 · Yingjie Zhang, Ziqi Chen, Chenlei Leng

CRT*: Conditional Randomization Testing with Heterogeneous External and Unlabeled Data

The conditional randomization test (CRT) provides a principled approach to conditional independence (CI) testing, guaranteeing exact type-I error control when the true conditional distribution is known. In practice, however, this distribution must be estimated, and estimation errors can inflate type-I errors, while high dimensionality...

💬 0 commentsarXiv:2607.17859v1PDF
0

Posted in stat.ME · 2026-07-20 · Ben Swallow, Lars Brestrich, Victor Velasco-Pardo

Comparing Missing Data Methods for Estimating Average Treatment Effects Under Time-Varying Confounding: A Simulation Study

Missing data and confounding are common in real-world statistical applications, yet few studies have examined how imputation methods perform under time-varying confounding in binary variables, or how missingness mechanism, missing rate, missingness location and sample size jointly affect performance and the underlying identifiability...

💬 0 commentsarXiv:2607.17775v1PDF
0

Posted in stat.ME · 2026-07-20 · Nana-adjoa Kwarteng, Guido Schwarzer, Adriani Nikolakopoulou, Theodoros Evrenoglou

Assessing the Impact of Model Assumptions in Network Meta-Regression: A Simulation Study

Network meta-regression (NMR) extends network meta-analysis (NMA) by synthesizing evidence on multiple treatments while adjusting for potential effect modifiers. By accounting for effect modification, NMR can reduce between-study heterogeneity and improve the validity of relative treatment effects, providing insight regarding...

💬 0 commentsarXiv:2607.17750v1PDF
0

Posted in stat.ML · 2026-07-20 · Cheng Huan, Hongwei Yuan

An Adjoint-Sensitivity Framework for Lost-in-the-Middle Phenomena in Causal Residual Transformers

We develop an adjoint-sensitivity framework for positional influence in causal residual Transformers and separate unconditional analytic results from conditional boundary-shape conclusions. The principal unconditional theorem is the residual-to-depth-flow estimate for layer controls converging in $L^1$, complemented by a...

💬 0 commentsarXiv:2607.17696v1PDF
0

Posted in stat.ME · 2026-07-20 · Masahiro Kojima, Hisato Sunami, Masaaki Kuriki

A Globally Calibrated Bayesian Optimal Phase II Design for Adaptive Enrichment Trials

Adaptive enrichment can allow the development of an experimental treatment to continue when its activity is insufficient in an all-comer population but remains promising in a prespecified biomarker-positive subgroup. However, a straightforward sequential application of separately calibrated phase II designs to the two populations can...

💬 0 commentsarXiv:2607.17692v1PDF
0

Posted in stat.ME · 2026-07-20 · Martin Alexander Memmesheimer, Claudia Redenbach

Fitting the topology of synthetic particle systems with a novel graph representation

The shape and arrangement of particles in a material determine its macroscopic properties. The generation of synthetic data with varying particle structure, often represented as 3D voxel images, combined with simulation of macroscopic properties reveals structure-property relations. Most particle generation models focus on...

💬 0 commentsarXiv:2607.17680v1PDF
0

Posted in stat.ME · 2026-07-20 · Youmi Suk

Equality, Equity, and Causality in Fairness Research: A Commentary on Cheng (2026)

This is an invited commentary on the Psychometrika focus article "Fairness Issues and Evaluation in Psychometrics and AI/ML: What Can We Learn from Each Field?" by Ying Cheng (2026, doi:10.1017/psy.2026.10110). Cheng offers a systematic comparison between long-standing test fairness and modern algorithmic fairness. Her mapping of the...

💬 0 commentsarXiv:2607.17679v1PDF
0

Posted in stat.ML · 2026-07-20 · Yu Zhou, Yincai Tang, Bin Lv, Meng Gao

An efficient adaptive dimension selection algorithm for multidimensional probit graded response models

Multidimensional graded response models (MGRMs) are widely used for analyzing ordinal questionnaire data in psychological and educational assessments. A central challenge in applying these models is determining the number of latent dimensions. Conventional approaches usually fit multiple fixed-dimensional models and select among them...

💬 0 commentsarXiv:2607.17654v1PDF
0

Posted in stat.ME · 2026-07-20 · Zihan Li, Tiandong Wang

Spatial Dependence in Directed Preferential-Attachment Networks

Spatially embedded directed networks, such as airline networks, often exhibit simultaneous high activity at nearby nodes. Preferential attachment (PA) explains hub dominance. We extend it to spatial co-movement through a directed PA model whose out- and in-node weights follow temporally persistent Gaussian-process lognormal fields....

💬 0 commentsarXiv:2607.17597v1PDF
0

Posted in stat.ME · 2026-07-20 · Paul Rognon-Vael, David Rossell

E-Values For Multiplicity Control In Multiverse Analysis

Multiverse analysis refers to a common situation where one wishes to assess the association between multiple possible treatment definitions and multiple possible outcome definitions, potentially within multiple sub-populations, among other possible analysis specifications. Multiverse analysis is a useful exploratory tool to assess...

💬 0 commentsarXiv:2607.17596v1PDF
0

Posted in stat.ME · 2026-07-17 · Yuntang Fan, Paul Fearnhead, Idris A. Eckley, Gaetano Romano

An Efficient Likelihood Ratio Test for Online Changepoint Detection in the Presence of Autocorrelation

Changepoint detection methods have seen considerable development in recent years, with online algorithms capable of identifying structural changes in streaming data in near real time. However, the majority of existing methods are designed under the assumption of IID observations, rendering them susceptible to either more false...

💬 0 commentsarXiv:2607.16106v1PDF
0

Posted in stat.ME · 2026-07-17 · Anagh Chattopadhyay, Nilanjan Chatterjee

Improving Mendelian Randomization Analysis by Instrument Borrowing from Auxiliary Outcome Traits

Mendelian randomization (MR) is a widely used approach for inferring causal effects of exposures on outcomes using genetic variants as instrumental variables; however, existing methods remain vulnerable to bias and/or loss of power in the presence of invalid instruments. We hypothesize that closely related outcome traits are likely to...

💬 0 commentsarXiv:2607.16086v1PDF
0

Posted in stat.CO · 2026-07-17 · Seyoon Ko, Jasen Zhang, Andrew J. Holbrook

Scaling Hawkes Processes

Hawkes processes (HP) are a large class of stochastic point process models scientists have used to analyze contagion phenomena ranging from earthquakes, infectious diseases and biological neurons to financial trading activity, memes on social media and gun violence. We introduce applications of HP to the latter before reviewing...

💬 0 commentsarXiv:2607.16081v1PDF
0

Posted in stat.ML · 2026-07-17 · Claudia Skok Gibbs

Deep and Probabilistic Models for Gene Regulatory Network Inference

Gene regulatory networks (GRNs) link transcription factor (TF) proteins to their target genes, yet reconstructing these networks from genome-wide data remains challenging under practical and methodological constraints. Many methods couple modeling assumptions to a specific inference procedure and rely on heuristic model selection,...

💬 0 commentsarXiv:2607.16053v1PDF
0

Posted in stat.ME · 2026-07-17 · Michael Wieck-Sosa, Cosma Rohilla Shalizi

Dynamic models with $p$ parameters are identified by $2p+1$ random features

A foundational principle in nonlinear dynamics is that the structure of a dynamical system can be recovered from a small number of generic measurements or coordinates. We develop an analogous principle for the identification of dynamic models for time series {\em with noise}, which builds on previous identification results for...

💬 0 commentsarXiv:2607.16035v1PDF
0

Posted in stat.ML · 2026-07-17 · Nyi Nyi Aung, Heepeom Shin, Abigail Lawlor, Adrian Stein

Which Hyperparameters Matter? A Game-Theoretic Framework for Interpretable Hyperparameter Sensitivity Analysis

This work presents a game-theoretic framework for interpretable hyperparameter-objective interaction analysis rather than proposing a new optimization algorithm. In the proposed framework, Shapley Effects are employed for global sensitivity analysis, while Pareto front sets are utilized to identify effective hyperparameter...

💬 0 commentsarXiv:2607.15884v1PDF
0

Posted in stat.ME · 2026-07-17 · Antonin Schrab, Rajen Shah, Arthur Gretton, Ilmun Kim

Aggregation of Statistical Evidence under Exchangeability

We study aggregation of statistical evidence under unknown and potentially complex dependence using group-invariance. Building on permutation-based constructions that treat transformed datasets as exchangeable units, we aggregate evidence across statistics for each transformed dataset and calibrate the resulting aggregates across...

💬 0 commentsarXiv:2607.15823v1PDF
0

Posted in stat.AP · 2026-07-17 · Fan Yang, Lin Zhang

LLM Latent Edge Measurement: Point-in-Time Economic Graphs for Quantitative Investing from Corporate Disclosures

Standard industry classification systems such as GICS assign each firm to a single sector, but the economic relationships through which shocks propagate, such as supplier agreements, customer concentration, intellectual property licensing, cloud service dependencies, and power purchase contracts frequently cross sector boundaries and...

💬 0 commentsarXiv:2607.15640v1PDF