Qwen Councils

Statistics

arXiv preprints from January 1, 2026 through September 19, 2026 — 06:25:35 EST

0

Posted in stat.ME · 2026-08-21 · Piotr Fryzlewicz

LABS: Extending the scope of binary segmentation via a look-ahead device

Binary segmentation is widely used for multiple change-point detection because it is fast, simple to describe, and simple to implement. Its validity rests on the requirement that, at each recursive stage, the procedure identifies one of the true change-points when several are present in the current interval. This holds for detecting...

💬 0 commentsarXiv:2608.21122v1PDF
0

Posted in stat.ME · 2026-08-21 · Rok Spruk

Public Signals, Concealed Choices: Dynamic Measurement without Behavioral Identification

Members of collective institutions may leave public traces while their individual choices remain concealed. This paper separates a corpus-conditional public position from the behavioral rule linking that position to participation and secret choice. I measure the first with a dynamic ordinal state-space model and establish a...

💬 0 commentsarXiv:2608.21077v1PDF
0

Posted in stat.AP · 2026-08-21 · Yipeng Wei, Zahra Hoodbhoy, Emily R. Smith, Fang Jin, Muhammad Imran Nisar, Muhammad Farrukh Qazi, Christopher Mores, Victor Akelo, Caleb Sagam, Florence Aweyo, Charlotte Tawiah, Veronica Agyemang, Kwaku Poku Asante, Sam Newton, Santosh Joseph Benjamin, Anne George Cherian, Devakumar Devadhas, James A, Margaret P. Kasaro, Augustine Tunga, Sarmila Mazumder, Neeraj Sharma, Wilbroad Mutale, Mae Bridget Spelke, Qing Pan

Knowledge-guided Transfer Prediction In Underrepresented Populations: A GRU-D-Static Framework For Maternal And Neonatal Outcomes

Integrating summary-level scientific knowledge into neural network models provides a practical strategy for transferring prediction models trained on adequately sampled source cohorts to underrepresented target populations, where individual-level data in the target domain are often limited or unavailable. In this study, we propose...

💬 0 commentsarXiv:2608.21073v1PDF
0

Posted in stat.AP · 2026-08-21 · Charu Gupta, Gabriel Innocenzi, Christina Yap, Daniel Jackson, Fabio Rigat

Calibration of clinical trial sample size based on design utility

Clinical trial design relies on both statistical and clinical considerations for pre-specification of potentially practice-changing target treatment effects. As larger trials tend to be associated with high power and modest minimal detectable benefit, trial sample size is typically calibrated with reference to relevant precedents to...

💬 0 commentsarXiv:2608.20997v1PDF
0

Posted in stat.ME · 2026-08-21 · Zern Ke, Mingshi Cui, Feng Dai, Birol Emir, Javier Cabrera, Demissie Alemayehu

From Cumulative Weights to Marginal Density Ratios: Per-Protocol Estimation in Sequential Target Trial Emulation

Sequential target trial emulation evaluates eligibility at multiple baseline times to emulate a sequence of randomized trials using observational data. Estimating per-protocol effects in this setting is challenging because treatment deviations and loss to follow-up induce selection among individuals who remain observed and adherent...

💬 0 commentsarXiv:2608.20976v1PDF
0

Posted in stat.ME · 2026-08-21 · Eylul Fidan, Ufuk Beyaztas, Soutir Bandyopadhyay

Spatial function-on-function quantile regression

This paper introduces a novel penalized spatial function-on-function quantile regression framework for analyzing spatially indexed functional data, bridging a critical gap between spatial functional models and quantile regression. Our work makes three key contributions. First, we propose the first spatial function-on-function quantile...

💬 0 commentsarXiv:2608.20919v1PDF
0

Posted in stat.ME · 2026-08-21 · Samhita Pal, Jared D Huling

Heterogeneous Effects of Continuous Treatments via Conditional Modified Treatment Policies

For continuous treatments such as drug dose or ventilator intensity, a key clinically actionable question is whether a modest, patient-specific adjustment to the current dose would help or harm, rather than whether to treat at all. Standard estimands such as average or conditional dose-response functions require positivity across a...

💬 0 commentsarXiv:2608.20744v1PDF
0

Posted in stat.ME · 2026-08-21 · Stephany Lima de Oliveira, Frederico Machado Almeida

A modified score function for monotone likelihood in promotion time cure rate models

Survival models that incorporate a cure fraction provide a flexible framework for jointly modeling the cure and the survival distributions. However, when the data comprise a high proportion of censored observations or highly unbalanced binary covariates, maximum likelihood estimation may become unstable, leading to parameter estimates...

💬 0 commentsarXiv:2608.20641v1PDF
0

Posted in stat.ML · 2026-08-21 · Cholyeon Cho, Yuchen Wu

Minimax Optimality of Score-Entropy Discrete Diffusion

Discrete diffusion models have demonstrated strong performance across a range of datasets, including natural language data and graph-structured data. Among many variants, score-entropy discrete diffusion (SEDD) has achieved particularly strong empirical results. In SEDD, new samples are generated by iteratively evaluating a sequence...

💬 0 commentsarXiv:2608.20635v1PDF
0

Posted in stat.ME · 2026-08-20 · Soonhong Cho

Let Time Tell: Identification and Gaussian Process Estimation for Interrupted Time Series

We study causal inference in interrupted time series designs where a treatment affects every unit simultaneously, so that the contemporaneous controls used by difference-in-differences and synthetic control are unavailable and the counterfactual must be extrapolated from a unit's own pre-treatment history. We establish identification...

💬 0 commentsarXiv:2608.20610v1PDF
0

Posted in stat.ME · 2026-08-20 · Hyungjoon Kim, Andee Kaplan, Matthew D. Koslovsky

A Comprehensive Bayesian Approach to Entity Resolution for Data with Multiple Truths

In many applications, from government to ecology, integrating data from diverse and noisy sources is critical for downstream inference. However, a unique identifier to link records cleanly from the same entity may not exist. Entity resolution (also referred to as de-duplication or record linkage) merges such databases to identify...

💬 0 commentsarXiv:2608.20601v1PDF
0

Posted in stat.ME · 2026-08-20 · David McCoy, Yi Li

Targeted Deep Survival Contrasts: Valid Inference for Treatment-Specific Survival Benefit with Neural Networks

Neural survival models are increasingly asked to support counterfactual claims---how much a treatment would change survival in a population---rather than only prognostic risk scores. Answering such questions from observational data requires valid inference for treatment-specific survival contrasts under confounding and...

💬 0 commentsarXiv:2608.20598v1PDF
0

Posted in stat.ME · 2026-08-21 · Simon Rudkin, Wanling Rudkin

A Multiscale Ball Test for Conditional Mean Independence

Tests of conditional mean independence can lose power when departures are confined to a bounded part of a multivariate predictor space and the relevant spatial scale is unknown. We propose a Multiscale Ball Conditional Mean Independence (MBCMI) test that aggregates support-weighted local mean contrasts in an outcome variable across...

💬 0 commentsarXiv:2608.20727v1PDF
0

Posted in stat.ME · 2026-08-21 · Deepra Ghosh, Sanat K. Sarkar

Controlling the False Discovery Rate Control in Two-Sided Gaussian Mean Testing Under Arbitrary Dependence

The recent work of Sarkar and Zhang (2025) introduced Positive Tail Dependence Under the Null (PTDN) and developed Generalized Shifted Benjamini-Hochberg (BH) procedures for two-sided Gaussian $z$- and $t$-testing under known covariance structures. This paper develops further consequences of that framework. First, we derive explicit...

💬 0 commentsarXiv:2608.21267v1PDF
0

Posted in stat.AP · 2026-08-21 · Niklas Heusch

A Synthetic Benchmark Dataset with Endogenous Marketing Spend for Validating Marketing Mix Models

Marketing Mix Models (MMMs) estimate the incremental sales effect of advertising from observational time series, yet they are rarely validated against ground truth, because ground truth is unobservable in real data. Synthetic data closes that gap in principle, but existing generators produce marketing spend exogenously - omitting the...

💬 0 commentsarXiv:2608.21130v1PDF
0

Posted in stat.AP · 2026-08-21 · Niklas Heusch

Structural Estimation of Marketing Mix Model Parameters from Geo-Experiments

Marketing Mix Models (MMMs) are widely used for marketing measurement and budget allocation, but face fundamental identification challenges: due to endogenous marketing spend decisions, MMM estimation on observational time-series data cannot recover the true causal effects of marketing. On the other hand, geo-experiments provide...

💬 0 commentsarXiv:2608.21128v1PDF
0

Posted in stat.ME · 2026-08-20 · Montserrat Fuentes, Veronica B. Patterson

From Kriging to Spatial AI: Fifty Years of Spatial Statistics for Complex Dependent Data

Spatial statistics has grown from kriging for spatial prediction into a broad framework for learning from complex dependent data. This article traces that development from random fields and spectral methods to Bayesian hierarchical models and scalable computation. It then connects these foundations to Spatial AI, where graph learning...

💬 0 commentsarXiv:2608.20260v1PDF
0

Posted in stat.ML · 2026-08-20 · Junpeng Ren, Carlos Misael Madrid Padilla, Yanzhen Chen, Oscar Hernan Madrid Padilla

Transfer Learning in Nonparametric Regression with Deep ReLU Networks

This paper develops a general transfer learning framework for nonparametric regression with data consisting of multiple groups. Under the assumption that groups share a common structure along with group-specific deviations in additive form, the proposed method employs a two-stage offset learning procedure: the first stage pools data...

💬 0 commentsarXiv:2608.20255v1PDF
0

Posted in stat.ME · 2026-08-20 · Montserrat Fuentes, Veronica B. Patterson

A Bayesian Edge-Space Framework for Whole-Connectome Inference in Multisite Autism Neuroimaging

Autism spectrum disorder (ASD) is associated with heterogeneous alterations across distributed brain systems, creating challenges for whole-connectome inference. The difficulty arises not only from the large number of connections, but also from dependence among effects indexed by anatomically and functionally related region pairs. We...

💬 0 commentsarXiv:2608.20243v1PDF
0

Posted in stat.ME · 2026-08-20 · Manish Gupta, Dipanjan De

Multi-Method Causal Evidence Synthesis: Ranking Candidate Drivers by Convergent Cross-Method Evidence from Observational Data

Practitioners inferring causality from observational data usually rely on a single method and treat its output as causal truth. Recent tools select an optimal method for a dataset, and recent ensembles aggregate multiple causal-discovery algorithms into one graph, but little work pools evidence across different mathematical...

💬 0 commentsarXiv:2608.20187v1PDF
0

Posted in stat.ML · 2026-08-20 · Lohithsai Yadala Chanchu, Hany Abdulsamad, Christian A. Naesseth

Discrete Diffusion Inference-Time Control with Nested Sequential Monte Carlo

We study inference-time control for text generation in discrete diffusion language models, where the goal is to steer sampling toward sequence-level rewards without retraining. Prior work in this domain has focused on particle-based methods such as best-of-$n$ sampling and bootstrap sequential Monte Carlo, which may suffer from...

💬 0 commentsarXiv:2608.20123v1PDF
0

Posted in stat.ME · 2026-08-20 · Yan Liu, Anita Koushik, Philippe Boileau, Cong Jiang, Miceline Mésidor, Claudia Waddingham, Denis Talbot, Mireille E. Schnitzer

Causal inference via propensity scores for case-control studies

Propensity score methods for causal inference are increasingly being used in cohort and experimental designs, but their development and uptake in outcome-dependent sampling schemes, such as case-control studies, remains limited. Case-control studies involve the sampling of individuals with and without an outcome of interest with the...

💬 0 commentsarXiv:2608.20080v1PDF
0

Posted in stat.AP · 2026-08-20 · Alejandro Rozo Posada, Maxime Fajgenblat, Christel Faes, James Colborn, Emanuele Giorgi, Baltazar Candrinho, Thomas Neyens

Integrating Temporal Disaggregation and Distributed Lag Nonlinear Models for Bayesian Spatio-Temporal Disease Mapping with High-Resolution Environmental Exposures

Environmental conditions are major drivers of malaria transmission, but epidemiological analyses are often constrained by temporal misalignment between health outcomes reported at coarse time scales and environmental exposures available at finer resolutions. Conventional approaches aggregate environmental data to match health...

💬 0 commentsarXiv:2608.20046v1PDF
0

Posted in stat.ME · 2026-08-20 · Rianne de Heide

Where Does the Union Bound Go? Best-Arm Identification and Strong FWER Control

In fixed-confidence best-arm identification, proofs often use a union bound across the competing arms. From a multiple-testing point of view this can look puzzling: if the best arm is unique, only one hypothesis of the form ``arm $i$ is best'' can be true. Why then should there be a Bonferroni-type factor of $K-1$? The answer is that...

💬 0 commentsarXiv:2608.19903v1PDF
0

Posted in stat.ME · 2026-08-20 · Mark Cary, Charles Bokor

A Repeated Measurements Approach to $SoH$ Battery Modelling of Cyclic Aged Data in a Laboratory Environment

This document describes the application of a first order linearised nonlinear repeated measurements approach to the analysis of battery cell ageing profiles generated under controlled conditions in a laboratory. The primary advantage of the model is it reflects the obvious structure in the data. Consequently, it is a two-component of...

💬 0 commentsarXiv:2608.19879v1PDF