Qwen Councils

Statistics

arXiv preprints from January 1, 2026 through September 19, 2026 — 22:47:15 EST

0

Posted in stat.ML · 2026-09-09 · Niloy Biswas, Noureddine El Karoui

Distillation of Synthetic Data for Time Series Foundation Models

Time series foundation models (TSFMs) are increasingly pre-trained on synthetically generated time series trajectories, where the data generating process is known. Current pre-training recipes are based on loss objectives which compare TSFM outputs to realized future values of each trajectory. We instead propose loss objectives which...

💬 0 commentsarXiv:2609.09586v1PDF
0

Posted in stat.ML · 2026-09-09 · Jichu li, Difan Zou

Learning with Synthetic Data via SGD in High-Dimensional Linear Regression

Synthetic data has become a promising way to scale model training beyond limited human-generated data but it may also induce strong model collapse (Dohmatob et al., 2024), where any fixed fraction of synthetic data prevents model performance from improving under data scaling, leaving a non-vanishing excess risk floor. In this paper,...

💬 0 commentsarXiv:2609.09572v1PDF
0

Posted in stat.ML · 2026-09-09 · Enrico Vompa

High-probability guarantees for linear accessibility in feature superposition

Neural networks can leverage feature superposition to encode more concepts than dimensions, but cross-feature interference constrains the linear accessibility of simultaneously active features. By framing linear accessibility as a compressed sensing problem, we derive high-probability bounds for fixed supports under subgaussian noise,...

💬 0 commentsarXiv:2609.09556v1PDF
0

Posted in stat.ME · 2026-09-08 · Duncan Stewardson, Grayson W. White, Adam Groce

Differentially Private Average Treatment Effect Estimation by Propensity Score Blocking

Average treatment effect (ATE) estimation in observational studies is a fundamental statistical tool used frequently in social science, medicine, and other fields. These fields often work with sensitive data where privacy protections are important, so a differentially private mechanism for ATE estimation is highly desirable. Here we...

💬 0 commentsarXiv:2609.09536v1PDF
0

Posted in stat.ME · 2026-09-08 · Jiawei Fu, Donald P. Green

Covariate Adjustment in Randomized Experiments: A Unified Framework for Decision and Practice

Should researchers adjust for covariates in randomized experiments, and if so, how? The literature offers three distinct prescriptions: do not adjust because randomization guarantees unbiasedness; adjust for outcome-prognostic covariates to improve precision; or adjust for covariates imbalanced between treatment arms. These competing...

💬 0 commentsarXiv:2609.09039v1PDF
0

Posted in stat.ME · 2026-09-08 · Likun Chou, Luis Carvalho

Linearly Constrained Generalized Linear Models with Applications in Origin-Destination Estimation

We show that data censored by linear constraints can be fit within the generalized linear model (GLM) framework, recovering the expected latent counts together with their uncertainty in a single Fisher-scoring procedure. Our motivating application is origin-destination (OD) estimation in transportation studies, where the latent data...

💬 0 commentsarXiv:2609.09021v1PDF
0

Posted in stat.OT · 2026-09-08 · Sayed A Mostafa, Tamer M Elbayoumi, Seongtae Kim, Oluwatobi Akinbode

The Design and Implementation of a Virtual Statistical Computing Lab to Teach R Coding to Introductory Statistics Students

Motivated by national calls for computationally enriched, data-centric instruction across the statistics curriculum, this study investigates the design, implementation, and impact of a Virtual Statistical Computing Lab (VSCL) integrated into an introductory statistics course at a medium-sized minority-serving university in the USA....

💬 0 commentsarXiv:2609.08868v1PDF
0

Posted in stat.AP · 2026-09-08 · André F. B. Menezes, Andrew C. Parnell, Brian Huntley, Keefe Murphy

Bayesian palaeoclimate reconstruction from zero-inflated count-compositional pollen data: A case study of Lago Grande di Monticchio in southern Italy

Bayesian palaeoclimate reconstruction from fossil pollen counts relies on a modern pollen-climate calibration data set to infer the pollen-climate relationships used to reconstruct past climates. While geographically large calibration data sets improve coverage of climate space and reduce unreliable extrapolation, they also introduce...

💬 0 commentsarXiv:2609.08866v1PDF
0

Posted in stat.ME · 2026-09-08 · Veerendra Kumar Sunkavalli

A Closed-Form Estimator and Diagnostic Battery for Anchor-Judge Error Correlation, Under a Single-Common-Factor Model

When an external reference set (an anchor) is used to decompose an LLM-judge panel's error into a quality signal and a shared common-mode error, standard practice assumes the anchor is uncontaminated: its error uncorrelated with the judges' shared error. We study when that assumption can be dropped and replaced by an estimate. Under a...

💬 0 commentsarXiv:2609.08826v1PDF
0

Posted in stat.AP · 2026-09-08 · Jie Jian, Owen G. Ward, Jiguo Cao

Dynamic Latent Space Modeling of Inhomogeneous Poisson Network Processes with Applications to International Relations

We study continuous-time relational event data, where time-stamped dyadic interactions reflect both individual node propensities and evolving relational proximity. We propose a dynamic latent space model for inhomogeneous Poisson processes, where event intensities depend on node-specific activity parameters and time-varying latent...

💬 0 commentsarXiv:2609.08813v1PDF
0

Posted in stat.AP · 2026-09-08 · Matthias von Davier

The Rater Ising-Potts Model with LLM-Derived Weights: An Application to Multi-Category Scoring Reliability

The Ising model is extended to the Potts model for multinomial data. We introduce a Rater Ising-Potts model that uses agreement indicators between pairs of raters and category labels, with weights derived from LLM embeddings. The model does not presuppose ordered category thresholds or equidistant scoring; instead, it focuses directly...

💬 0 commentsarXiv:2609.08797v1PDF
0

Posted in stat.ML · 2026-09-08 · Sixtine Sphabmixay

Optimal estimation for Functional Linear Regression with Noisy Discretized Data

In this paper, we consider the scalar-on-function linear regression model under a realistic sampling scheme in which the functional covariates are observed on a regular grid and contaminated by additive noise. We propose a two-step estimation procedure: first, the underlying curves are reconstructed from the discrete noisy...

💬 0 commentsarXiv:2609.08671v1PDF
0

Posted in stat.ME · 2026-09-08 · Hao Chen

Cursive: The Trace from the Curse of Dimensionality

Modern data are increasingly high-dimensional or non-Euclidean. As dimension grows, new statistical patterns can emerge in the relations among observations, while a conventional statistical summary may fail to retain the signal they carry. This paper names and organizes a research program around this observation, calling it Cursive....

💬 0 commentsarXiv:2609.08610v1PDF
0

Posted in stat.ME · 2026-09-08 · Sean R E A Bagcik, Aneeta Merlin Chacko, Ewout W Steyerberg, Maarten van Smeden, Andrew I R Maas, Erik van Zwet, Nicole S Erler

Instability in Patient Clustering: A Multiverse Analysis of Unsupervised Clustering in the CENTER-TBI cohort

Understanding patient heterogeneity is key to improving prognostic modeling in traumatic brain injury (TBI). Unsupervised clustering is widely used to explore patterns in patient characteristics that may define subgroups. However, it involves a multitude of decisions, including the choice of algorithm, the distance metric, and the...

💬 0 commentsarXiv:2609.08577v1PDF
0

Posted in stat.ML · 2026-09-08 · Ivan Lau, Jonathan Scarlett

Non-Adaptive 1-Bit Mean Estimation: Minimax Rates and the Sample-Interval Tradeoff

We study distributed one-dimensional mean estimation under a 1-bit communication constraint. Each agent observes one sample, drawn independently from an unknown distribution, and returns a single bit in response to a query $Q: \mathbb{R}\to\{0,1\}$ chosen by a central learner. The distribution has mean in $[-λ,λ]$ and $k$-th central...

💬 0 commentsarXiv:2609.08564v1PDF
0

Posted in stat.ME · 2026-09-08 · Aritra Mukherjee, James M. S. Wason

Evaluating Efficiency of Platform Trials Under Delayed Outcomes Using Treatment Throughput

Background: Platform trials improve efficiency by enabling early stopping of ineffective arms, shared controls, and addition of new treatments without compromising statistical properties. However, delays in observing primary outcomes may reduce these benefits. Investigators must either pause recruitment, delaying identification of...

💬 0 commentsarXiv:2609.08560v1PDF
0

Posted in stat.ME · 2026-09-08 · Zejing Zheng, Rui Huang, Junlong Zhao

Transfer Learning with Heterogeneous Feature Spaces in Linear Regression

Transfer learning improves target-task performance by leveraging related source data. Most methods assume shared feature spaces, yet in many applications, each source observes only a subset of target covariates. Classical imputation fails here due to block missingness, and standard imputation matrices are not optimized for target...

💬 0 commentsarXiv:2609.08526v1PDF
0

Posted in stat.AP · 2026-09-07 · Min-Ren Guan, Shen-Ning Tung

Pre-game paired-comparison modeling of professional League of Legends map outcomes

We build and evaluate a pre-game win-probability forecaster for individual maps (``games'') in professional \emph{League of Legends} (LoL). The proposed model is a one-stage logistic regression fit end-to-end on the win/loss log-loss: each team's exponentially-weighted moving average of past same-side results, a ridge-shrunk stable...

💬 0 commentsarXiv:2609.08060v1PDF
0

Posted in stat.AP · 2026-09-05 · Lei Liu

From Discrete Trailing Returns to a Continuous Graphical Profile: Return-to-Present Curves

Investment performance is commonly presented either as a conventional cumulative-return chart, which fixes a historical starting date and traces performance forward, or as a trailing-return table, which fixes the current endpoint but reports only a small set of prespecified horizons. These two displays have complementary limitations:...

💬 0 commentsarXiv:2609.06267v1PDF
0

Posted in stat.AP · 2026-09-04 · Jasmine Aherne, Ines Henriques-Cadby, Thomas House

Symptom clusters in Long COVID in the UK: prospective community-based cohort study using unsupervised machine learning

Long COVID is a condition usually defined by persisting symptoms following infection by the SARS-CoV-2 virus beyond the acute phase of infection. The condition has a significant impact on healthcare systems, the economy, and the individuals living with it. Due to the diverse and extensive symptomatology of long COVID, symptom...

💬 0 commentsarXiv:2609.05213v1PDF
0

Posted in stat.ML · 2026-09-04 · Chloé Hashimoto-Cullen, Ghislain Agoua, Benjamin Guedj, Sylvain Le Corff

PAC-Bayesian Reconstruction Guarantees for Time Series Variational Autoencoders

Forecasting time series accurately is critical for applications with complex data ranging from energy systems to healthcare and finance. Among current state of the art models, generative latent variable models are increasingly implemented; yet principled generalisation guarantees for modern latent variable models remain limited. In...

💬 0 commentsarXiv:2609.05212v1PDF
0

Posted in stat.ML · 2026-09-04 · Cassandra Durr, Alvaro Köhn-Luque, Chris Jewell, Lloyd A. C. Chapman

FluxDisco: Symbolic Regression for Stoichiometric Dynamical Systems via Monte Carlo Graph Search

Dynamical symbolic regression methods identify governing differential equations from noisy data, balancing interpretability and predictive accuracy. However, standard methods often produce expressions that violate known physical laws. To address this, we propose FluxDisco, a physics-informed framework tailored for flux-based,...

💬 0 commentsarXiv:2609.05207v1PDF
0

Posted in stat.ME · 2026-09-04 · Yi Xu, Xinye Chen, Sheng Jiang

Post-Corrected Raw-Score Martingale Posterior Sampling for von Mises-Fisher Models

We develop a finite-horizon calibration method for raw-score martingale posteriors, with von Mises--Fisher models as the main worked example. Starting from the maximum likelihood estimator, predictive paths are generated by simulating future observations from the current fitted model and updating the natural parameter by...

💬 0 commentsarXiv:2609.05096v1PDF
0

Posted in stat.ME · 2026-09-04 · Dennis Oestmann, Thorsten Dickhaus

On E-Backtesting: Generalizations and Sample Size Determination

We present an approach for determining sample sizes required to detect underestimations of the expected shortfall with a prescribed power when applying the recently proposed e-backtesting procedure. We consider scenarios in which the value-at-risk at level $p$ is always estimated correctly, while the difference between the true...

💬 0 commentsarXiv:2609.05089v1PDF
0

Posted in stat.ML · 2026-09-04 · Maximilian Fleissner, Debarghya Ghoshdastidar, Samory Kpotufe

An Analysis of Self-supervised Pre-training with Dependent Samples

Self-supervised learning relies on so-called data augmentations $φ(x)$ of unlabeled datapoints $x$ --- for example, masking random pixels in an image $x$ --- that should leave the label of $x$ invariant and are often used to learn a lower-complexity invariant subspace $\cal V$ for downstream tasks. In practice, such augmentations $\{...

💬 0 commentsarXiv:2609.05031v1PDF