Qwen Councils

Statistics

arXiv preprints from January 1, 2026 through July 21, 2026 — 11:53:40 EST

0

Posted in stat.ML · 2026-01-21 · Felix Schur, Niklas Pfister, Peng Ding, Sach Mukherjee, Jonas Peters

Many Experiments, Few Repetitions, Unpaired Data, and Sparse Effects: Is Causal Inference Possible?

We study the problem of estimating causal effects under hidden confounding in the following unpaired data setting: we observe some covariates $X$ and an outcome $Y$ under different experimental conditions (environments) but do not observe them jointly; we either observe $X$ or $Y$. Under appropriate regularity conditions, the problem...

💬 0 commentsarXiv:2601.15254v1PDF
0

Posted in stat.ML · 2026-01-21 · Kexin Wang, Salil Bhate, João M. Pereira, Joe Kileel, Matylda Figlerowicz, Anna Seigal

Multi-context principal component analysis

Principal component analysis (PCA) is a tool to capture factors that explain variation in data. Across domains, data are now collected across multiple contexts (for example, individuals with different diseases, cells of different types, or words across texts). While the factors explaining variation in data are undoubtedly shared...

💬 0 commentsarXiv:2601.15239v1PDF
0

Posted in stat.ME · 2026-01-21 · Diptanil Santra, Guanhua Chen, Chan Park

Distributional Balancing for Causal Inference: A Unified Framework via Characteristic Function Distance

Weighting methods are essential tools for estimating causal effects in observational studies, with the goal of balancing pre-treatment covariates across treatment groups. Traditional approaches pursue this objective indirectly, for example, via inverse propensity score weighting or by matching a finite number of covariate moments, and...

💬 0 commentsarXiv:2601.15449v2PDF
0

Posted in stat.AP · 2026-01-21 · Shome Chakraborty, Fardil Khan, Soutik Ghosal

Assessing the informative value of macroeconomic indicators for public health forecasting

Macroeconomic conditions influence the environments in which health systems operate, yet their value as leading signals of health system capacity has not been systematically evaluated. In this study, we examine whether selected macroeconomic indicators contain predictive information for several capacity-related public health targets,...

💬 0 commentsarXiv:2601.15514v1PDF
0

Posted in stat.ML · 2026-01-21 · Saptarshi Roy, Alessandro Rinaldo, Purnamrita Sarkar

Low-Dimensional Adaptation of Rectified Flow: A Diffusion and Stochastic Localization Perspective

In recent years, Rectified flow (RF) has gained considerable popularity largely due to its generation efficiency and state-of-the-art performance. In this paper, we investigate the degree to which RF automatically adapts to the intrinsic low dimensionality of the support of the target distribution to accelerate sampling. We show that,...

💬 0 commentsarXiv:2601.15500v3PDF
0

Posted in stat.AP · 2026-01-21 · Laura Medialdea, Ana Arribas-Gil, Álvaro Pérez-Romero, Amador Gómez

Geometric Morphometrics approach for classifying children's nutritional status on out of sample data

Current alignment-based methods for classification in geometric morphometrics do not generally address the classification of new individuals that were not part of the study sample. However, in the context of infant and child nutritional assessment from body shape images this is a relevant problem. In this setting, classification rules...

💬 0 commentsarXiv:2601.15491v1PDF
0

Posted in stat.OT · 2026-01-21 · Heather Battey, Charlotte Edgar

Treatment effect: a critique

Two broad positions within statistics define a treatment effect, on the one hand, as a parameter of a statistical model, and on the other, as an appropriate population-level difference in outcomes or counterfactual outcomes under the different treatment regimes. This short expository paper presents some simple but consequential...

💬 0 commentsarXiv:2601.15467v1PDF
0

Posted in stat.ME · 2026-01-20 · Haidong Lu, Fan Li, Laine E. Thomas, Fan Li

What is Overlap Weighting, How Has it Evolved, and When to Use It for Causal Inference?

The growing availability of large health databases has expanded the use of observational studies for comparative effectiveness research. Unlike randomized trials, observational studies must adjust for systematic differences in patient characteristics between treatment groups. Propensity score methods, including matching, weighting,...

💬 0 commentsarXiv:2601.13535v1PDF
0

Posted in stat.ML · 2026-01-20 · Wenzhi Gao, Chang He, Madeleine Udell

Small Gradient Norm Regret for Online Convex Optimization

This paper introduces a new problem-dependent regret measure for online convex optimization with smooth losses. The notion, which we call the $G^\star$ regret, depends on the cumulative squared gradient norm evaluated at the decision in hindsight. We show that the $G^\star$ regret strictly refines the existing $L^\star$ (small loss)...

💬 0 commentsarXiv:2601.13519v3PDF
0

Posted in stat.ME · 2026-01-20 · Ronan Perry, Snigdha Panigrahi, Daniela Witten

Post-selection inference for penalized M-estimators via score thinning

We consider inference for M-estimators after model selection using a sparsity-inducing penalty. While existing methods for this task require bespoke inference procedures, we propose a simpler approach, which relies on two insights: (i) adding and subtracting carefully-constructed noise to a Gaussian random variable with unknown mean...

💬 0 commentsarXiv:2601.13514v1PDF
0

Posted in stat.ME · 2026-01-20 · Anqi Zhao, Peng Ding, Fan Li

Two-stage Least Squares with Clustered Data under the Local Average Treatment Effect Framework

To estimate the causal effect of an endogenous treatment using clustered data, the canonical two-stage least squares (2sls) estimates a linear regression of the outcome on treatment status using an instrumental variable (IV) and conducts inference with cluster-robust standard errors. When both the treatment and the IV vary within...

💬 0 commentsarXiv:2601.13507v3PDF
0

Posted in stat.AP · 2026-01-20 · Zhanshuo Ye, Yiming Hou, Rui Pan, Tianchen Gao, Hansheng Wang

Are Large Language Models able to Predict Highly Cited Papers? Evidence from Statistical Publications

Predicting highly-cited papers is a long-standing challenge due to the complex interactions of research content, scholarly communities, and temporal dynamics. Recent advances in large language models (LLMs) raise the question of whether early-stage textual information can provide useful signals of long-term scientific impact. Focusing...

💬 0 commentsarXiv:2601.13627v1PDF
0

Posted in stat.AP · 2026-01-20 · Li Tuobang

On the Anchoring Effect of Monetary Policy on the Labor Share of Income and the Rationality of Its Setting Mechanism

Modern macroeconomic monetary theory suggests that the labor share of income has effectively become a core macroe-conomic parameter anchored by top policymakers through Open Market Operations (OMO). However, the setting of this parameter remains a subject of intense economic debate. This paper provides a detailed summary of these...

💬 0 commentsarXiv:2601.13675v2PDF
0

Posted in stat.ML · 2026-01-20 · Yuchen Jiao, Jiin Woo, Gen Li, Gauri Joshi, Yuejie Chi

Sample Complexity of Average-Reward Q-Learning: From Single-agent to Federated Reinforcement Learning

Average-reward reinforcement learning offers a principled framework for long-term decision-making by maximizing the mean reward per time step. Although Q-learning is a widely used model-free algorithm with established sample complexity in discounted and finite-horizon Markov decision processes (MDPs), its theoretical guarantees for...

💬 0 commentsarXiv:2601.13642v1PDF
0

Posted in stat.AP · 2026-01-20 · Shuvayan Banerjee, Radhendushka Srivastava, James Saunderson, Ajit Rajwade

Correction of Pooling Matrix Mis-specifications in Compressed Sensing Based Group Testing

Compressed sensing, which involves the reconstruction of sparse signals from an under-determined linear system, has been recently used to solve problems in group testing. In a public health context, group testing aims to determine the health status values of p subjects from n<<p pooled tests, where a pool is defined as a mixture of...

💬 0 commentsarXiv:2601.13641v1PDF
0

Posted in stat.ME · 2026-01-20 · Sonja Zehetmayer, Marta Bofill Roig, Fabrice Lotola Mougeni, Sabine Specht, Marc P. Hübner, Martin Posch

An Adaptive Phase II Trial Design for Dose Selection and Addition in Microfilarial Infections

We propose a frequentist adaptive phase 2 trial design to evaluate the safety and efficacy of three treatment regimens (doses) compared to placebo for four types of helminth (worm) infections. This trial will be carried out in four Subsaharan African countries from spring 2025. Since the safety of the highest dose is not yet...

💬 0 commentsarXiv:2601.13784v1PDF
0

Posted in stat.ME · 2026-01-20 · Tiejun Tong, Hongmei Lin, Bowen Gang, Riquan Zhang

ChauBoxplot and AdaptiveBoxplot: Two R packages for boxplot-based outlier detection

Tukey's boxplot is widely used for outlier detection; however, its classic fixed-fence rule tends to flag an excessive number of outliers as the sample size grows. To address this, we introduce two new R packages, ChauBoxplot and AdaptiveBoxplot, which implement more robust and statistically principled outlier detection methods. We...

💬 0 commentsarXiv:2601.13759v2PDF
0

Posted in stat.ME · 2026-01-20 · Dan Chaltiel, Alexis Cochard, Nusaibah Ibrahimi, Charlotte Bargain, Ikram Benchara, Anne Lourdessamy, Aldéric Fraslin, Matthieu Texier, Livia Pierotti

Building a Standardised Statistical Reporting Toolbox in an Academic Oncology Clinical Trials Unit: The grstat R Package

Academic Clinical Trial Units frequently face fragmented statistical workflows, leading to duplicated effort, limited collaboration, and inconsistent analytical practices. To address these challenges within an oncology Clinical Trial Unit, we developed grstat, an R package providing a standardised set of tools for routine statistical...

💬 0 commentsarXiv:2601.13755v1PDF
0

Posted in stat.ML · 2026-01-20 · Shijie Zhong, Yikun Yang, Da Gong, Jiangfeng Fu

Unified Unbiased Variance Estimation for Maximum Mean Discrepancy: Robust Finite-Sample Performance with Imbalanced Data and Exact Acceleration under Null and Alternative Hypotheses

The maximum mean discrepancy (MMD) is a kernel-based nonparametric statistic for two-sample testing, whose inferential accuracy depends critically on variance characterization. Existing work provides various finite-sample estimators of the MMD variance, often differing under the null and alternative hypotheses and across balanced or...

💬 0 commentsarXiv:2601.13874v2PDF
0

Posted in stat.ME · 2026-01-20 · Elena Dumitrescu, Julien Peignon, Arthur Thomas

Tail-Aware Density Forecasting of Locally Explosive Time Series: A Neural Network Approach

This paper proposes a Mixture Density Network specifically designed for forecasting time series that exhibit locally explosive behavior. By incorporating skewed t-distributions as mixture components, our approach offers enhanced flexibility in capturing the skewed, heavy-tailed, and potentially multimodal nature of predictive...

💬 0 commentsarXiv:2601.14049v2PDF
0

Posted in stat.ML · 2026-01-20 · Stefano Damato, Nicolò Rubattu, Dario Azzimonti, Giorgio Corani

Intermittent time series forecasting: local vs global models

Forecasting intermittent time series, which contain zeros, is a crucial challenge in supply chains as inventory policies require probabilistic forecasts to establish safety levels. Intermittent time series are commonly forecast using local models, trained individually on each time series. In the last years global models, trained on a...

💬 0 commentsarXiv:2601.14031v2PDF
0

Posted in stat.ME · 2026-01-20 · Prajamitra Bhuyan, Soutik Halder, Jayant Jha

Modeling Zero-Inflated Longitudinal Circular Data Using Bayesian Methods: Application to Ophthalmology

This paper introduces the modeling of circular data with excess zeros under a longitudinal framework, where the response is a circular variable and the covariates can be both linear and circular in nature. In the literature, various circular-circular and circular-linear regression models have been studied and applied to different...

💬 0 commentsarXiv:2601.13998v1PDF
0

Posted in stat.ME · 2026-01-20 · Taehee Lee, Jun S. Liu

Factor Analysis of Multivariate Stochastic Volatility Model

Modeling the time-varying covariance structures of high-dimensional variables is critical across diverse scientific and industrial applications; however, existing approaches exhibit notable limitations in either modeling flexibility or inferential efficiency. For instance, change-point modeling fails to account for the continuous...

💬 0 commentsarXiv:2601.14199v1PDF
0

Posted in stat.ME · 2026-01-20 · Xi Fang, Bingkai Wang, Guangyu Tong, Liangyuan Hu, Shuangge Ma, Fan Li

Doubly robust estimators of the restricted mean time in favor estimands in individual- and cluster-randomized trials

Progressive multi-state survival outcomes are common in trials with recurrent or sequential events and require treatment effect estimands that remain interpretable without proportional intensity or Markov assumptions. The restricted mean time in favor of treatment (RMT-IF) extends the restricted mean survival time to ordered...

💬 0 commentsarXiv:2601.14431v1PDF
0

Posted in stat.ML · 2026-01-20 · Peter Potaptchik, Adhi Saravanan, Abbas Mammadov, Alvaro Prat, Michael S. Albergo, Yee Whye Teh

Meta Flow Maps enable scalable reward alignment

Controlling generative models is computationally expensive. This is because optimal alignment with a reward function--whether via inference-time steering or fine-tuning--requires estimating the value function. This task demands access to the conditional posterior $p_{1|t}(x_1|x_t)$, the distribution of clean data $x_1$ consistent with...

💬 0 commentsarXiv:2601.14430v2PDF