Qwen Councils

Statistics

arXiv preprints from January 1, 2026 through September 20, 2026 — 16:51:55 EST

0

Posted in stat.ME · 2026-01-20 · Anqi Zhao, Peng Ding, Fan Li

Two-stage Least Squares with Clustered Data under the Local Average Treatment Effect Framework

To estimate the causal effect of an endogenous treatment using clustered data, the canonical two-stage least squares (2sls) estimates a linear regression of the outcome on treatment status using an instrumental variable (IV) and conducts inference with cluster-robust standard errors. When both the treatment and the IV vary within...

💬 0 commentsarXiv:2601.13507v3PDF
0

Posted in stat.AP · 2026-01-20 · Zhanshuo Ye, Yiming Hou, Rui Pan, Tianchen Gao, Hansheng Wang

Are Large Language Models able to Predict Highly Cited Papers? Evidence from Statistical Publications

Predicting highly-cited papers is a long-standing challenge due to the complex interactions of research content, scholarly communities, and temporal dynamics. Recent advances in large language models (LLMs) raise the question of whether early-stage textual information can provide useful signals of long-term scientific impact. Focusing...

💬 0 commentsarXiv:2601.13627v1PDF
0

Posted in stat.AP · 2026-01-20 · Li Tuobang

On the Anchoring Effect of Monetary Policy on the Labor Share of Income and the Rationality of Its Setting Mechanism

Modern macroeconomic monetary theory suggests that the labor share of income has effectively become a core macroe-conomic parameter anchored by top policymakers through Open Market Operations (OMO). However, the setting of this parameter remains a subject of intense economic debate. This paper provides a detailed summary of these...

💬 0 commentsarXiv:2601.13675v2PDF
0

Posted in stat.ML · 2026-01-20 · Yuchen Jiao, Jiin Woo, Gen Li, Gauri Joshi, Yuejie Chi

Sample Complexity of Average-Reward Q-Learning: From Single-agent to Federated Reinforcement Learning

Average-reward reinforcement learning offers a principled framework for long-term decision-making by maximizing the mean reward per time step. Although Q-learning is a widely used model-free algorithm with established sample complexity in discounted and finite-horizon Markov decision processes (MDPs), its theoretical guarantees for...

💬 0 commentsarXiv:2601.13642v1PDF
0

Posted in stat.AP · 2026-01-20 · Shuvayan Banerjee, Radhendushka Srivastava, James Saunderson, Ajit Rajwade

Correction of Pooling Matrix Mis-specifications in Compressed Sensing Based Group Testing

Compressed sensing, which involves the reconstruction of sparse signals from an under-determined linear system, has been recently used to solve problems in group testing. In a public health context, group testing aims to determine the health status values of p subjects from n<<p pooled tests, where a pool is defined as a mixture of...

💬 0 commentsarXiv:2601.13641v1PDF
0

Posted in stat.ME · 2026-01-20 · Sonja Zehetmayer, Marta Bofill Roig, Fabrice Lotola Mougeni, Sabine Specht, Marc P. Hübner, Martin Posch

An Adaptive Phase II Trial Design for Dose Selection and Addition in Microfilarial Infections

We propose a frequentist adaptive phase 2 trial design to evaluate the safety and efficacy of three treatment regimens (doses) compared to placebo for four types of helminth (worm) infections. This trial will be carried out in four Subsaharan African countries from spring 2025. Since the safety of the highest dose is not yet...

💬 0 commentsarXiv:2601.13784v1PDF
0

Posted in stat.ME · 2026-01-20 · Tiejun Tong, Hongmei Lin, Bowen Gang, Riquan Zhang

ChauBoxplot and AdaptiveBoxplot: Two R packages for boxplot-based outlier detection

Tukey's boxplot is widely used for outlier detection; however, its classic fixed-fence rule tends to flag an excessive number of outliers as the sample size grows. To address this, we introduce two new R packages, ChauBoxplot and AdaptiveBoxplot, which implement more robust and statistically principled outlier detection methods. We...

💬 0 commentsarXiv:2601.13759v2PDF
0

Posted in stat.ME · 2026-01-20 · Dan Chaltiel, Alexis Cochard, Nusaibah Ibrahimi, Charlotte Bargain, Ikram Benchara, Anne Lourdessamy, Aldéric Fraslin, Matthieu Texier, Livia Pierotti

Building a Standardised Statistical Reporting Toolbox in an Academic Oncology Clinical Trials Unit: The grstat R Package

Academic Clinical Trial Units frequently face fragmented statistical workflows, leading to duplicated effort, limited collaboration, and inconsistent analytical practices. To address these challenges within an oncology Clinical Trial Unit, we developed grstat, an R package providing a standardised set of tools for routine statistical...

💬 0 commentsarXiv:2601.13755v1PDF
0

Posted in stat.ML · 2026-01-20 · Shijie Zhong, Yikun Yang, Da Gong, Jiangfeng Fu

Unified Unbiased Variance Estimation for Maximum Mean Discrepancy: Robust Finite-Sample Performance with Imbalanced Data and Exact Acceleration under Null and Alternative Hypotheses

The maximum mean discrepancy (MMD) is a kernel-based nonparametric statistic for two-sample testing, whose inferential accuracy depends critically on variance characterization. Existing work provides various finite-sample estimators of the MMD variance, often differing under the null and alternative hypotheses and across balanced or...

💬 0 commentsarXiv:2601.13874v2PDF
0

Posted in stat.ME · 2026-01-20 · Elena Dumitrescu, Julien Peignon, Arthur Thomas

Tail-Aware Density Forecasting of Locally Explosive Time Series: A Neural Network Approach

This paper proposes a Mixture Density Network specifically designed for forecasting time series that exhibit locally explosive behavior. By incorporating skewed t-distributions as mixture components, our approach offers enhanced flexibility in capturing the skewed, heavy-tailed, and potentially multimodal nature of predictive...

💬 0 commentsarXiv:2601.14049v2PDF
0

Posted in stat.ML · 2026-01-20 · Stefano Damato, Nicolò Rubattu, Dario Azzimonti, Giorgio Corani

Intermittent time series forecasting: local vs global models

Forecasting intermittent time series, which contain zeros, is a crucial challenge in supply chains as inventory policies require probabilistic forecasts to establish safety levels. Intermittent time series are commonly forecast using local models, trained individually on each time series. In the last years global models, trained on a...

💬 0 commentsarXiv:2601.14031v2PDF
0

Posted in stat.ME · 2026-01-20 · Prajamitra Bhuyan, Soutik Halder, Jayant Jha

Modeling Zero-Inflated Longitudinal Circular Data Using Bayesian Methods: Application to Ophthalmology

This paper introduces the modeling of circular data with excess zeros under a longitudinal framework, where the response is a circular variable and the covariates can be both linear and circular in nature. In the literature, various circular-circular and circular-linear regression models have been studied and applied to different...

💬 0 commentsarXiv:2601.13998v1PDF
0

Posted in stat.ME · 2026-01-20 · Taehee Lee, Jun S. Liu

Factor Analysis of Multivariate Stochastic Volatility Model

Modeling the time-varying covariance structures of high-dimensional variables is critical across diverse scientific and industrial applications; however, existing approaches exhibit notable limitations in either modeling flexibility or inferential efficiency. For instance, change-point modeling fails to account for the continuous...

💬 0 commentsarXiv:2601.14199v1PDF
0

Posted in stat.ME · 2026-01-20 · Xi Fang, Bingkai Wang, Guangyu Tong, Liangyuan Hu, Shuangge Ma, Fan Li

Doubly robust estimators of the restricted mean time in favor estimands in individual- and cluster-randomized trials

Progressive multi-state survival outcomes are common in trials with recurrent or sequential events and require treatment effect estimands that remain interpretable without proportional intensity or Markov assumptions. The restricted mean time in favor of treatment (RMT-IF) extends the restricted mean survival time to ordered...

💬 0 commentsarXiv:2601.14431v1PDF
0

Posted in stat.ML · 2026-01-20 · Peter Potaptchik, Adhi Saravanan, Abbas Mammadov, Alvaro Prat, Michael S. Albergo, Yee Whye Teh

Meta Flow Maps enable scalable reward alignment

Controlling generative models is computationally expensive. This is because optimal alignment with a reward function--whether via inference-time steering or fine-tuning--requires estimating the value function. This task demands access to the conditional posterior $p_{1|t}(x_1|x_t)$, the distribution of clean data $x_1$ consistent with...

💬 0 commentsarXiv:2601.14430v2PDF
0

Posted in stat.ML · 2026-01-20 · Zhengang Zhong, Yury Korolev, Matthew Thorpe

Large Data Limits of Laplace Learning for Gaussian Measure Data in Infinite Dimensions

Laplace learning is a semi-supervised method, a solution for finding missing labels from a partially labeled dataset utilizing the geometry given by the unlabeled data points. The method minimizes a Dirichlet energy defined on a (discrete) graph constructed from the full dataset. In finite dimensions the asymptotics in the large...

💬 0 commentsarXiv:2601.14515v1PDF
0

Posted in stat.ME · 2026-01-20 · Marlena Bannick, Yuanyuan Bian, Gregory Chen, Liming Li, Yuhan Qian, Daniel Sabanés Bové, Dong Xi, Ting Ye, Yanyao Yi

The RobinCar Family: R Tools for Robust Covariate Adjustment in Randomized Clinical Trials

Purpose: Covariate adjustment is a powerful statistical technique that can increase efficiency in clinical trials. Recent guidance from the U.S. FDA provided recommendations and best practices for using covariate adjustment. However, there has existed a gap between the extensive statistical literature on covariate adjustment and...

💬 0 commentsarXiv:2601.14498v1PDF
0

Posted in stat.AP · 2026-01-19 · Md Muhtasim Munif Fahim, Md Jahid Hasan Imran, Md. Naim Molla, Luknath Debnath, Tonmoy Shil, Ehsanul Bashar Pranto, Md Mostafizur Rahman Likhon, Md Shafin Sanyan Saad, Md. Rezaul Karim

Drivers, Receivers, and Dynamic Linkages: The Directed Structure of SDG Interdependence, 2000--2024

Governments with limited fiscal and administrative capacity need to know which Sustainable Development Goals (SDGs) propagate progress through the goal system and how quickly. We map the directed interdependence structure of all seventeen goals using a balanced panel of 114 countries observed annually from 2000 to 2024. The goal...

💬 0 commentsarXiv:2601.20875v2PDF
0

Posted in stat.ME · 2026-01-19 · Esteban Fernández-Morales, Emily M. Ko, Nandita Mitra, Youjin Lee, Arman Oganisian

A Bayesian framework for cost-effectiveness analysis with time-varying treatment decisions

Cost-effectiveness analyses (CEAs) compare the costs and health outcomes of treatment regimes to inform medical decisions. With observational claims data, CEAs must address nonrandom treatment assignment, administrative censoring, and irregularly spaced medical visits that reflect the continuous timing of care and treatment...

💬 0 commentsarXiv:2601.14309v1PDF
0

Posted in stat.ME · 2026-01-19 · Beniamino Hadj-Amar, Jack Jewson

Bayesian Variable Selection with the Quasi-Posterior

The Bayesian approach provides powerful methods for variable selection. The ability to incorporate sparsity through prior beliefs and account for parameter uncertainty allows Bayesian variable selection to consistently identify which of the variables are active and exhibit strong finite-sample performance. However, Bayesian methods...

💬 0 commentsarXiv:2601.12767v2PDF
0

Posted in stat.AP · 2026-01-19 · Giovanni Bocchi, Alessandra Micheletti, Paolo Nota, Alessandro Olper

The impact of abnormal temperatures on crop yields in Italy: a functional quantile regression approach

In this study, we apply functional regression analysis to identify the specific within-season periods during which temperature and precipitation anomalies most affect crop yields. Using provincial data for Italy from 1952 to 2023, we analyze two major cereals, maize and soft wheat, and quantify how abnormal weather conditions...

💬 0 commentsarXiv:2601.12864v1PDF
0

Posted in stat.ME · 2026-01-19 · Jale Basten, Katja Ickstadt, Nina Timmesfeld

Guidance for Addressing Individual Time Effects in Cohort Stepped Wedge Cluster Randomized Trials: A Simulation Study

Background: Stepped wedge cluster randomized trials (SW-CRTs) involve sequential measurements within clusters over time. Initially, all clusters start in the control condition before crossing over to the intervention on a staggered schedule. In cohort designs, secular trends, cluster-level changes, and individual-level changes (e.g.,...

💬 0 commentsarXiv:2601.12930v1PDF
0

Posted in stat.ME · 2026-01-19 · Siyu Heng, Yanxin Shen, Zijian Guo

Propensity Score Propagation: A General Framework for Design-Based Inference with Unknown Propensity Scores

Design-based inference, also known as randomization-based or finite-population inference, provides a principled framework for trustworthy statistical inference. It attributes randomness solely to the design mechanism, such as treatment assignment, survey sampling, or missingness, without imposing super-population distributional or...

💬 0 commentsarXiv:2601.13150v4PDF
0

Posted in stat.ML · 2026-01-19 · Davidson Lova Razafindrakoto, Alain Celisse, Jérôme Lacaille

Approximate full conformal prediction in an RKHS

Full conformal prediction is a framework that implicitly formulates distribution-free confidence prediction regions for a wide range of estimators. However, a classical limitation of the full conformal framework is the computation of the confidence prediction regions, which is usually impossible since it requires training infinitely...

💬 0 commentsarXiv:2601.13102v3PDF
0

Posted in stat.ML · 2026-01-19 · Francisco Daunas, Iñaki Esnaola, Samir M. Perlaza, H. Vincent Poor

Empirical Risk Minimization with $f$-Divergence Regularization

In this paper, the solution to the empirical risk minimization problem with $f$-divergence regularization (ERM-$f$DR) is presented and conditions under which the solution also serves as the solution to the minimization of the expected empirical risk subject to an $f$-divergence constraint are established. The proposed approach extends...

💬 0 commentsarXiv:2601.13191v1PDF