Qwen Councils

Statistics

arXiv preprints from January 1, 2026 through September 20, 2026 — 18:31:11 EST

0

Posted in stat.ME · 2026-01-17 · Yang Lu, Nandini Dendukuri

Using Directed Acyclic Graphs to Illustrate Common Biases in Diagnostic Test Accuracy Studies

Background: Diagnostic test accuracy (DTA) studies, like etiological studies, are susceptible to various biases including reference standard error bias, partial verification bias, spectrum effect, confounding, and bias from misassumption of conditional independence. While directed acyclic graphs (DAGs) are widely used in etiological...

💬 0 commentsarXiv:2601.12167v1PDF
0

Posted in stat.ME · 2026-01-17 · Danielle Tsao, Krikamol Muandet, Frederick Eberhardt, Emilija Perković

Lost in Aggregation: The Causal Interpretation of the IV Estimand

Instrumental variable based estimation of a causal effect has emerged as a standard approach to mitigate confounding bias in the social sciences and epidemiology, where conducting randomized experiments can be too costly or impossible. However, justifying the validity of the instrument often poses a significant challenge. In this...

💬 0 commentsarXiv:2601.12120v1PDF
0

Posted in stat.AP · 2026-01-16 · Michael T. Gorczyca

A Note on Harmonic Underspecification in Log-Normal Trigonometric Regression

Analysis of biological rhythm data often involves performing least squares trigonometric regression, which models the oscillations of a response over time as a sum of sinusoidal components. When the response is not normally distributed, an investigator will either transform the response before applying least squares trigonometric...

💬 0 commentsarXiv:2601.10919v1PDF
0

Posted in stat.ML · 2026-01-16 · Fenglin Zhang, Jie Wang

Contextual Distributionally Robust Optimization with Causal and Continuous Structure: An Interpretable and Tractable Approach

In this paper, we introduce a framework for contextual distributionally robust optimization (DRO) that considers the causal and continuous structure of the underlying distribution by developing interpretable and tractable decision rules that prescribe decisions using covariates. We first introduce the causal Sinkhorn discrepancy...

💬 0 commentsarXiv:2601.11016v2PDF
0

Posted in stat.ME · 2026-01-16 · Xiaojing Sun, Bingxin Zhao, Fei Xue

Generalized Heterogeneous Functional Model with Applications to Large-scale Mobile Health Data

Physical activity is crucial for human health. With the increasing availability of large-scale mobile health data, strong associations have been found between physical activity and various diseases. However, accurately capturing this complex relationship is challenging, possibly because it varies across different subgroups of...

💬 0 commentsarXiv:2601.10994v1PDF
0

Posted in stat.ML · 2026-01-16 · Minseo Kang, Seunghwan Park, Dongha Kim

Memorize Early, Then Query: Inlier-Memorization-Guided Active Outlier Detection

Outlier detection (OD) aims to identify abnormal instances, known as outliers or anomalies, by learning typical patterns of normal data, or inliers. Performing OD under an unsupervised regime-without any information about anomalous instances in the training data-is challenging. A recently observed phenomenon, known as the...

💬 0 commentsarXiv:2601.10993v2PDF
0

Posted in stat.AP · 2026-01-16 · Shi Feng, B. Brian Park, Andrew Mondschein

Analyzing Residential Speeding Using Connected Vehicle Data: A Case Study in Charlottesville, VA Area

This study uses connected vehicle data to analyze speeding behavior on residential roads. A scalable pipeline processes trajectory data and supplements missing speed limits to generate summaries at OpenStreetMap's way ID level. The findings reveal a highly skewed distribution of both aggressive and reckless speeding. Based on a case...

💬 0 commentsarXiv:2601.10974v1PDF
0

Posted in stat.ME · 2026-01-16 · Soma Nikai, Yuichi Goto, Koji Tsukuda

Robust $M$-Estimation of Scatter Matrices via Precision Structure Shrinkage

Maronna's and Tyler's $M$-estimators are among the most widely used robust estimators for scatter matrices. However, when the dimension of observations is relatively high, their performance can substantially deteriorate in certain situations, particularly in the presence of clustered outliers. To address this issue, we propose an...

💬 0 commentsarXiv:2601.11099v2PDF
0

Posted in stat.ML · 2026-01-16 · Hangjin Jiang, Yuzhou Li, Zhaoxing Gao

Split-and-Conquer: Distributed Factor Modeling for High-Dimensional Matrix-Variate Time Series

In this paper, we propose a distributed framework for reducing the dimensionality of high-dimensional, large-scale, heterogeneous matrix-variate time series data using a factor model. The data are first partitioned column-wise (or row-wise) and allocated to node servers, where each node estimates the row (or column) loading matrix via...

💬 0 commentsarXiv:2601.11091v1PDF
0

Posted in stat.CO · 2026-01-16 · Sebastiano Grazzi, Sifan Liu, Gareth O. Roberts, Jun Yang

Sub-Cauchy Sampling: Escaping the Dark Side of the Moon

We introduce a Markov chain Monte Carlo algorithm based on Sub-Cauchy Projection, a geometric transformation that generalizes stereographic projection by mapping Euclidean space into a spherical cap of a hyper-sphere, referred to as the complement of the dark side of the moon. We prove that our proposed method is uniformly ergodic for...

💬 0 commentsarXiv:2601.11066v1PDF
0

Posted in stat.ME · 2026-01-16 · Michael C. Sachs, Erin E. Gabriel, Robin J. Evans, Arvid Sjölander

Deriving Complete Constraints in Hidden Variable Models

Hidden variable graphical models can sometimes imply constraints on the observable distribution that are more complex than simple conditional independence relations. These observable constraints can falsify assumptions of the model that would otherwise be untestable due to the unobserved variables and can be used to constrain...

💬 0 commentsarXiv:2601.11242v3PDF
0

Posted in stat.ME · 2026-01-16 · Pierre Alquier, Jean-David Fermanian, Benjamin Poignard

Estimation of time series by Maximum Mean Discrepancy

We define two minimum distance estimators for dependent data by minimizing some approximated Maximum Mean Discrepancy distances between the true empirical distribution of observations and their assumed (parametric) model distribution. When the latter one is intractable, it is approximated by simulation, allowing to accommodate most...

💬 0 commentsarXiv:2601.11233v1PDF
0

Posted in stat.ME · 2026-01-16 · Yuki Toyoda

ThSQCA: Threshold-Sweep Qualitative Comparative Analysis in R

Qualitative Comparative Analysis (QCA) requires researchers to choose calibration and dichotomization thresholds, and these choices can substantially affect truth tables, minimization, and resulting solution formulas. Despite this dependency, threshold sensitivity is often examined only in an ad hoc manner because repeated analyses...

💬 0 commentsarXiv:2601.11229v4PDF
0

Posted in stat.CO · 2026-01-16 · Radhika Kulkarni, Aluisio Pinheiro, Brani Vidakovic, Abdourrahmane M. Atto

Smooth SCAD: A Raised Cosine SCAD Type Thresholding Rule for Wavelet Denoising

We introduce a smooth variant of the SCAD thresholding rule for wavelet denoising by replacing its piecewise linear transition with a raised cosine. The resulting shrinkage function is odd, continuous on R, and continuously differentiable away from the main threshold, yet retains the hallmark SCAD properties of sparsity for small...

💬 0 commentsarXiv:2601.11461v1PDF
0

Posted in stat.CO · 2026-01-16 · Yiping Hong, Sameh Abdulah, Marc G. Genton, Ying Sun

Fisher Scoring for Exact Matérn Covariance Estimation through Stable Smoothness Optimization

Gaussian Random Fields (GRFs) with Matérn covariance functions have emerged as a powerful framework for modeling spatial processes due to their flexibility in capturing different features of the spatial field. However, the smoothness parameter is challenging to estimate using maximum likelihood estimation (MLE), which involves...

💬 0 commentsarXiv:2601.11437v1PDF
0

Posted in stat.ME · 2026-01-16 · Ricardo J. Sandoval, Sivaraman Balakrishnan, Avi Feller, Michael I. Jordan, Ian Waudby-Smith

On Nonasymptotic Confidence Intervals for Treatment Effects in Randomized Experiments

We study nonasymptotic (finite-sample) confidence intervals for treatment effects in randomized experiments. In the existing literature, the effective sample sizes of nonasymptotic confidence intervals tend to be looser than the corresponding central-limit-theorem-based confidence intervals by a factor depending on the square root of...

💬 0 commentsarXiv:2601.11744v2PDF
0

Posted in stat.ME · 2026-01-16 · Xinlei Xu, Caitlin H Daly, Audrey Béliveau

Identifying Conditions Favouring Multiplicative Heterogeneity Models in Network Meta-Analysis

Explicit modelling of between-study heterogeneity is essential in network meta-analysis (NMA) to ensure valid inference and avoid overstating precision. While the additive random-effects (RE) model is the conventional approach, the multiplicative-effect (ME) model remains underexplored. The ME model inflates within-study variances by...

💬 0 commentsarXiv:2601.11735v2PDF
0

Posted in stat.ML · 2026-01-16 · Guerlain Lambert, Céline Helbert, Claire Lauvernet

Gradient-based Active Learning with Gaussian Processes for Global Sensitivity Analysis

Global sensitivity analysis of complex numerical simulators is often limited by the small number of model evaluations that can be afforded. In such settings, surrogate models built from a limited set of simulations can substantially reduce the computational burden, provided that the design of computer experiments is enriched...

💬 0 commentsarXiv:2601.11790v1PDF
0

Posted in stat.ME · 2026-01-15 · Yang Ou, Lan Xue, Carmen Tekwe, Kedir N. Turi, Roger S. Zoh

Estimating the effect of lymphovascular invasion on 2-year survival probability under endogeneity: a recursive copula-based approach

Lymphovascular invasion (LVI) is an important prognostic marker for head and neck squamous cell carcinoma (HNSC), but the true effect of LVI on survival may be distorted by endogeneity arising from unmeasured confounding. Conventional one-stage conditional models and instrument-based two-stage estimators are prone to bias under...

💬 0 commentsarXiv:2601.09984v1PDF
0

Posted in stat.ME · 2026-01-15 · Faruk Muritala, Austin Brown, Dhrubajyoti Ghosh, Sherry Ni

Derivations for the Cumulative Standardized Binomial EWMA (CSB-EWMA) Control Chart

This paper presents the exact mathematical derivation of the mean and variance properties for the Exponentially Weighted Moving Average (EWMA) statistic applied to binomial proportion monitoring in Multiple Stream Processes (MSPs). We develop a Cumulative Standardized Binomial EWMA (CSB-EWMA) formulation that provides adaptive control...

💬 0 commentsarXiv:2601.09968v1PDF
0

Posted in stat.ME · 2026-01-15 · Lei Huang, Chengyue Liu, Li Wang

Weighted least squares estimation by multivariate-dependent weights for linear regression models

Multivariate linear regression models often face the problem of heteroscedasticity caused by multiple explanatory variables. The weighted least squares estimation with univariate-dependent weights has limitations in constructing weight functions. Therefore, this paper proposes a multivariate dependent weighted least squares estimation...

💬 0 commentsarXiv:2601.10049v1PDF
0

Posted in stat.AP · 2026-01-15 · Peter Maurice Catt

An Information-Theoretic Diagnostic Analytics Framework for Mapping Past-Future Dependence in Horizon-Specific Forecastability

In many systems, the true data-generating process is unknown, requiring forecasters to rely on observed time series. This study proposes a pre-modeling diagnostic framework for horizon-specific forecastability assessment that evaluates forecastability before model selection begins. Forecastability is operationalized using auto-mutual...

💬 0 commentsarXiv:2601.10006v4PDF
0

Posted in stat.ME · 2026-01-15 · Mayukh Choudhury, Debraj Das, Sujit Ghosh

Asymptotic Theory of Tail Dependence and Bootstrap for Checkerboard Copulas

A comprehensive asymptotic and bootstrap theory is established for checkerboard-based estimation of the copula and its lower and upper tail copula counterparts under unknown marginal distributions. The proposed estimator of the tail copula extends a local bilinear interpolation of the empirical copula to the tail region, providing a...

💬 0 commentsarXiv:2601.10252v3PDF
0

Posted in stat.ME · 2026-01-15 · Yue Yu, Guanghui Wang, Liu Liu, Changliang Zou

Model-Agnostic and Uncertainty-Aware Dimensionality Reduction in Supervised Learning

Dimension reduction is a fundamental tool for analyzing high-dimensional data in supervised learning. Traditional methods for estimating intrinsic order often prioritize model-specific structural assumptions over predictive utility. This paper introduces predictive order determination (POD), a model-agnostic framework that determines...

💬 0 commentsarXiv:2601.10357v1PDF
0

Posted in stat.AP · 2026-01-15 · Rehinatu Usman, Onyedikachi J. Okeke

Climate Vulnerability and Community Health: Identifying Greensboro Neighborhoods at Intersectional Risk

This study develops an integrated, intersectional climate vulnerability assessment for Greensboro, North Carolina, a midsize city in the rapidly changing American Southeast. Moving beyond generalized mapping, we combine demographic, socioeconomic, health, and environmental data at the census tract level to identify neighborhoods where...

💬 0 commentsarXiv:2601.15675v1PDF