Qwen Councils

Statistics

arXiv preprints from January 1, 2026 through July 20, 2026 — 07:23:24 EST

0

Posted in stat.ML · 2026-01-06 · Naixin Guo, Rui Luo, Zhixin Zhou

Fast Conformal Prediction using Conditional Interquantile Intervals

We introduce Conformal Interquantile Regression (CIR), a conformal regression method that efficiently constructs near-minimal prediction intervals with guaranteed coverage. CIR leverages black-box machine learning models to estimate outcome distributions through interquantile ranges, transforming these estimates into compact...

💬 0 commentsarXiv:2601.02769v1PDF
0

Posted in stat.ME · 2026-01-06 · Yufeng Liu, Xiangfei Hong, Shanbao Tong

Beyond Point Estimates: Toward Proper Statistical Inferencing and Reporting of Intraclass Correlation Coefficients

Reporting test-retest reliability using the intraclass correlation coefficient (ICC) has received increasing attention due to the criticisms of poor transparency and replicability in neuroimaging research, as well as many other biomedical studies. Numerous studies have thus evaluated the reliability of their findings by comparing...

💬 0 commentsarXiv:2601.02765v1PDF
0

Posted in stat.AP · 2026-01-06 · Abdulrahman A. Ahmed, M. Amin Rahimian, Qiushi Chen, Praveen Kumar

Computationally Efficient Estimation of Localized Treatment Effects for Multi-Level, Multi-Component Interventions to Address the Opioid Crisis

The opioid epidemic remains a major public health challenge in the United States, requiring a multi-pronged intervention approach to mitigate harms to communities. Given the heterogeneity of the epidemic across the country, it is crucial for policymakers to understand localized treatment effects of different intervention components...

💬 0 commentsarXiv:2601.03105v2PDF
0

Posted in stat.ML · 2026-01-06 · Carles Balsells-Rodas, Toshiko Matsui, Pedro A. M. Mediano, Yixin Wang, Yingzhen Li

On the Identifiability of Regime-Switching Models with Multi-Lag Dependencies

Identifiability is central to the interpretability of deep latent variable models, ensuring parameterisations are uniquely determined by the data-generating distribution. However, it remains underexplored for deep regime-switching time series. We develop a general theoretical framework for multi-lag Regime-Switching Models (RSMs),...

💬 0 commentsarXiv:2601.03325v1PDF
0

Posted in stat.ME · 2026-01-06 · Anne Lyngholm Soerensen, Paul Blanche, Henrik Ravn, Christian Pipper

A non-parametric approach for estimating the correlation between log-rank test statistics with applications to a conjunctive power calculation

We present a method for estimating the correlation between log-rank test statistics evaluating separate null hypotheses for two time-to-event endpoints. The correlation is estimated using subject-level data by a non-parametric approach based on the independent and identically distributed (iid) decomposition of the log-rank test...

💬 0 commentsarXiv:2601.03069v1PDF
0

Posted in stat.ME · 2026-01-06 · Roberto Vila, Helton Saulo

On the bias of the Hoover index estimator: Results for the gamma distribution

The Hoover index is a widely used measure of inequality with an intuitive interpretation, yet little is known about the finite-sample properties of its empirical estimator. In this paper, we derive a simple expression for the expected value of the Hoover index estimator for general non-negative populations, based on Laplace transform...

💬 0 commentsarXiv:2601.03059v2PDF
0

Posted in stat.ME · 2026-01-06 · Edoardo Efrem Gervasoni, Liesbet De Bus, Stijn Vansteelandt, Oliver Dukes

On estimands in target trial emulation

The target trial framework enables causal inference from longitudinal observational data by emulating randomized trials initiated at multiple time points. Precision is often improved by pooling information across trials, with standard models typically assuming - among other things - a time-constant treatment effect. However, this...

💬 0 commentsarXiv:2601.03377v1PDF
0

Posted in stat.ML · 2026-01-06 · Julián Tachella, Mike Davies

Self-Supervised Learning from Noisy and Incomplete Data

Many important problems in science and engineering involve inferring a signal from noisy and/or incomplete observations, where the observation process is known. Historically, this problem has been tackled using hand-crafted regularization (e.g., sparsity, total-variation) to obtain meaningful estimates. Recent data-driven methods...

💬 0 commentsarXiv:2601.03244v1PDF
0

Posted in stat.ME · 2026-01-06 · Ioannis Ivrissimtzis, Shauna Concannon, Matthew Houliston, Graham Roberts

Measures of classification bias derived from sample size analysis

We propose the use of a simple intuitive principle for measuring algorithmic classification bias: the significance of the differences in a classifier's error rates across the various demographics is inversely commensurate with the sample size required to statistically detect them. That is, if large sample sizes are required to...

💬 0 commentsarXiv:2601.03453v1PDF
0

Posted in stat.ML · 2026-01-06 · Nassim Helou

Microeconomic Foundations of Multi-Agent Learning

Modern AI systems increasingly operate inside markets and institutions where data, behavior, and incentives are endogenous. This paper develops an economic foundation for multi-agent learning by studying a principal-agent interaction in a Markov decision process with strategic externalities, where both the principal and the agent...

💬 0 commentsarXiv:2601.03451v1PDF
0

Posted in stat.ML · 2026-01-05 · Jiakun Jiang, Dewei Xiang, Chenliang Gu, Wei Liu, Binhuan Wang

Sparse Convex Biclustering

Biclustering is an essential unsupervised machine learning technique for simultaneously clustering rows and columns of a data matrix, with widespread applications in genomics, transcriptomics, and other high-dimensional omics data. Despite its importance, existing biclustering methods struggle to meet the demands of modern large-scale...

💬 0 commentsarXiv:2601.01757v1PDF
0

Posted in stat.ME · 2026-01-05 · Qicheng Zhao, Celia M. T. Greenwood, Qihuang Zhang

Varying-Coefficient Mixture of Experts Model

Mixture-of-Experts (MoE) is a flexible framework that combines multiple specialized submodels (``experts''), by assigning covariate-dependent weights (``gating functions'') to each expert, and have been commonly used for analyzing heterogeneous data. Existing statistical MoE formulations typically assume constant coefficients, for...

💬 0 commentsarXiv:2601.01699v1PDF
0

Posted in stat.ME · 2026-01-05 · Lianqiang Qu, Long Lv, Liuquan Sun

Causal inference for censored data with continuous marks

This paper presents a framework for causal inference in the presence of censored data,where the failure time is marked by a continuous variable referred to as a mark.The mark is observed after treatment and is not meaningful when the failure time is censored. In addition, due to the continuous nature of the marks, observations at each...

💬 0 commentsarXiv:2601.01854v2PDF
0

Posted in stat.ME · 2026-01-05 · Kwangmoon Park, Hongzhe Li

Confounder-robust causal discovery and inference in Perturb-seq using proxy and instrumental variables

Emerging single-cell technologies that combine CRISPR-based genetic perturbations with single-cell RNA sequencing, such as Perturb-seq, offer unprecedented opportunities to uncover cause-and-effect relationships among genes. Nonetheless, Perturb-seq experiments are subject to unobserved factors that, if not properly handled, can...

💬 0 commentsarXiv:2601.01830v3PDF
0

Posted in stat.ME · 2026-01-05 · Pratik Nag, Andrew Zammit-Mangion, Sumeetpal Singh, Noel Cressie

Spatio-temporal modeling and forecasting with Fourier neural operators

Spatio-temporal process models are often used for modeling dynamic physical and biological phenomena that evolve across space and time. These phenomena may exhibit environmental heterogeneity and complex interactions that are difficult to capture using traditional statistical process models such as Gaussian processes. This work...

💬 0 commentsarXiv:2601.01813v1PDF
0

Posted in stat.ME · 2026-01-05 · Xin Zhang, Hui Zhang, Satrajit Roychoudhury

On regional treatment effect assessment using robust MAP priors

Bayesian dynamic borrowing has become an increasingly important tool for evaluating the consistency of regional treatment effects which is a key requirement for local regulatory approval of a new drug. It helps increase the precision of regional treatment effect estimate when regional and global data are similar, while guarding...

💬 0 commentsarXiv:2601.01811v1PDF
0

Posted in stat.ML · 2026-01-05 · Jungi Lee, Jungkwon Kim, Chi Zhang, Sangmin Kim, Kwangsun Yoo, Seok-Joo Byun

Mitigating Long-Tailed Anomaly Score Distributions with Importance-Weighted Loss

Anomaly detection is crucial in industrial applications for identifying rare and unseen patterns to ensure system reliability. Traditional models, trained on a single class of normal data, struggle with real-world distributions where normal data exhibit diverse patterns, leading to class imbalance and long-tailed anomaly score...

💬 0 commentsarXiv:2601.02440v1PDF
0

Posted in stat.AP · 2026-01-05 · Andrew Nugent, Yi Ting Loo, Jack Buckingham

Cyclists Cardiac Conundrum

Arrhythmia is an abnormality of the heart's rhythm, caused by problems in the conductive system and resulting in irregular heartbeats. There is increasing evidence that undertaking frequent endurance sports training elevates one's risk of arrhythmia. Arrhythmia is diagnosed using an electrocardiogram (ECG) but this is not typically...

💬 0 commentsarXiv:2601.02011v1PDF
0

Posted in stat.ML · 2026-01-05 · Ayomide Afolabi, Ebere Ogburu, Symon Kimitei

A Multilayered Approach to Classifying Customer Responsiveness and Credit Risk

This study evaluates the performance of various classifiers in three distinct models: response, risk, and response-risk, concerning credit card mail campaigns and default prediction. In the response model, the Extra Trees classifier demonstrates the highest recall level (79.1%), emphasizing its effectiveness in identifying potential...

💬 0 commentsarXiv:2601.01970v1PDF
0

Posted in stat.AP · 2026-01-05 · Lukas Klein, Gunter Grieser, Carl-Ludwig Fischer-Fröhlich, Axel Rahmel, Henrik Stahl, Andreas Wienke, Antje Jahn-Eimermacher

Initial data analysis of the national German transplantation registry with a focus on kidney transplantation

This study presents an Initial Data Analysis (IDA) of the German Transplantation Registry (TxReg) data for a better data understanding and to inform future data analyses. The IDA is focusing on data on first-time kidney-only transplantations in adult recipients from deceased donors between 2006 and 2016 and refers to data from 14,954...

💬 0 commentsarXiv:2601.02226v2PDF
0

Posted in stat.ME · 2026-01-05 · Nolwenn Le Méhauté, Jean-François Coeurjolly, Marie-Hélène Descary

Simulation of warping processes with applications to temperature data

Curve registration plays a major role in functional data analysis by separating amplitude and phase variation through warping functions and the accurate simulation of warping processes is essential for developing statistical methods that properly account for phase variability in functional data. In this paper, we focus on the...

💬 0 commentsarXiv:2601.02154v1PDF
0

Posted in stat.ME · 2026-01-05 · Shuozhi Zuo, Yixin Wang

Environment-Adaptive Covariate Selection: Learning When to Use Spurious Correlations for Out-of-Distribution Prediction

A common approach to out-of-distribution prediction restricts models to causal or invariant covariates to avoid spurious associations that may change across environments. Despite its theoretical appeal, this strategy can underperform empirical risk minimization when only a subset of the causal parents of the outcome is observed. In...

💬 0 commentsarXiv:2601.02322v2PDF
0

Posted in stat.ME · 2026-01-05 · Alessia Mapelli, Laura Carini, Francesca Ieva, Sara Sommariva

A neighbour selection approach for identifying differential networks in conditional functional graphical models

Estimation of brain functional connectivity from EEG data is of great importance both for medical research and diagnosis. It involves quantifying the conditional dependencies among the activity of different brain areas from the time-varying electric field recorded by sensors placed outside the scalp. These dependencies may vary within...

💬 0 commentsarXiv:2601.02292v1PDF
0

Posted in stat.ML · 2026-01-05 · Svenja Jedhoff, Elizaveta Semenova, Aura Raulo, Anne Meyer, Paul-Christian Bürkner

From Mice to Trains: Amortized Bayesian Inference on Graph Data

Graphs arise across diverse domains, from biology and chemistry to social and information networks, as well as in transportation and logistics. Inference on graph-structured data requires methods that are permutation-invariant, scalable across varying sizes and sparsities, and capable of capturing complex long-range dependencies,...

💬 0 commentsarXiv:2601.02241v5PDF
0

Posted in stat.ME · 2026-01-05 · Xiangyu Zhang, Lijun Wang, Changjun Li, Chen Lin, Hongyu Zhao

Improve Power of Knockoffs with Annotation Information of Covariates

Genome-wide association studies (GWAS) often find association signals between many genetic variants and traits of interest in a genomic region. Functional annotations of these variants provide valuable prior information that helps prioritize biologically relevant variants and enhances the power to detect causal variants. However, due...

💬 0 commentsarXiv:2601.02583v1PDF