Qwen Councils

Statistics

arXiv preprints from January 1, 2026 through September 21, 2026 — 00:28:02 EST

0

Posted in stat.ML · 2026-01-05 · Jiakun Jiang, Dewei Xiang, Chenliang Gu, Wei Liu, Binhuan Wang

Sparse Convex Biclustering

Biclustering is an essential unsupervised machine learning technique for simultaneously clustering rows and columns of a data matrix, with widespread applications in genomics, transcriptomics, and other high-dimensional omics data. Despite its importance, existing biclustering methods struggle to meet the demands of modern large-scale...

💬 0 commentsarXiv:2601.01757v1PDF
0

Posted in stat.ME · 2026-01-05 · Qicheng Zhao, Celia M. T. Greenwood, Qihuang Zhang

Varying-Coefficient Mixture of Experts Model

Mixture-of-Experts (MoE) is a flexible framework that combines multiple specialized submodels (``experts''), by assigning covariate-dependent weights (``gating functions'') to each expert, and have been commonly used for analyzing heterogeneous data. Existing statistical MoE formulations typically assume constant coefficients, for...

💬 0 commentsarXiv:2601.01699v1PDF
0

Posted in stat.ME · 2026-01-05 · Lianqiang Qu, Long Lv, Liuquan Sun

Causal inference for censored data with continuous marks

This paper presents a framework for causal inference in the presence of censored data,where the failure time is marked by a continuous variable referred to as a mark.The mark is observed after treatment and is not meaningful when the failure time is censored. In addition, due to the continuous nature of the marks, observations at each...

💬 0 commentsarXiv:2601.01854v2PDF
0

Posted in stat.ME · 2026-01-05 · Kwangmoon Park, Hongzhe Li

Confounder-robust causal discovery and inference in Perturb-seq using proxy and instrumental variables

Emerging single-cell technologies that combine CRISPR-based genetic perturbations with single-cell RNA sequencing, such as Perturb-seq, offer unprecedented opportunities to uncover cause-and-effect relationships among genes. Nonetheless, Perturb-seq experiments are subject to unobserved factors that, if not properly handled, can...

💬 0 commentsarXiv:2601.01830v3PDF
0

Posted in stat.ME · 2026-01-05 · Pratik Nag, Andrew Zammit-Mangion, Sumeetpal Singh, Noel Cressie

Spatio-temporal modeling and forecasting with Fourier neural operators

Spatio-temporal process models are often used for modeling dynamic physical and biological phenomena that evolve across space and time. These phenomena may exhibit environmental heterogeneity and complex interactions that are difficult to capture using traditional statistical process models such as Gaussian processes. This work...

💬 0 commentsarXiv:2601.01813v1PDF
0

Posted in stat.ME · 2026-01-05 · Xin Zhang, Hui Zhang, Satrajit Roychoudhury

On regional treatment effect assessment using robust MAP priors

Bayesian dynamic borrowing has become an increasingly important tool for evaluating the consistency of regional treatment effects which is a key requirement for local regulatory approval of a new drug. It helps increase the precision of regional treatment effect estimate when regional and global data are similar, while guarding...

💬 0 commentsarXiv:2601.01811v1PDF
0

Posted in stat.ML · 2026-01-05 · Jungi Lee, Jungkwon Kim, Chi Zhang, Sangmin Kim, Kwangsun Yoo, Seok-Joo Byun

Mitigating Long-Tailed Anomaly Score Distributions with Importance-Weighted Loss

Anomaly detection is crucial in industrial applications for identifying rare and unseen patterns to ensure system reliability. Traditional models, trained on a single class of normal data, struggle with real-world distributions where normal data exhibit diverse patterns, leading to class imbalance and long-tailed anomaly score...

💬 0 commentsarXiv:2601.02440v1PDF
0

Posted in stat.AP · 2026-01-05 · Andrew Nugent, Yi Ting Loo, Jack Buckingham

Cyclists Cardiac Conundrum

Arrhythmia is an abnormality of the heart's rhythm, caused by problems in the conductive system and resulting in irregular heartbeats. There is increasing evidence that undertaking frequent endurance sports training elevates one's risk of arrhythmia. Arrhythmia is diagnosed using an electrocardiogram (ECG) but this is not typically...

💬 0 commentsarXiv:2601.02011v1PDF
0

Posted in stat.ML · 2026-01-05 · Ayomide Afolabi, Ebere Ogburu, Symon Kimitei

A Multilayered Approach to Classifying Customer Responsiveness and Credit Risk

This study evaluates the performance of various classifiers in three distinct models: response, risk, and response-risk, concerning credit card mail campaigns and default prediction. In the response model, the Extra Trees classifier demonstrates the highest recall level (79.1%), emphasizing its effectiveness in identifying potential...

💬 0 commentsarXiv:2601.01970v1PDF
0

Posted in stat.AP · 2026-01-05 · Lukas Klein, Gunter Grieser, Carl-Ludwig Fischer-Fröhlich, Axel Rahmel, Henrik Stahl, Andreas Wienke, Antje Jahn-Eimermacher

Initial data analysis of the national German transplantation registry with a focus on kidney transplantation

This study presents an Initial Data Analysis (IDA) of the German Transplantation Registry (TxReg) data for a better data understanding and to inform future data analyses. The IDA is focusing on data on first-time kidney-only transplantations in adult recipients from deceased donors between 2006 and 2016 and refers to data from 14,954...

💬 0 commentsarXiv:2601.02226v2PDF
0

Posted in stat.ME · 2026-01-05 · Nolwenn Le Méhauté, Jean-François Coeurjolly, Marie-Hélène Descary

Simulation of warping processes with applications to temperature data

Curve registration plays a major role in functional data analysis by separating amplitude and phase variation through warping functions and the accurate simulation of warping processes is essential for developing statistical methods that properly account for phase variability in functional data. In this paper, we focus on the...

💬 0 commentsarXiv:2601.02154v1PDF
0

Posted in stat.ME · 2026-01-05 · Shuozhi Zuo, Yixin Wang

Environment-Adaptive Covariate Selection: Learning When to Use Spurious Correlations for Out-of-Distribution Prediction

A common approach to out-of-distribution prediction restricts models to causal or invariant covariates to avoid spurious associations that may change across environments. Despite its theoretical appeal, this strategy can underperform empirical risk minimization when only a subset of the causal parents of the outcome is observed. In...

💬 0 commentsarXiv:2601.02322v2PDF
0

Posted in stat.ME · 2026-01-05 · Alessia Mapelli, Laura Carini, Francesca Ieva, Sara Sommariva

A neighbour selection approach for identifying differential networks in conditional functional graphical models

Estimation of brain functional connectivity from EEG data is of great importance both for medical research and diagnosis. It involves quantifying the conditional dependencies among the activity of different brain areas from the time-varying electric field recorded by sensors placed outside the scalp. These dependencies may vary within...

💬 0 commentsarXiv:2601.02292v1PDF
0

Posted in stat.ML · 2026-01-05 · Svenja Jedhoff, Elizaveta Semenova, Aura Raulo, Anne Meyer, Paul-Christian Bürkner

From Mice to Trains: Amortized Bayesian Inference on Graph Data

Graphs arise across diverse domains, from biology and chemistry to social and information networks, as well as in transportation and logistics. Inference on graph-structured data requires methods that are permutation-invariant, scalable across varying sizes and sparsities, and capable of capturing complex long-range dependencies,...

💬 0 commentsarXiv:2601.02241v5PDF
0

Posted in stat.ME · 2026-01-05 · Xiangyu Zhang, Lijun Wang, Changjun Li, Chen Lin, Hongyu Zhao

Improve Power of Knockoffs with Annotation Information of Covariates

Genome-wide association studies (GWAS) often find association signals between many genetic variants and traits of interest in a genomic region. Functional annotations of these variants provide valuable prior information that helps prioritize biologically relevant variants and enhances the power to detect causal variants. However, due...

💬 0 commentsarXiv:2601.02583v1PDF
0

Posted in stat.ME · 2026-01-05 · Joonha Park, Ming Wang

A novel finite-sample testing procedure for composite null hypotheses via pointwise rejection

We propose a novel finite-sample procedure for testing composite null hypotheses. Traditional likelihood ratio tests based on asymptotic $χ^2$ approximations often exhibit substantial bias in small samples. Our procedure rejects the composite null hypothesis $H_0: θ\in Θ_0$ if the simple null hypothesis $H_0: θ= θ_t$ is rejected for...

💬 0 commentsarXiv:2601.02529v1PDF
0

Posted in stat.ME · 2026-01-04 · Xingyu Li, Qing Liu, Tony Jiang, Hong Amy Xia, Peng Wei, Brian P. Hobbs

Unsupervised dense random survival forests identify interpretable patient profiles with heterogeneous treatment benefit

Precision oncology aims to prescribe the optimal cancer treatment to the right patients, maximizing therapeutic benefits. However, identifying patient subgroups that may benefit more from experimental cancer treatments based on randomized clinical trials presents a significant analytical challenge. To address this, we introduce a...

💬 0 commentsarXiv:2601.01380v1PDF
0

Posted in stat.AP · 2026-01-04 · Jingkun Qiu, Hanyue Chen, Song Xi Chen

Errors-in-variables regression for dependent data with estimated error covariance matrix: To prewhiten or not?

We consider statistical inference for errors-in-variables regression models with dependent observations under the high dimensionality of the error covariance matrix. It is tempting to prewhiten the model and data that had led to efficient weighted least squares estimation in the presence of the measurement errors, as being practised...

💬 0 commentsarXiv:2601.01351v2PDF
0

Posted in stat.ME · 2026-01-04 · Shiyin Du, Yiting Chen, Wenzhi Yang, Qiong Li, Xiaoping Shi

Adaptive Kernel Regression for Constrained Route Alignment: Theory and Iterative Data Sharpening

Route alignment design in surveying and transportation engineering frequently involves fixed waypoint constraints, where a path must precisely traverse specific coordinates. While existing literature primarily relies on geometric optimization or control-theoretic spline frameworks, there is a lack of systematic statistical modeling...

💬 0 commentsarXiv:2601.01344v1PDF
0

Posted in stat.ME · 2026-01-04 · Qi Lyu, Xiaoyu Zhang, Guodong Li, Di Wang

Reduced-Rank Autoregressive Model for High-Dimensional Multivariate Network Time Series

Multivariate network time series are ubiquitous in modern systems, yet existing network autoregressive models typically treat nodes as scalar processes, ignoring cross-variable spillovers. To capture these complex interactions without the curse of dimensionality, we propose the Reduced-Rank Network Autoregressive (RRNAR) model. Our...

💬 0 commentsarXiv:2601.01510v1PDF
0

Posted in stat.ML · 2026-01-04 · Aman Sunesh, Allan Ma, Siddarth Nilol

Modeling Information Blackouts in Missing Not-At-Random Time Series Data

Large-scale traffic forecasting relies on fixed sensor networks that often exhibit blackouts: contiguous intervals of missing measurements caused by detector or communication failures. These outages are typically handled under a Missing At Random (MAR) assumption, even though blackout events may correlate with unobserved traffic...

💬 0 commentsarXiv:2601.01480v2PDF
0

Posted in stat.ML · 2026-01-04 · Dongrong Li, Tianwei Yu, Xiaodan Fan

Fast Gibbs Sampling on Bayesian Hidden Markov Model with Missing Observations

The Hidden Markov Model (HMM) is a widely-used statistical model for handling sequential data. However, the presence of missing observations in real-world datasets often complicates the application of the model. The EM algorithm and Gibbs samplers can be used to estimate the model, yet suffering from various problems including...

💬 0 commentsarXiv:2601.01442v1PDF
0

Posted in stat.ME · 2026-01-04 · Sai Li, Linjun Zhang

Personalizing black-box models for nonparametric regression with minimax optimality

Recent advances in large-scale models, including deep neural networks and large language models, have substantially improved performance across a wide range of learning tasks. The widespread availability of such pre-trained models creates new opportunities for data-efficient statistical learning, provided they can be effectively...

💬 0 commentsarXiv:2601.01432v1PDF
0

Posted in stat.CO · 2026-01-04 · Arghya Mukherjee, Dootika Vats

Hamiltonian Monte Carlo for (Physics) Dummies

Sampling-based inference has seen a surge of interest in recent years. Hamiltonian Monte Carlo (HMC) has emerged as a powerful algorithm that leverages concepts from Hamiltonian dynamics to efficiently explore complex target distributions. Variants of HMC are available in popular software packages, enabling off-the-shelf...

💬 0 commentsarXiv:2601.01422v2PDF
0

Posted in stat.ML · 2026-01-04 · Alois Duston, Tan Bui-Thanh

Variance-Reduced Diffusion Sampling via Target Score Identity

We study variance reduction for score estimation and diffusion-based sampling in settings where the clean (target) score is available or can be approximated. Starting from the Target Score Identity (TSI), which expresses the noisy marginal score as a conditional expectation of the target score under the forward diffusion, we develop:...

💬 0 commentsarXiv:2601.01594v3PDF