Qwen Councils

Statistics

arXiv preprints from January 1, 2026 through July 20, 2026 — 02:30:54 EST

0

Posted in stat.ME · 2026-01-11 · Roman Hornung, Alexander Hapfelmeier

Unity Forests: Improving Interaction Modelling and Interpretability in Random Forests

Random forests (RFs) are widely used for prediction and variable importance analysis and are often believed to capture any types of interactions via recursive splitting. However, since the splits are chosen locally, interactions are only reliably captured when at least one involved covariate has a marginal effect. We introduce unity...

💬 0 commentsarXiv:2601.07003v1PDF
0

Posted in stat.ME · 2026-01-11 · Hao-Xuan Sun, Song Xi Chen, Yumou Qiu

Localization Estimator for High Dimensional Tensor Covariance Matrices

This paper considers covariance matrix estimation of tensor data under high dimensionality. A multi-bandable covariance class is established to accommodate the need for complex covariance structures of multi-layer lattices and general covariance decay patterns. We propose a high dimensional covariance localization estimator for tensor...

💬 0 commentsarXiv:2601.06989v1PDF
0

Posted in stat.ML · 2026-01-11 · Zhiyuan Tang, Wanning Chen, Kan Xu

Match Made with Matrix Completion: Efficient Learning under Matching Interference

Matching markets face increasing needs to learn the matching qualities between demand and supply for effective design of matching policies. In practice, the matching rewards are high-dimensional due to the growing diversity of participants. We leverage a natural low-rank matrix structure of the matching rewards in these two-sided...

💬 0 commentsarXiv:2601.06982v1PDF
0

Posted in stat.ML · 2026-01-11 · Taishi Watanabe, Ryo Karakida, Jun-nosuke Teramae

The Impact of Anisotropic Covariance Structure on the Training Dynamics and Generalization Error of Linear Networks

The success of deep neural networks largely depends on the statistical structure of the training data. While learning dynamics and generalization on isotropic data are well-established, the impact of pronounced anisotropy on these crucial aspects is not yet fully understood. We examine the impact of data anisotropy, represented by a...

💬 0 commentsarXiv:2601.06961v1PDF
0

Posted in stat.ME · 2026-01-11 · Jiguang Li, Hengrui Luo

Robust Bayesian Optimization via Tempered Posteriors

Bayesian optimization (BO) iteratively fits a Gaussian process (GP) surrogate to accumulated evaluations and selects new queries via an acquisition function such as expected improvement (EI). In practice, BO often concentrates evaluations near the current incumbent, causing the surrogate to become overconfident and to understate...

💬 0 commentsarXiv:2601.07094v1PDF
0

Posted in stat.ML · 2026-01-11 · Pedro Abdalla, Junren Chen

Robust Mean Estimation under Quantization

We consider the problem of mean estimation under quantization and adversarial corruption. We construct multivariate robust estimators that are optimal up to logarithmic factors in two different settings. The first is a one-bit setting, where each bit depends only on a single sample, and the second is a partial quantization setting, in...

💬 0 commentsarXiv:2601.07074v1PDF
0

Posted in stat.CO · 2026-01-11 · Eric Feltham

FormulaCompiler.jl and Margins.jl: Efficient Marginal Effects in Julia

Marginal effects analysis is fundamental to interpreting statistical models, yet existing implementations face computational constraints that limit analysis at scale. We introduce two Julia packages that address this gap. Margins.jl provides a clean two-function API organizing analysis around a 2-by-2 framework: evaluation context...

💬 0 commentsarXiv:2601.07065v1PDF
0

Posted in stat.ML · 2026-01-11 · Alex Kokot, Anand Hemmady, Vydhourie Thiyageswaran, Marina Meila

Local EGOP for Continuous Index Learning

We introduce the setting of continuous index learning, in which a function of many variables varies only along a small number of directions at each point. For efficient estimation, it is beneficial for a learning algorithm to adapt, near each point $x$, to the subspace that captures the local variability of the function $f$. We pose...

💬 0 commentsarXiv:2601.07061v4PDF
0

Posted in stat.ME · 2026-01-11 · Yuhao Deng, Donglin Zeng, Yuanjia Wang

Semiparametric Analysis of Interval-Censored Data Subject to Inaccurate Diagnoses with A Terminal Event

Interval-censoring frequently occurs in studies of chronic diseases where disease status is inferred from intermittently collected biomarkers. Although many methods have been developed to analyze such data, they typically assume perfect disease diagnosis, which often does not hold in practice due to the inherent imperfect clinical...

💬 0 commentsarXiv:2601.07044v1PDF
0

Posted in stat.ME · 2026-01-10 · Qianqian Yao

Empirical Likelihood Test for Common Invariant Subspace of Multilayer Networks based on Monte Carlo Approximation

Multilayer (or multiple) networks are widely used to represent diverse patterns of relationships among objects in increasingly complex real-world systems. Identifying a common invariant subspace across network layers has become an active area of research, as such a subspace can filter out layer-specific noise, facilitate cross-network...

💬 0 commentsarXiv:2601.06390v1PDF
0

Posted in stat.CO · 2026-01-10 · Foo Hui-Mean, Yuan-chin Ivan Chang

Efficient Data Reduction Via PCA-Guided Quantile Based Sampling

In large-scale statistical modeling, reducing data size through subsampling is essential for balancing computational efficiency and statistical accuracy. We propose a new method, Principal Component Analysis guided Quantile Sampling (PCA-QS), which projects data onto principal components and applies quantile-based sampling to retain...

💬 0 commentsarXiv:2601.06375v1PDF
0

Posted in stat.ML · 2026-01-10 · Jinyuan Chang, Chenguang Duan, Yuling Jiao, Yi Xu, Jerry Zhijian Yang

Inference-Time Alignment for Diffusion Models via Variationally Stable Doob's Matching

Inference-time alignment for diffusion models aims to adapt a pre-trained reference diffusion model toward a target distribution without retraining the reference score network, thereby preserving the generative capacity of the reference model while enforcing desired properties at the inference time. A central mechanism for achieving...

💬 0 commentsarXiv:2601.06514v2PDF
0

Posted in stat.ME · 2026-01-10 · Qunqiang Feng, Yaru Tian, Ting Yan

Triple-dyad ratio estimation for the $p_1$ model

Although the $p_1$ model was proposed 40 years ago, little progress has been made to address asymptotic theories in this model, that is, neither consistency of the maximum likelihood estimator (MLE) nor other parameter estimation with statistical guarantees is understood. This problem has been acknowledged as a long-standing open...

💬 0 commentsarXiv:2601.06481v1PDF
0

Posted in stat.ML · 2026-01-10 · Tianming Bai, Jiannan Yang

Physics-informed Gaussian Process Regression in Solving Eigenvalue Problem of Linear Operators

Applying Physics-Informed Gaussian Process Regression to the eigenvalue problem $(\mathcal{L}-λ)u = 0$ poses a fundamental challenge, where the null source term results in a trivial predictive mean and a degenerate marginal likelihood. Drawing inspiration from system identification, we construct a transfer function-type indicator for...

💬 0 commentsarXiv:2601.06462v1PDF
0

Posted in stat.ME · 2026-01-10 · Monika S. Dhull

Mittag Leffler Distributions Estimation and Autoregressive Framework

This work deals with the estimation of parameters of Mittag-Leffler (ML($α, σ$)) distribution. We estimate the parameters of ML($α, σ$) using empirical Laplace transform method. The simulation study indicates that the proposed method provides satisfactory results. The real life application of ML($α, σ$) distribution on high frequency...

💬 0 commentsarXiv:2601.06610v1PDF
0

Posted in stat.ME · 2026-01-10 · Mengta Chung

A Symmetric Random Scan Collapsed Gibbs Sampler for Fully Bayesian Variable Selection with Spike-and-Slab Priors

We introduce a symmetric random scan Gibbs sampler for scalable Bayesian variable selection that eliminates storage of the full cross-product matrix by computing required quantities on-the-fly. Data-informed proposal weights, constructed from marginal correlations, concentrate sampling effort on promising candidates while a uniform...

💬 0 commentsarXiv:2601.07864v1PDF
0

Posted in stat.ME · 2026-01-10 · Genshiro Kitagawa

Bayesian Optimization of Noisy Log-Likelihoods Evaluated by Particle Filters -- One Parameter Case --

Likelihood functions evaluated using particle filters are typically noisy, computationally expensive, and non-differentiable due to Monte Carlo variability. These characteristics make conventional optimization methods difficult to apply directly or potentially unreliable. This paper investigates the use of Bayesian optimization for...

💬 0 commentsarXiv:2601.06545v1PDF
0

Posted in stat.ME · 2026-01-10 · Sphiwe B. Skhosana, Weixin Yao

Nonparametric contaminated Gaussian mixture of regressions

Semi- and non-parametric mixture of regressions are a very useful flexible class of mixture of regressions in which some or all of the parameters are non-parametric functions of the covariates. These models are, however, based on the Gaussian assumption of the component error distributions. Thus, their estimation is sensitive to...

💬 0 commentsarXiv:2601.06695v1PDF
0

Posted in stat.ME · 2026-01-10 · Glen A. Satten, Mo Li, Ni Zhao, Robert L. Strawderman

R-Estimation with Right-Censored Data

This paper considers the problem of directly generalizing the R-estimator under a linear model formulation with right-censored outcomes. We propose a natural generalization of the rank and corresponding estimating equation for the R-estimator in the case of the Wilcoxon (i.e., linear-in-ranks) score function, and show how it can...

💬 0 commentsarXiv:2601.06685v1PDF
0

Posted in stat.ME · 2026-01-10 · The Tien Mai, Sayantan Banerjee

Censored Graphical Horseshoe: Bayesian sparse precision matrix estimation with censored and missing data

Gaussian graphical models provide a powerful framework for studying conditional dependencies in multivariate data, with widespread applications spanning biomedical, environmental sciences, and other data-rich scientific domains. While the Graphical Horseshoe (GHS) method has emerged as a state-of-the-art Bayesian method for sparse...

💬 0 commentsarXiv:2601.06671v1PDF
0

Posted in stat.ML · 2026-01-09 · Getachew K. Befekadu

A brief note on learning problem with global perspectives

This brief note considers the problem of learning with dynamic-optimizing principal-agent setting, in which the agents are allowed to have global perspectives about the learning process, i.e., the ability to view things according to their relative importances or in their true relations based-on some aggregated information shared by...

💬 0 commentsarXiv:2601.05441v1PDF
0

Posted in stat.ME · 2026-01-09 · Kaiyuan Zhou, Xiaoyu Zhang, Wenyang Zhang, Di Wang

Two-Stage Robust Sparse Gradient Methods for Regression Under Heavy-Tailed Designs

We study high-dimensional sparse regression under simultaneous heavy-tailed covariates and noise. Heavy-tailed data affect sparse optimization in two different ways: extreme covariates can destabilize the gradient field during global localization, while heavy-tailed noise limits the final statistical accuracy during local refinement....

💬 0 commentsarXiv:2601.05669v2PDF
0

Posted in stat.ME · 2026-01-09 · Jiayi Wang

Conditional Cauchy-Schwarz Divergence for Time Series Analysis: Kernelized Estimation and Applications in Clustering and Fraud Detection

We study the conditional Cauchy-Schwarz divergence (C-CSD) as a symmetric and density-free measure for time series analysis. We derive a practical kernel based estimator using radial basis function kernels on both the condition and output spaces, together with numerical stabilizations including a symmetric logarithmic form with an...

💬 0 commentsarXiv:2601.05711v1PDF
0

Posted in stat.AP · 2026-01-09 · Joseph Marsh, Nathan A. Judd, Lax Chan, Rowland G. Seymour

Neural Methods for Multiple Systems Estimation Models

Estimating the size of hidden populations using Multiple Systems Estimation (MSE) is a critical task in quantitative sociology; however, practical application is often hindered by imperfect administrative data and computational constraints. Real-world datasets frequently suffer from censoring and missingness due to privacy concerns,...

💬 0 commentsarXiv:2601.05859v1PDF
0

Posted in stat.AP · 2026-01-09 · Jonathan F. Kunst, Killian A. C. Melsen, Willem Kruijer, José Crossa, Chris Maliepaard, Fred A. van Eeuwijk, Carel F. W. Peeters

A latent factor approach to hyperspectral time series data for multivariate genomic prediction of grain yield in wheat

High-dimensional time series phenotypic data is becoming increasingly common within plant breeding programmes. However, analysing and integrating such data for genetic analysis and genomic prediction remains difficult. Here we show how factor analysis with Procrustes rotation on the genetic correlation matrix of hyperspectral...

💬 0 commentsarXiv:2601.05842v1PDF