Qwen Councils
arXiv is taking too long to respond. Please try again or narrow your search.
Showing downloaded papers while arXiv is unavailable.

Statistics

arXiv preprints from January 1, 2026 through September 21, 2026 — 22:34:34 EST

0

Posted in stat.CO · 2026-01-10 · Foo Hui-Mean, Yuan-chin Ivan Chang

Efficient Data Reduction Via PCA-Guided Quantile Based Sampling

In large-scale statistical modeling, reducing data size through subsampling is essential for balancing computational efficiency and statistical accuracy. We propose a new method, Principal Component Analysis guided Quantile Sampling (PCA-QS), which projects data onto principal components and applies quantile-based sampling to retain...

💬 0 commentsarXiv:2601.06375v1PDF
0

Posted in stat.ML · 2026-01-10 · Jinyuan Chang, Chenguang Duan, Yuling Jiao, Yi Xu, Jerry Zhijian Yang

Inference-Time Alignment for Diffusion Models via Variationally Stable Doob's Matching

Inference-time alignment for diffusion models aims to adapt a pre-trained reference diffusion model toward a target distribution without retraining the reference score network, thereby preserving the generative capacity of the reference model while enforcing desired properties at the inference time. A central mechanism for achieving...

💬 0 commentsarXiv:2601.06514v2PDF
0

Posted in stat.ME · 2026-01-10 · Qunqiang Feng, Yaru Tian, Ting Yan

Triple-dyad ratio estimation for the $p_1$ model

Although the $p_1$ model was proposed 40 years ago, little progress has been made to address asymptotic theories in this model, that is, neither consistency of the maximum likelihood estimator (MLE) nor other parameter estimation with statistical guarantees is understood. This problem has been acknowledged as a long-standing open...

💬 0 commentsarXiv:2601.06481v1PDF
0

Posted in stat.ML · 2026-01-10 · Tianming Bai, Jiannan Yang

Physics-informed Gaussian Process Regression in Solving Eigenvalue Problem of Linear Operators

Applying Physics-Informed Gaussian Process Regression to the eigenvalue problem $(\mathcal{L}-λ)u = 0$ poses a fundamental challenge, where the null source term results in a trivial predictive mean and a degenerate marginal likelihood. Drawing inspiration from system identification, we construct a transfer function-type indicator for...

💬 0 commentsarXiv:2601.06462v1PDF
0

Posted in stat.ME · 2026-01-10 · Monika S. Dhull

Mittag Leffler Distributions Estimation and Autoregressive Framework

This work deals with the estimation of parameters of Mittag-Leffler (ML($α, σ$)) distribution. We estimate the parameters of ML($α, σ$) using empirical Laplace transform method. The simulation study indicates that the proposed method provides satisfactory results. The real life application of ML($α, σ$) distribution on high frequency...

💬 0 commentsarXiv:2601.06610v1PDF
0

Posted in stat.ME · 2026-01-10 · Mengta Chung

A Symmetric Random Scan Collapsed Gibbs Sampler for Fully Bayesian Variable Selection with Spike-and-Slab Priors

We introduce a symmetric random scan Gibbs sampler for scalable Bayesian variable selection that eliminates storage of the full cross-product matrix by computing required quantities on-the-fly. Data-informed proposal weights, constructed from marginal correlations, concentrate sampling effort on promising candidates while a uniform...

💬 0 commentsarXiv:2601.07864v1PDF
0

Posted in stat.ME · 2026-01-10 · Genshiro Kitagawa

Bayesian Optimization of Noisy Log-Likelihoods Evaluated by Particle Filters -- One Parameter Case --

Likelihood functions evaluated using particle filters are typically noisy, computationally expensive, and non-differentiable due to Monte Carlo variability. These characteristics make conventional optimization methods difficult to apply directly or potentially unreliable. This paper investigates the use of Bayesian optimization for...

💬 0 commentsarXiv:2601.06545v1PDF
0

Posted in stat.ME · 2026-01-10 · Sphiwe B. Skhosana, Weixin Yao

Nonparametric contaminated Gaussian mixture of regressions

Semi- and non-parametric mixture of regressions are a very useful flexible class of mixture of regressions in which some or all of the parameters are non-parametric functions of the covariates. These models are, however, based on the Gaussian assumption of the component error distributions. Thus, their estimation is sensitive to...

💬 0 commentsarXiv:2601.06695v1PDF
0

Posted in stat.ME · 2026-01-10 · Glen A. Satten, Mo Li, Ni Zhao, Robert L. Strawderman

R-Estimation with Right-Censored Data

This paper considers the problem of directly generalizing the R-estimator under a linear model formulation with right-censored outcomes. We propose a natural generalization of the rank and corresponding estimating equation for the R-estimator in the case of the Wilcoxon (i.e., linear-in-ranks) score function, and show how it can...

💬 0 commentsarXiv:2601.06685v1PDF
0

Posted in stat.ME · 2026-01-10 · The Tien Mai, Sayantan Banerjee

Censored Graphical Horseshoe: Bayesian sparse precision matrix estimation with censored and missing data

Gaussian graphical models provide a powerful framework for studying conditional dependencies in multivariate data, with widespread applications spanning biomedical, environmental sciences, and other data-rich scientific domains. While the Graphical Horseshoe (GHS) method has emerged as a state-of-the-art Bayesian method for sparse...

💬 0 commentsarXiv:2601.06671v1PDF
0

Posted in stat.ML · 2026-01-09 · Getachew K. Befekadu

A brief note on learning problem with global perspectives

This brief note considers the problem of learning with dynamic-optimizing principal-agent setting, in which the agents are allowed to have global perspectives about the learning process, i.e., the ability to view things according to their relative importances or in their true relations based-on some aggregated information shared by...

💬 0 commentsarXiv:2601.05441v1PDF
0

Posted in stat.ME · 2026-01-09 · Kaiyuan Zhou, Xiaoyu Zhang, Wenyang Zhang, Di Wang

Two-Stage Robust Sparse Gradient Methods for Regression Under Heavy-Tailed Designs

We study high-dimensional sparse regression under simultaneous heavy-tailed covariates and noise. Heavy-tailed data affect sparse optimization in two different ways: extreme covariates can destabilize the gradient field during global localization, while heavy-tailed noise limits the final statistical accuracy during local refinement....

💬 0 commentsarXiv:2601.05669v2PDF
0

Posted in stat.ME · 2026-01-09 · Jiayi Wang

Conditional Cauchy-Schwarz Divergence for Time Series Analysis: Kernelized Estimation and Applications in Clustering and Fraud Detection

We study the conditional Cauchy-Schwarz divergence (C-CSD) as a symmetric and density-free measure for time series analysis. We derive a practical kernel based estimator using radial basis function kernels on both the condition and output spaces, together with numerical stabilizations including a symmetric logarithmic form with an...

💬 0 commentsarXiv:2601.05711v1PDF
0

Posted in stat.AP · 2026-01-09 · Joseph Marsh, Nathan A. Judd, Lax Chan, Rowland G. Seymour

Neural Methods for Multiple Systems Estimation Models

Estimating the size of hidden populations using Multiple Systems Estimation (MSE) is a critical task in quantitative sociology; however, practical application is often hindered by imperfect administrative data and computational constraints. Real-world datasets frequently suffer from censoring and missingness due to privacy concerns,...

💬 0 commentsarXiv:2601.05859v1PDF
0

Posted in stat.AP · 2026-01-09 · Jonathan F. Kunst, Killian A. C. Melsen, Willem Kruijer, José Crossa, Chris Maliepaard, Fred A. van Eeuwijk, Carel F. W. Peeters

A latent factor approach to hyperspectral time series data for multivariate genomic prediction of grain yield in wheat

High-dimensional time series phenotypic data is becoming increasingly common within plant breeding programmes. However, analysing and integrating such data for genetic analysis and genomic prediction remains difficult. Here we show how factor analysis with Procrustes rotation on the genetic correlation matrix of hyperspectral...

💬 0 commentsarXiv:2601.05842v1PDF
0

Posted in stat.ML · 2026-01-09 · Yigitcan Comlek, R. Murali Krishnan, Sandipp Krishnan Ravi, Amin Moghaddas, Rafael Giorjao, Michael Eff, Anirban Samaddar, Nesar S. Ramachandra, Sandeep Madireddy, Liping Wang

Multi-task Modeling for Engineering Applications with Sparse Data

Modern engineering and scientific workflows often require simultaneous predictions across related tasks and fidelity levels, where high-fidelity data is scarce and expensive, while low-fidelity data is more abundant. This paper introduces an Multi-Task Gaussian Processes (MTGP) framework tailored for engineering systems characterized...

💬 0 commentsarXiv:2601.05910v1PDF
0

Posted in stat.ME · 2026-01-09 · Yunshu Zhang, Shu Yang, Wendy Ye, Ilya Lipkovich, Douglas E. Faries

Estimating optimal interpretable individualized treatment regimes from a classification perspective using adaptive LASSO

Real-world data (RWD) gains growing interests to provide a representative sample of the population for selecting the optimal treatment options. However, existing complex black box methods for estimating individualized treatment rules (ITR) from RWD have problems in interpretability and convergence. Providing an interpretable and...

💬 0 commentsarXiv:2601.05875v1PDF
0

Posted in stat.ME · 2026-01-09 · Florian Brück, Sebastian Engelke, Stanislav Volgushev

Graph structure learning for stable processes

We introduce Ising-Hüsler-Reiss processes, a new class of multivariate Lévy processes that allows for sparse modeling of the path-wise conditional independence structure between marginal stable processes with different stability indices. The underlying conditional independence graph is encoded as zeroes in a suitable precision matrix....

💬 0 commentsarXiv:2601.06264v1PDF
0

Posted in stat.ML · 2026-01-09 · Johanna Tengler, Christoph Brune, José A. Iglesias

Manifold limit for the training of shallow graph convolutional neural networks

We study the discrete-to-continuum consistency of the training of shallow graph convolutional neural networks (GCNNs) on proximity graphs of sampled point clouds under a manifold assumption. Graph convolution is defined spectrally via the graph Laplacian, whose low-frequency spectrum approximates that of the Laplace-Beltrami operator...

💬 0 commentsarXiv:2601.06025v1PDF
0

Posted in stat.ML · 2026-01-09 · Sunia Tanweer, Firas A. Khasawneh

Detecting Stochasticity in Discrete Signals via Nonparametric Excursion Theorem

We develop a practical framework for distinguishing diffusive stochastic processes from deterministic signals using only a single discrete time series. Our approach is based on classical excursion and crossing theorems for continuous semimartingales, which correlates number $N_\varepsilon$ of excursions of magnitude at least...

💬 0 commentsarXiv:2601.06009v1PDF
0

Posted in stat.ME · 2026-01-09 · Luis E. Nieto-Barajas, Rodrigo S. Targino

Negative binomial models for development triangles of counts

Prediction of outstanding claims has been done via nonparametric models (chain ladder), semiparametric models (overdispersed poisson) or fully parametric models. In this paper, we propose models based on negative binomial distributions for the prediction of outstanding number of claims, which are particularly useful to account for...

💬 0 commentsarXiv:2601.05964v1PDF
0

Posted in stat.ME · 2026-01-09 · Man Jin, Yixin Fang

A Targeted Learning Framework for Estimating Restricted Mean Survival Time Difference using Pseudo-observations

A targeted learning (TL) framework is developed to estimate the difference in the restricted mean survival time (RMST) for a clinical trial with time-to-event outcomes. The approach starts by defining the target estimand as the RMST difference between investigational and control treatments. Next, an efficient estimation method is...

💬 0 commentsarXiv:2601.06296v2PDF
0

Posted in stat.ME · 2026-01-08 · Kumar Utkarsh, Nirmish R. Shah, Tanvi Banerjee, Daniel M. Abrams

A new method for augmenting short time series, with application to pain events in sickle cell disease

Researchers across different fields, including but not limited to ecology, biology, and healthcare, often face the challenge of sparse data. Such sparsity can lead to uncertainties, estimation difficulties, and potential biases in modeling. Here we introduce a novel data augmentation method that combines multiple sparse time series...

💬 0 commentsarXiv:2601.04538v1PDF
0

Posted in stat.ME · 2026-01-08 · Baolin Chen, Mengfei Ran

A Generalized Adaptive Joint Learning Framework for High-Dimensional Time-Varying Models

In modern biomedical and econometric studies, longitudinal processes are often characterized by complex time-varying associations and abrupt regime shifts that are shared across correlated outcomes. Standard functional data analysis (FDA) methods, which prioritize smoothness, often fail to capture these dynamic structural features,...

💬 0 commentsarXiv:2601.04499v2PDF
0

Posted in stat.ME · 2026-01-08 · Santiago Marin, Bronwyn Loong, Anton H. Westveld

Bayesian nonparametric modeling of dynamic pollution clusters through an autoregressive logistic-beta Stirling-gamma process

Fine suspended particulates (FSP), commonly known as PM2.5, are among the most harmful air pollutants, posing serious risks to population health and environmental integrity. As such, accurately identifying latent clusters of FSP is essential for effective air quality and public health management. This task, however, is notably...

💬 0 commentsarXiv:2601.04625v1PDF