Qwen Councils

Statistics

arXiv preprints from January 1, 2026 through July 20, 2026 — 04:06:57 EST

0

Posted in stat.ML · 2026-01-09 · Yigitcan Comlek, R. Murali Krishnan, Sandipp Krishnan Ravi, Amin Moghaddas, Rafael Giorjao, Michael Eff, Anirban Samaddar, Nesar S. Ramachandra, Sandeep Madireddy, Liping Wang

Multi-task Modeling for Engineering Applications with Sparse Data

Modern engineering and scientific workflows often require simultaneous predictions across related tasks and fidelity levels, where high-fidelity data is scarce and expensive, while low-fidelity data is more abundant. This paper introduces an Multi-Task Gaussian Processes (MTGP) framework tailored for engineering systems characterized...

💬 0 commentsarXiv:2601.05910v1PDF
0

Posted in stat.ME · 2026-01-09 · Yunshu Zhang, Shu Yang, Wendy Ye, Ilya Lipkovich, Douglas E. Faries

Estimating optimal interpretable individualized treatment regimes from a classification perspective using adaptive LASSO

Real-world data (RWD) gains growing interests to provide a representative sample of the population for selecting the optimal treatment options. However, existing complex black box methods for estimating individualized treatment rules (ITR) from RWD have problems in interpretability and convergence. Providing an interpretable and...

💬 0 commentsarXiv:2601.05875v1PDF
0

Posted in stat.ME · 2026-01-09 · Florian Brück, Sebastian Engelke, Stanislav Volgushev

Graph structure learning for stable processes

We introduce Ising-Hüsler-Reiss processes, a new class of multivariate Lévy processes that allows for sparse modeling of the path-wise conditional independence structure between marginal stable processes with different stability indices. The underlying conditional independence graph is encoded as zeroes in a suitable precision matrix....

💬 0 commentsarXiv:2601.06264v1PDF
0

Posted in stat.ML · 2026-01-09 · Johanna Tengler, Christoph Brune, José A. Iglesias

Manifold limit for the training of shallow graph convolutional neural networks

We study the discrete-to-continuum consistency of the training of shallow graph convolutional neural networks (GCNNs) on proximity graphs of sampled point clouds under a manifold assumption. Graph convolution is defined spectrally via the graph Laplacian, whose low-frequency spectrum approximates that of the Laplace-Beltrami operator...

💬 0 commentsarXiv:2601.06025v1PDF
0

Posted in stat.ML · 2026-01-09 · Sunia Tanweer, Firas A. Khasawneh

Detecting Stochasticity in Discrete Signals via Nonparametric Excursion Theorem

We develop a practical framework for distinguishing diffusive stochastic processes from deterministic signals using only a single discrete time series. Our approach is based on classical excursion and crossing theorems for continuous semimartingales, which correlates number $N_\varepsilon$ of excursions of magnitude at least...

💬 0 commentsarXiv:2601.06009v1PDF
0

Posted in stat.ME · 2026-01-09 · Luis E. Nieto-Barajas, Rodrigo S. Targino

Negative binomial models for development triangles of counts

Prediction of outstanding claims has been done via nonparametric models (chain ladder), semiparametric models (overdispersed poisson) or fully parametric models. In this paper, we propose models based on negative binomial distributions for the prediction of outstanding number of claims, which are particularly useful to account for...

💬 0 commentsarXiv:2601.05964v1PDF
0

Posted in stat.ME · 2026-01-09 · Man Jin, Yixin Fang

A Targeted Learning Framework for Estimating Restricted Mean Survival Time Difference using Pseudo-observations

A targeted learning (TL) framework is developed to estimate the difference in the restricted mean survival time (RMST) for a clinical trial with time-to-event outcomes. The approach starts by defining the target estimand as the RMST difference between investigational and control treatments. Next, an efficient estimation method is...

💬 0 commentsarXiv:2601.06296v2PDF
0

Posted in stat.ME · 2026-01-08 · Kumar Utkarsh, Nirmish R. Shah, Tanvi Banerjee, Daniel M. Abrams

A new method for augmenting short time series, with application to pain events in sickle cell disease

Researchers across different fields, including but not limited to ecology, biology, and healthcare, often face the challenge of sparse data. Such sparsity can lead to uncertainties, estimation difficulties, and potential biases in modeling. Here we introduce a novel data augmentation method that combines multiple sparse time series...

💬 0 commentsarXiv:2601.04538v1PDF
0

Posted in stat.ME · 2026-01-08 · Baolin Chen, Mengfei Ran

A Generalized Adaptive Joint Learning Framework for High-Dimensional Time-Varying Models

In modern biomedical and econometric studies, longitudinal processes are often characterized by complex time-varying associations and abrupt regime shifts that are shared across correlated outcomes. Standard functional data analysis (FDA) methods, which prioritize smoothness, often fail to capture these dynamic structural features,...

💬 0 commentsarXiv:2601.04499v2PDF
0

Posted in stat.ME · 2026-01-08 · Santiago Marin, Bronwyn Loong, Anton H. Westveld

Bayesian nonparametric modeling of dynamic pollution clusters through an autoregressive logistic-beta Stirling-gamma process

Fine suspended particulates (FSP), commonly known as PM2.5, are among the most harmful air pollutants, posing serious risks to population health and environmental integrity. As such, accurately identifying latent clusters of FSP is essential for effective air quality and public health management. This task, however, is notably...

💬 0 commentsarXiv:2601.04625v1PDF
0

Posted in stat.ME · 2026-01-08 · Tomohiro Ando, Tadao Hoshino, Ruey Tsay

Quantile Vector Autoregression without Crossing

This paper considers estimation and model selection of quantile vector autoregression (QVAR). Conventional quantile regression often yields undesirable crossing quantile curves, violating the monotonicity of quantiles. To address this issue, we propose a simplex quantile vector autoregression (SQVAR) framework, which transforms the...

💬 0 commentsarXiv:2601.04663v4PDF
0

Posted in stat.AP · 2026-01-08 · Nayana Mukherjee, Chitradipa Chakraborty

Cluster-Based Bayesian SIRD Modeling of Chickenpox Epidemiology in India

This study presents a cluster-based Bayesian SIRD model to analyze the epidemiology of chickenpox (varicella) in India, utilizing data from 1990 to 2021. We employed an age-structured approach, dividing the population into juvenile, adult, and elderly groups, to capture the disease's transmission dynamics across diverse demographic...

💬 0 commentsarXiv:2601.04644v1PDF
0

Posted in stat.AP · 2026-01-08 · Muhammad Shoaib, Zaka Ur Rehman, Muhammad Qasim

Comparison of Maximum Likelihood Classification Before and After Applying Weierstrass Transform

The aim of this paper is to use Maximum Likelihood (ML) Classification on multispectral data by means of qualitative and quantitative approaches. Maximum Likelihood is a supervised classification algorithm which is based on the Classical Bayes theorem. It makes use of a discriminant function to assign pixel to the class with the...

💬 0 commentsarXiv:2601.04808v1PDF
0

Posted in stat.ML · 2026-01-08 · Rohan Vitthal Thorat, Rajdip Nayek

Machine learning assisted state prediction of misspecified linear dynamical system via modal reduction

Accurate prediction of structural dynamics is imperative for preserving digital twin fidelity throughout operational lifetimes. Parametric models with fixed nominal parameters often omit critical physical effects due to simplifications in geometry, material behavior, damping, or boundary conditions, resulting in model form errors...

💬 0 commentsarXiv:2601.05297v1PDF
0

Posted in stat.ME · 2026-01-08 · Jan Martin Wenkel, Michael Stanley Smith, Nadja Klein

Bayesian Additive Regression Tree Copula Processes for Scalable Distributional Prediction

We show how to construct the implied copula process of response values from a Bayesian additive regression tree (BART) model with prior on the leaf node variances. This copula process, defined on the covariate space, can be paired with any marginal distribution for the dependent variable to construct a flexible distributional BART...

💬 0 commentsarXiv:2601.04913v2PDF
0

Posted in stat.AP · 2026-01-08 · Zixuan Feng, Qiushi Chen, Paul Griffin, Le Bao

A Bayesian Multi-State Data Integration Approach for Estimating County-level Prevalence of Opioid Misuse in the United States

Drug overdose deaths, including from opioids, remain a significant public health threat to the United States (US). To abate the harms of opioid misuse, understanding its prevalence at the local level is crucial for stakeholders in communities to develop response strategies that effectively use limited resources. Although there exist...

💬 0 commentsarXiv:2601.04966v1PDF
0

Posted in stat.ML · 2026-01-08 · Anastasiia Bakhmach, Paul Dufossé, Simon Charpigny, Florence Monville, Laurent Greillier, Fabrice Barlési, Sébastien Benzekry

ROOFS: RObust biOmarker Feature Selection

Feature selection (FS) is essential for biomarker discovery and clinical predictive modeling. Over the past decades, methodological literature on FS has become rich and mature, offering a wide spectrum of algorithmic approaches. However, much of this methodological progress has not fully translated into applied biomedical research....

💬 0 commentsarXiv:2601.05151v3PDF
0

Posted in stat.ME · 2026-01-08 · Alex Ocampo, Enrico Giudice, Zachary R. McCaw, Tim P. Morris

Revealing the Truth: Calculating True Values in Causal Inference Simulation Studies via Gaussian Quadrature

Simulation studies are used to understand the properties of statistical methods. A key luxury in many simulation studies is knowledge of the true value (i.e. the estimand) being targeted. With this oracle knowledge in-hand, the researcher conducting the simulation study can assess across repeated realizations of the data how well a...

💬 0 commentsarXiv:2601.05128v1PDF
0

Posted in stat.ML · 2026-01-08 · James Rice

Stochastic Deep Learning: A Probabilistic Framework for Modeling Uncertainty in Structured Temporal Data

I propose a novel framework that integrates stochastic differential equations (SDEs) with deep generative models to improve uncertainty quantification in machine learning applications involving structured and temporal data. This approach, termed Stochastic Latent Differential Inference (SLDI), embeds an Itô SDE in the latent space of...

💬 0 commentsarXiv:2601.05227v1PDF
0

Posted in stat.ML · 2026-01-08 · Maja Waldron

CAOS: Conformal Aggregation of One-Shot Predictors

One-shot prediction enables rapid adaptation of pretrained foundation models to new tasks using only one labeled example, but lacks principled uncertainty quantification. While conformal prediction provides finite-sample coverage guarantees, standard split conformal methods are inefficient in the one-shot setting due to data splitting...

💬 0 commentsarXiv:2601.05219v2PDF
0

Posted in stat.AP · 2026-01-08 · Mellissa Meisels, Melody Huang, Tiffany M. Tang

Estimating Consensus Ideal Points Using Multi-Source Data

In the advent of big data and machine learning, researchers now have a wealth of congressional candidate ideal point estimates at their disposal for theory testing. Weak relationships raise questions about the extent to which they capture a shared quantity -- rather than idiosyncratic, domain-specific factors -- yet different measures...

💬 0 commentsarXiv:2601.05213v1PDF
0

Posted in stat.ME · 2026-01-08 · Yuchao Wang, Tianying Wang

Multi-Group Quadratic Discriminant Analysis via Projection

Multi-group classification arises in many prediction and decision-making problems, including applications in epidemiology, genomics, finance, and image recognition. Although classification methods have advanced considerably, much of the literature focuses on binary problems, and available extensions often provide limited flexibility...

💬 0 commentsarXiv:2601.05415v1PDF
0

Posted in stat.AP · 2026-01-08 · Aleix Alcacer, Irene Epifanio

Representing asymmetric relationships by h-plots. Discovering the archetypal patterns of cross-journal citation relationships

This work approaches the multidimensional scaling problem from a novel angle. We introduce a scalable method based on the h-plot, which inherently accommodates asymmetric proximity data. Instead of embedding the objects themselves, the method embeds the variables that define the proximity to or from each object. It is straightforward...

💬 0 commentsarXiv:2601.05400v1PDF
0

Posted in stat.ME · 2026-01-08 · Yezhuo Li, Fan Zhang, Dhanashree Shinde, Qiong Zhang, Sai Pradeep, Srikanth Pilla, Gang Li

Uncertainty Analysis of Experimental Parameters for Reducing Warpage in Injection Molding

Injection molding is a critical manufacturing process, but controlling warpage remains a major challenge due to complex thermomechanical interactions. Simulation-based optimization is widely used to address this, yet traditional methods often overlook the uncertainty in model parameters. In this paper, we propose a data-driven...

💬 0 commentsarXiv:2601.05396v1PDF
0

Posted in stat.ME · 2026-01-08 · Aleix Alcacer, Irene Epifanio

Archetypal cases for questionnaires with nominal multiple choice questions

Archetypal analysis serves as an exploratory tool that interprets a collection of observations as convex combinations of pure (extreme) patterns. When these patterns correspond to actual observations within the sample, they are termed archetypoids. For the first time, we propose applying archetypoid analysis to nominal observations,...

💬 0 commentsarXiv:2601.05392v1PDF