Qwen Councils

Statistics

arXiv preprints from January 1, 2026 through September 21, 2026 — 19:17:47 EST

0

Posted in stat.AP · 2026-01-15 · Mikkel Meyer Andersen, Nicole Huber, Kimberly S Andreaggi, Tóra Oluffa Stenberg Olsen, Walther Parson, Charla Marshall

MitoFREQ: A Novel Approach for Mitogenome Frequency Estimation from Top-level Haplogroups and Single Nucleotide Variants

Lineage marker population frequencies can serve as one way to express evidential value in forensic genetics. However, for high-quality whole mitochondrial DNA genome sequences (mitogenomes), population data remain limited. In this paper, we offer a new method, MitoFREQ, for estimating the population frequencies of mitogenomes....

💬 0 commentsarXiv:2601.10464v1PDF
0

Posted in stat.AP · 2026-01-15 · Glenna Nightingale, Karthik Mohan, Eloi Ribe, Valentin Popov, Shakes Wang, Clara Calia, Luciana Brondi, Sohan Seth

Modeling mental health trajectories during the COVID-19 pandemic using UK-wide data in the presence of sociodemographic variables

Background: The negative effects of the COVID-19 pandemic on the mental health and well-being of populations are an important public health issue. Our study aims to determine the underlying factors shaping mental health trajectories during the COVID-19 pandemic in the UK. Methods: Data from the Understanding Society COVID-19 Study...

💬 0 commentsarXiv:2601.10445v1PDF
0

Posted in stat.ME · 2026-01-15 · Yingying Ma, Chenlei Leng

A Propagation Framework for Network Regression

We introduce a unified and computationally efficient framework for regression on network data, addressing limitations of existing models that require specialized estimation procedures or impose restrictive decay assumptions. Our Network Propagation Regression (NPR) models outcomes as functions of covariates propagated through network...

💬 0 commentsarXiv:2601.10533v1PDF
0

Posted in stat.ML · 2026-01-15 · Francisco Madaleno, Pratik Misra, Alex Markham

Coarsening Causal DAG Models

Directed acyclic graphical (DAG) models are a powerful tool for representing causal relationships among jointly distributed random variables, especially concerning data from across different experimental settings. However, it is not always practical or desirable to estimate a causal model at the granularity of given features in a...

💬 0 commentsarXiv:2601.10531v2PDF
0

Posted in stat.ML · 2026-01-15 · Luke W. Yerbury, Ricardo J. G. B. Campello, G. C. Livingston, Mark Goldsworthy, Lachlan O'Neil

CROCS: A Two-Stage Clustering Framework for Behaviour-Centric Consumer Segmentation with Smart Meter Data

With grid operators confronting rising uncertainty from renewable integration and a broader push toward electrification, Demand-Side Management (DSM) -- particularly Demand Response (DR) -- has attracted significant attention as a cost-effective mechanism for balancing modern electricity systems. Unprecedented volumes of consumption...

💬 0 commentsarXiv:2601.10494v3PDF
0

Posted in stat.CO · 2026-01-15 · Constantin Vaillant Tenzer

Mesh Denoising

In this paper, we study four mesh denoising methods: linear filtering, a heat diffusion method, Sobolev regularization, and, to a lesser extent, a barycentric approach based on the Sinkhorn algorithm. We illustrate that, for a simple image denoising task, a naive choice of a Gibbs kernel can lead to unsatisfactory results. We...

💬 0 commentsarXiv:2601.10487v1PDF
0

Posted in stat.ME · 2026-01-15 · William L. Lippitt, Edward J. Bedrick, Nichole E. Carlson

Adjusted Similarity Measures and a Violation of Expectations

Adjusted similarity measures, such as Cohen's kappa for inter-rater reliability and the adjusted Rand index used to compare clustering algorithms, are a vital tool for comparing discrete labellings. These measures are intended to have the property of 0 expectation under a null distribution and maximum value 1 under maximal similarity...

💬 0 commentsarXiv:2601.10641v1PDF
0

Posted in stat.ML · 2026-01-15 · Eric Xia, Jason M. Klusowski

Classification Imbalance as Transfer Learning

Classification imbalance arises when one class is much rarer than the other. We frame this setting as transfer learning under label (prior) shift between an imbalanced source distribution induced by the observed data and a balanced target distribution under which performance is evaluated. Within this framework, we study a family of...

💬 0 commentsarXiv:2601.10630v1PDF
0

Posted in stat.ML · 2026-01-15 · Mihailo Stojnic

Parametric RDT approach to computational gap of symmetric binary perceptron

We study potential presence of statistical-computational gaps (SCG) in symmetric binary perceptrons (SBP) via a parametric utilization of \emph{fully lifted random duality theory} (fl-RDT) [96]. A structural change from decreasingly to arbitrarily ordered $c$-sequence (a key fl-RDT parametric component) is observed on the second...

💬 0 commentsarXiv:2601.10628v1PDF
0

Posted in stat.ME · 2026-01-15 · Yongzhen Feng, Weiwei Wang, Raymond K. W. Wong, Xianyang Zhang

Fair Regression under Demographic Parity: A Unified Framework

We propose a unified framework for fair regression tasks formulated as risk minimization problems subject to a demographic parity constraint. Unlike many existing approaches that are limited to specific loss functions or rely on challenging non-convex optimization, our framework is applicable to a broad spectrum of regression tasks....

💬 0 commentsarXiv:2601.10623v1PDF
0

Posted in stat.ME · 2026-01-15 · Paramahansa Pramanik, Arnab Kumar Maity, Anjan Mandal, Haley Kate Robinson

A Bayesian Discrete Framework for Enhancing Decision-Making Processes in Clinical Trial Designs and Evaluations

This study examines the application of Bayesian approach in the context of clinical trials, emphasizing their increasing importance in contemporary biomedical research. While conventional frequentist approach provides a foundational basis for analysis, it often lacks the flexibility to integrate prior knowledge, which can constrain...

💬 0 commentsarXiv:2601.10615v1PDF
0

Posted in stat.ME · 2026-01-15 · Zhangyi He, Feng Yu, Suzie Cro, Laurent Billot

From aggressive to conservative early stopping in Bayesian group sequential designs

Group sequential designs (GSDs) are widely used in confirmatory trials to allow interim monitoring while preserving control of the type I error rate. In the frequentist framework, O'Brien-Fleming-type stopping boundaries dominate practice because they impose highly conservative early stopping while allowing more liberal decisions as...

💬 0 commentsarXiv:2601.10590v1PDF
0

Posted in stat.ME · 2026-01-15 · Salvador V. Balkus, Hasan Laith, Nima S. Hejazi

On the use of cross-fitting in causal machine learning with correlated units

In causal machine learning, the fitting and evaluation of nuisance models are often performed on separate partitions, or folds, of the observed data. This technique, called cross-fitting, eliminates bias introduced by the use of black-box predictive algorithms. When study units may be correlated, such as in spatial, clustered, or...

💬 0 commentsarXiv:2601.10899v2PDF
0

Posted in stat.ME · 2026-01-15 · Simon Fontaine, Nisha J. D'Silva, Marcell Costa de Medeiros, Grace Y. Chen, Ji Zhu, Gen Li

Locally sparse varying coefficient mixed model with application to longitudinal microbiome differential abundance

Differential abundance (DA) analysis in microbiome studies has recently been used to uncover a plethora of associations between microbial composition and various health conditions. While current approaches to DA typically apply only to cross-sectional data, many studies feature a longitudinal design to better understand the underlying...

💬 0 commentsarXiv:2601.10872v1PDF
0

Posted in stat.ML · 2026-01-14 · Nick Polson, Vadim Sokolov

Horseshoe Mixtures-of-Experts (HS-MoE)

Horseshoe mixtures-of-experts (HS-MoE) models provide a Bayesian framework for sparse expert selection in mixture-of-experts architectures. We combine the horseshoe prior's adaptive global-local shrinkage with input-dependent gating, yielding data-adaptive sparsity in expert usage. Our primary methodological contribution is a particle...

💬 0 commentsarXiv:2601.09043v1PDF
0

Posted in stat.ME · 2026-01-14 · Kyusoon Kim, Hee-Seok Oh

Graph Canonical Coherence Analysis

We propose graph canonical coherence analysis (gCChA), a novel framework that extends canonical correlation analysis to multivariate graph signals in the graph frequency domain. The proposed method addresses challenges posed by the inherent features of graphs: discreteness, finiteness, and irregularity. It identifies pairs of...

💬 0 commentsarXiv:2601.09038v1PDF
0

Posted in stat.ML · 2026-01-14 · Kai Ming Ting, Ye Zhu, Hang Zhang, Tianrun Liang

Mass Distribution versus Density Distribution in the Context of Clustering

This paper investigates two fundamental descriptors of data, i.e., density distribution versus mass distribution, in the context of clustering. Density distribution has been the de facto descriptor of data distribution since the introduction of statistics. We show that density distribution has its fundamental limitation --...

💬 0 commentsarXiv:2601.10759v2PDF
0

Posted in stat.ME · 2026-01-14 · Pratim Guha Niyogi, Muraleetharan Sanjayan, Kathryn C. Fitzgerald, Ellen M. Mowry, Vadim Zipunnikov

Scalar-on-distribution regression via generalized odds with applications to accelerometry-assessed disability in multiple sclerosis

Distributional representations of data collected using digital health technologies have been shown to outperform scalar summaries for clinical prediction, with carefully quantified tail-behavior often driving the gains. Motivated by these findings, we propose a unified generalized odds (GO) framework that represents subject-specific...

💬 0 commentsarXiv:2601.09126v1PDF
0

Posted in stat.ME · 2026-01-14 · Dapeng Shi, Haoran Zhang, Tiandong Wang, Junhui Wang

A Multilayer Probit Network Model for Community Detection with Dependent Edges and Layers

Community detection in multilayer networks, which aims to identify groups of nodes exhibiting similar connectivity patterns across multiple network layers, has attracted considerable attention in recent years. Most existing methods are based on the assumption that different layers are either independent or follow specific dependence...

💬 0 commentsarXiv:2601.09161v2PDF
0

Posted in stat.AP · 2026-01-14 · Guodong Xu, Juan Du, Hui Yang

3D-SONAR: Self-Organizing Network for 3D Anomaly Ranking

Surface anomaly detection using 3D point cloud data has gained increasing attention in industrial inspection. However, most existing methods rely on deep learning techniques that are highly dependent on large-scale datasets for training, which are difficult and expensive to acquire in real-world applications. To address this...

💬 0 commentsarXiv:2601.09294v1PDF
0

Posted in stat.ME · 2026-01-14 · Ángel López-Oriona, Ying Sun, Hanlin Shang

White noise testing for functional time series via functional quantile autocorrelation

We introduce a novel class of nonlinear tests for serial dependence in functional time series, grounded in the functional quantile autocorrelation framework. Unlike traditional approaches based on the classical autocovariance kernel, the functional quantile autocorrelation framework leverages quantile-based excursion sets to robustly...

💬 0 commentsarXiv:2601.09371v2PDF
0

Posted in stat.ME · 2026-01-14 · Joanna Hindley, Charlotte Hartley, Jennifer Hellier, Kate Sturgeon, Sophie Greenwood, Ian Newsome, Katherine Barrett, Debs Smith, Tra My Pham, Dongquan Bi, Beatriz Goulao, Suzie Cro, Brennan C Kahan

Tools to help patients and other stakeholders' input into choice of estimand and intercurrent event strategy in randomised trials

Estimands can help to clarify the research questions being addressed in randomised trials. Because the choice of estimand can affect how relevant trial results are to patients and other stakeholders, such as clinicians or policymakers, it is important for them to be involved in these decisions. However, there are barriers to having...

💬 0 commentsarXiv:2601.09442v1PDF
0

Posted in stat.ME · 2026-01-14 · David Bamio, Jacobo de Uña-Álvarez

Smoothing spline density estimation from doubly truncated data

In Astronomy, Survival Analysis and Epidemiology, among many other fields, doubly truncated data often appear. Double truncation generally induces a sampling bias, so ordinary estimators may be inconsistent. In this paper, smoothing spline density estimation from doubly truncated data is investigated. For this purpose, an appropriate...

💬 0 commentsarXiv:2601.09576v1PDF
0

Posted in stat.ME · 2026-01-14 · Jonathan W. Bartlett, Dominic Magirr, Tim P. Morris

How to interpret hazard ratios

The hazard ratio, typically estimated using Cox's famous proportional hazards model, is the most common effect measure used to describe the association or effect of a covariate on a time-to-event outcome. In recent years the hazard ratio has been argued by some to lack a causal interpretation, even in randomised trials, and even if...

💬 0 commentsarXiv:2601.09571v1PDF
0

Posted in stat.ME · 2026-01-14 · Rongqian Zhang, Elena Tuzhilina, Jun Young Park

Sparse covariate-driven factorization of high-dimensional brain connectivity with application to site effect correction

Large-scale neuroimaging studies often collect data from multiple scanners across different sites, where variations in scanners, scanning procedures, and other conditions across sites can introduce artificial site effects. These effects may bias brain connectivity measures, such as functional connectivity (FC), which quantify...

💬 0 commentsarXiv:2601.09525v2PDF