Qwen Councils

Statistics

arXiv preprints from January 1, 2026 through September 19, 2026 — 13:15:11 EST

0

Posted in stat.ML · 2026-07-20 · Cheng Huan, Hongwei Yuan

An Adjoint-Sensitivity Framework for Lost-in-the-Middle Phenomena in Causal Residual Transformers

We develop an adjoint-sensitivity framework for positional influence in causal residual Transformers and separate unconditional analytic results from conditional boundary-shape conclusions. The principal unconditional theorem is the residual-to-depth-flow estimate for layer controls converging in $L^1$, complemented by a...

💬 0 commentsarXiv:2607.17696v1PDF
0

Posted in stat.ME · 2026-07-20 · Masahiro Kojima, Hisato Sunami, Masaaki Kuriki

A Globally Calibrated Bayesian Optimal Phase II Design for Adaptive Enrichment Trials

Adaptive enrichment can allow the development of an experimental treatment to continue when its activity is insufficient in an all-comer population but remains promising in a prespecified biomarker-positive subgroup. However, a straightforward sequential application of separately calibrated phase II designs to the two populations can...

💬 0 commentsarXiv:2607.17692v1PDF
0

Posted in stat.ME · 2026-07-20 · Martin Alexander Memmesheimer, Claudia Redenbach

Fitting the topology of synthetic particle systems with a novel graph representation

The shape and arrangement of particles in a material determine its macroscopic properties. The generation of synthetic data with varying particle structure, often represented as 3D voxel images, combined with simulation of macroscopic properties reveals structure-property relations. Most particle generation models focus on...

💬 0 commentsarXiv:2607.17680v1PDF
0

Posted in stat.ME · 2026-07-20 · Youmi Suk

Equality, Equity, and Causality in Fairness Research: A Commentary on Cheng (2026)

This is an invited commentary on the Psychometrika focus article "Fairness Issues and Evaluation in Psychometrics and AI/ML: What Can We Learn from Each Field?" by Ying Cheng (2026, doi:10.1017/psy.2026.10110). Cheng offers a systematic comparison between long-standing test fairness and modern algorithmic fairness. Her mapping of the...

💬 0 commentsarXiv:2607.17679v1PDF
0

Posted in stat.ML · 2026-07-20 · Yu Zhou, Yincai Tang, Bin Lv, Meng Gao

An efficient adaptive dimension selection algorithm for multidimensional probit graded response models

Multidimensional graded response models (MGRMs) are widely used for analyzing ordinal questionnaire data in psychological and educational assessments. A central challenge in applying these models is determining the number of latent dimensions. Conventional approaches usually fit multiple fixed-dimensional models and select among them...

💬 0 commentsarXiv:2607.17654v1PDF
0

Posted in stat.ME · 2026-07-20 · Zihan Li, Tiandong Wang

Spatial Dependence in Directed Preferential-Attachment Networks

Spatially embedded directed networks, such as airline networks, often exhibit simultaneous high activity at nearby nodes. Preferential attachment (PA) explains hub dominance. We extend it to spatial co-movement through a directed PA model whose out- and in-node weights follow temporally persistent Gaussian-process lognormal fields....

💬 0 commentsarXiv:2607.17597v1PDF
0

Posted in stat.ME · 2026-07-20 · Paul Rognon-Vael, David Rossell

E-Values For Multiplicity Control In Multiverse Analysis

Multiverse analysis refers to a common situation where one wishes to assess the association between multiple possible treatment definitions and multiple possible outcome definitions, potentially within multiple sub-populations, among other possible analysis specifications. Multiverse analysis is a useful exploratory tool to assess...

💬 0 commentsarXiv:2607.17596v1PDF
0

Posted in stat.ME · 2026-07-17 · Yuntang Fan, Paul Fearnhead, Idris A. Eckley, Gaetano Romano

An Efficient Likelihood Ratio Test for Online Changepoint Detection in the Presence of Autocorrelation

Changepoint detection methods have seen considerable development in recent years, with online algorithms capable of identifying structural changes in streaming data in near real time. However, the majority of existing methods are designed under the assumption of IID observations, rendering them susceptible to either more false...

💬 0 commentsarXiv:2607.16106v1PDF
0

Posted in stat.ME · 2026-07-17 · Anagh Chattopadhyay, Nilanjan Chatterjee

Improving Mendelian Randomization Analysis by Instrument Borrowing from Auxiliary Outcome Traits

Mendelian randomization (MR) is a widely used approach for inferring causal effects of exposures on outcomes using genetic variants as instrumental variables; however, existing methods remain vulnerable to bias and/or loss of power in the presence of invalid instruments. We hypothesize that closely related outcome traits are likely to...

💬 0 commentsarXiv:2607.16086v1PDF
0

Posted in stat.CO · 2026-07-17 · Seyoon Ko, Jasen Zhang, Andrew J. Holbrook

Scaling Hawkes Processes

Hawkes processes (HP) are a large class of stochastic point process models scientists have used to analyze contagion phenomena ranging from earthquakes, infectious diseases and biological neurons to financial trading activity, memes on social media and gun violence. We introduce applications of HP to the latter before reviewing...

💬 0 commentsarXiv:2607.16081v1PDF
0

Posted in stat.ML · 2026-07-17 · Claudia Skok Gibbs

Deep and Probabilistic Models for Gene Regulatory Network Inference

Gene regulatory networks (GRNs) link transcription factor (TF) proteins to their target genes, yet reconstructing these networks from genome-wide data remains challenging under practical and methodological constraints. Many methods couple modeling assumptions to a specific inference procedure and rely on heuristic model selection,...

💬 0 commentsarXiv:2607.16053v1PDF
0

Posted in stat.ME · 2026-07-17 · Michael Wieck-Sosa, Cosma Rohilla Shalizi

Dynamic models with $p$ parameters are identified by $2p+1$ random features

A foundational principle in nonlinear dynamics is that the structure of a dynamical system can be recovered from a small number of generic measurements or coordinates. We develop an analogous principle for the identification of dynamic models for time series {\em with noise}, which builds on previous identification results for...

💬 0 commentsarXiv:2607.16035v1PDF
0

Posted in stat.ML · 2026-07-17 · Nyi Nyi Aung, Heepeom Shin, Abigail Lawlor, Adrian Stein

Which Hyperparameters Matter? A Game-Theoretic Framework for Interpretable Hyperparameter Sensitivity Analysis

This work presents a game-theoretic framework for interpretable hyperparameter-objective interaction analysis rather than proposing a new optimization algorithm. In the proposed framework, Shapley Effects are employed for global sensitivity analysis, while Pareto front sets are utilized to identify effective hyperparameter...

💬 0 commentsarXiv:2607.15884v1PDF
0

Posted in stat.ME · 2026-07-17 · Antonin Schrab, Rajen Shah, Arthur Gretton, Ilmun Kim

Aggregation of Statistical Evidence under Exchangeability

We study aggregation of statistical evidence under unknown and potentially complex dependence using group-invariance. Building on permutation-based constructions that treat transformed datasets as exchangeable units, we aggregate evidence across statistics for each transformed dataset and calibrate the resulting aggregates across...

💬 0 commentsarXiv:2607.15823v1PDF
0

Posted in stat.AP · 2026-07-17 · Fan Yang, Lin Zhang

LLM Latent Edge Measurement: Point-in-Time Economic Graphs for Quantitative Investing from Corporate Disclosures

Standard industry classification systems such as GICS assign each firm to a single sector, but the economic relationships through which shocks propagate, such as supplier agreements, customer concentration, intellectual property licensing, cloud service dependencies, and power purchase contracts frequently cross sector boundaries and...

💬 0 commentsarXiv:2607.15640v1PDF
0

Posted in stat.ME · 2026-07-17 · Marie Turčičová, Patrícia Martinková

Asymptotically exact threshold for detecting anomalies in multivariate Gaussian data with application to time series

In this paper, we propose a new thresholding technique for detecting anomalies in multivariate normal random samples, under the assumption that anomalous observations are sparse and differ from the rest of the data in their mean. The mean vector of the non-anomalous data is assumed to be zero, while the covariance matrix is unknown....

💬 0 commentsarXiv:2607.15637v1PDF
0

Posted in stat.ML · 2026-07-17 · Moritz Hardt

Retraining Seeks Stable Signals

Predictive models deployed at scale influence future data, a phenomenon called performativity. And there is always one way to cope: Train the model on new data, deploy it again, and repeat. This process, called retraining or repeated risk minimization, creates a feedback loop between model and data that real-world learning systems...

💬 0 commentsarXiv:2607.15623v1PDF
0

Posted in stat.CO · 2026-07-16 · Renny Doig, Liangliang Wang

Compound Auxiliary Metropolis: Incorporating Auxiliary Variables into Multi-Candidate MCMC

Multiple-try Metropolis (MTM) is a Markov chain Monte Carlo (MCMC) algorithm that improves local transition efficiency by evaluating multiple candidate draws at each iteration. However, for complicated target distributions exhibiting severely non-Gaussian topography or multiple well-separated modes, locally optimal transitions may be...

💬 0 commentsarXiv:2607.15499v1PDF
0

Posted in stat.ME · 2026-07-16 · Ebrahim Khaled Ebrahim, Ahmed El-Kotory

A directional Hosmer-Lemeshow goodness-of-fit test for sparse logistic regression

Goodness-of-fit assessment for the binary logistic regression model is difficult when covariates are continuous: the data are effectively sparse, the classical Pearson and deviance tests fail, and practitioners rely on partition-based tests, such as the Hosmer-Lemeshow test, that group observations before comparing observed and...

💬 0 commentsarXiv:2607.15454v1PDF
0

Posted in stat.AP · 2026-07-16 · QIan Cheng, Nilay Tanik Argon, Aniruddhan Ganesaraman, Serhan Ziya

Proactive Inpatient Bed Requests for Emergency Department Admissions

Emergency department (ED) boarding occurs when admitted patients remain in the ED while awaiting inpatient beds. Boarding is a major driver of ED crowding and has been associated with poor patient outcomes. We propose a framework to help EDs reduce boarding time and length of stay by using information about current patients and bed...

💬 0 commentsarXiv:2607.15432v1PDF
0

Posted in stat.ME · 2026-07-16 · Melissa Lynne Martin, Theodore D. Satterthwaite, Ian J. Barnett

Sequential Control of False Positives in Online Change Point Detection

Online change point detection is the process of identifying distributional changes in time-ordered data in real time. In applications such as mobile health (mHealth), repeated testing is often performed as new data arrive, creating a multiple testing problem. Traditional approaches for controlling the family-wise error rate (FWER) are...

💬 0 commentsarXiv:2607.15423v1PDF
0

Posted in stat.ML · 2026-07-16 · Robert Chew, Matthew R. Williams

Design-Based Supervised Learning with Noisy Human Labels

Researchers increasingly use automated classifiers to label unstructured data for statistical analysis. Existing rectification methods can correct errors in these automated labels using a probability-sampled audit set, but they usually treat the audit labels as correct. In practice, human audit labels are often noisy, and only some...

💬 1 commentsarXiv:2607.15455v1PDF
0

Posted in stat.ML · 2026-07-17 · Gabriel Samberg, YoonHaeng Hur, Yuehaw Khoo, Nir Sharon

Cluster-Aware Matching via Laplacian Optimal Transport

In many applications of matching, the point clouds to be matched are not merely unstructured sets of points but rather samples from distributions with an intrinsic cluster structure. In such cases, as individual points are often interchangeable within a coherent region, finding a robust region-to-region alignment is more desirable...

💬 0 commentsarXiv:2607.16178v1PDF
0

Posted in stat.ME · 2026-07-14 · Alberto Quaini, Chen Zhou

Anchored Geodesic Analysis for Multivariate Extremes

Extremal dependence is naturally described by the angular law of large multivariate observations. We introduce anchored geodesic component analysis (AGCA), a dimension-reduction method for extremal angular laws on the positive unit sphere. AGCA approximates angular variation by great subspheres constrained to pass through a chosen...

💬 0 commentsarXiv:2607.13112v1PDF
0

Posted in stat.AP · 2026-01-21 · Li Tuobang

Implementing Substance Over Form: A Novel Metric for Taxing E-commerce to Address Deterritorialization

Against the backdrop of e-commerce restructuring consumption patterns, last-mile delivery stations have substantially fulfilled the function of community retail distribution. However, the current tax system only levies a low labor service tax on delivery fees, resulting in a tax contribution from the massive circulating goods value...

💬 0 commentsarXiv:2601.14616v1PDF