Ann. Appl. Stat. 20 (3), (September 2026)
No abstract available
Ann. Appl. Stat. 20 (3), (September 2026)
No abstract available
Causal Inference, Design and Policy
Zhiyu Kang, Li Chen, Peng Wei, Zhichao Xu, Chunlin Li, Tianzhong Yang
Ann. Appl. Stat. 20 (3), 1789-1808, (September 2026) DOI: 10.1214/26-AOAS2161
KEYWORDS: causal mediation analysis, High-dimensional mediators, omics mediators, selection bias, total mediation effect
Mediation analysis helps uncover how exposures impact outcomes through intermediate variables. Traditional mean-based total mediation effect measures may suffer from the cancellation of opposite componentwise effects, and existing methods often lack the power to capture weak effects in high-dimensional mediators. Additionally, most high-dimensional mediation analysis methods have focused on continuous outcomes, with limited attention to binary outcomes, particularly in case-control studies. To fill this gap, we propose an total mediation effect measure within the liability framework that offers a clear and intuitive causal interpretation, provides additional insights beyond the mean-based measures, and is invariant to disease prevalence. We develop a cross-fitted, modified Haseman–Elston regression-based estimation procedure tailored for mediation analysis in case-control studies, which can also be applied to cohort studies. Our estimator remains consistent in the presence of nonmediators and weak effects, as demonstrated in extensive simulations. Theoretical justification for consistency is provided under mild conditions and without requiring exact mediator selection. In a case-control substudy of the Women’s Health Initiative involving 2150 individuals, we found that many metabolites were mediators with weak effects in the path from BMI to coronary heart disease, and we estimated that 89% (95% CI: 57%–100%) of the BMI-explained variation in underlying CHD liability is mediated by the measured metabolomics. The proposed estimation procedure is implemented in the R package “r2MedCausal,” available on GitHub.
Ying Yang, Chengshun Shi, Fang Yao, Shouyang Wang, Hongtu Zhu
Ann. Appl. Stat. 20 (3), 1809-1831, (September 2026) DOI: 10.1214/26-AOAS2173
KEYWORDS: A/B testing, policy evaluation, reinforcement learning, spatially randomized designs
This article studies the benefits of using spatially randomized experimental designs which partition the experimental area into nonoverlapping units with treatments assigned randomly to these units. Such designs improve policy evaluation in online experiments by providing more precise policy value estimators and more effective testing algorithms than traditional global designs, which apply the same treatment across all units simultaneously. We examine both parametric and nonparametric methods for estimating and inferring policy values based on the spatially randomized designs. Our analysis includes evaluating the mean squared error of the treatment effect estimator and the statistical power of the associated tests. Additionally, we extend our findings to the dynamic setting with spatiotemporal dependencies, where treatments are allocated sequentially over time, and account for potential temporal carryover effects. Our theoretical insights are supported by comprehensive numerical experiments.
Georgia Papadogeorgou, Srijata Samanta
Ann. Appl. Stat. 20 (3), 1832-1852, (September 2026) DOI: 10.1214/26-AOAS2174
KEYWORDS: Bayesian causal inference, interference, potential outcomes, spatial confounding, spatial causal inference, unmeasured confounding
This manuscript unites causal inference and spatial statistics, presenting novel insights for causal inference in spatial data analysis, and drawing from tools in spatial statistics to estimate causal effects. We introduce spatial causal graphs to highlight that spatial confounding and interference can be entangled, in that investigating the presence of one can lead to wrongful conclusions in the presence of the other. Moreover, we show that spatial dependence in the exposure variable can render standard analyses invalid. To remedy these issues, we propose a Bayesian parametric approach based on tools commonly used in spatial statistics. This approach simultaneously accounts for interference and mitigates bias from local and neighborhood unmeasured spatial confounding. From a Bayesian perspective, we show that incorporating an exposure model is necessary. Under a specific model formulation, we prove that all parameters are identifiable including the causal effects, even in the presence of unmeasured confounding. We illustrate the approach with a simulation study. We evaluate the effect of local and neighboring sulfur dioxide emissions from power plants on county-level cardiovascular mortality from observational spatial data in the United States, where unmeasured spatial confounding and interference might be present simultaneously.
Dongjae Son, Brian J. Reich, Erin M. Schliep, Shu Yang, David A. Gill
Ann. Appl. Stat. 20 (3), 1853-1872, (September 2026) DOI: 10.1214/26-AOAS2187
KEYWORDS: Poisson process, potential outcomes, propensity scores, spatial confounding
Marine Protected Areas (MPAs) have been established globally to conserve marine resources. Given their maintenance costs and impact on commercial fishing, it is critical to evaluate their effectiveness to support future conservation. In this paper, we use data collected from the Australian coast to estimate the effect of MPAs on biodiversity. Environmental studies such as these are often observational, and processes of interest exhibit spatial dependence, which present challenges in estimating the causal effects. Spatial data can also be subject to preferential sampling, where the sampling locations are related to the policy and the response variable, further complicating inference and prediction. To address these challenges, we propose a spatial causal inference method that simultaneously accounts for unmeasured spatial confounders in both the sampling process and the treatment allocation. We prove the identifiability of key parameters in the model and the consistency of the posterior distributions of those parameters. We show via simulation studies that the causal effect of interest can be reliably estimated under the proposed model. The proposed method is applied to assess the effect of MPAs on fish biomass. We find evidence of preferential sampling and that properly accounting for this source of bias impacts the estimate of the causal effect.
Samhita Pal, Dhrubajyoti Ghosh, Shu Yang
Ann. Appl. Stat. 20 (3), 1873-1893, (September 2026) DOI: 10.1214/26-AOAS2191
KEYWORDS: Causal discovery penalty, Directed acyclic graph, Sparsity, PPMI, proteomics
Parkinson’s disease (PD) is a progressive neurodegenerative disorder that lacks reliable early-stage biomarkers for diagnosis, prognosis, and therapeutic monitoring. While cerebrospinal fluid (CSF) biomarkers, such as α-synuclein seed amplification assays (αSyn-SAA), offer diagnostic potential, their clinical utility is limited by invasiveness and incomplete specificity. Plasma biomarkers provide a minimally invasive alternative, but their mechanistic role in PD remains unclear. A major challenge is distinguishing whether plasma biomarkers causally reflect primary neurodegenerative processes or are downstream consequences of disease progression. To address this, we leverage the Parkinson’s Progression Markers Initiative (PPMI) Project 9000, containing 2,924 plasma and CSF biomarkers, to systematically infer causal relationships with disease status. However, only a sparse subset of these biomarkers and their interconnections are actually relevant for the disease. Existing causal discovery algorithms, such as Fast Causal Inference (FCI) and its variants, struggle with the high dimensionality of biomarker datasets under sparsity, limiting their scalability. We propose Penalized Fast Causal Inference (PFCI), a novel approach that incorporates sparsity constraints to efficiently infer causal structures in large-scale biological datasets. By applying PFCI to PPMI data, we aim to identify biomarkers that are causally linked to PD pathology, enabling early diagnosis and patient stratification. Our findings will facilitate biomarker-driven clinical trials and contribute to the development of neuroprotective therapies.
Gabriel Durham, Anil Battalahalli, Amy M. Kilbourne, Andrew Quanbeck, Wenchu Pan, Tim Lycurgus, Daniel Almirall
Ann. Appl. Stat. 20 (3), 1894-1916, (September 2026) DOI: 10.1214/26-AOAS2200
KEYWORDS: Adaptive intervention, sequential multiple assignment randomized trial, cluster-randomized trial, longitudinal outcomes, Marginal Structural Models
Many health policies or programs can be conceptualized as adaptive interventions. An adaptive intervention is a sequence of decision rules that guide the provision of actions (intervention options) at critical decision points based on the evolving need of recipients, including their response to prior actions. In many health policy settings, adaptive interventions target a population of clusters (e.g., schools), with the ultimate intent of impacting outcomes at the level of individuals within the clusters (e.g., mental health care providers in the schools). Health policy researchers can use clustered, sequential, multiple assignment, randomized trials (SMARTs) to answer important scientific questions concerning clustered adaptive interventions. A common primary aim is to compare the mean of a nested, end-of-study outcome between two clustered adaptive interventions. However, existing methods are not suitable when the primary outcome in a clustered SMART is nested and longitudinal (e.g., repeated outcome measures nested within mental healthcare providers, and mental healthcare providers nested within schools). This manuscript proposes a three-level marginal mean modeling and estimation approach for comparing adaptive interventions in a clustered SMART. The proposed method enables policy analysts to answer a wider array of scientific questions in the marginal comparison of clustered adaptive interventions. Further, relative to using existing two-level methods with a nested, but nonlongitudinal, end-of-study outcome, the proposed method benefits from improved statistical efficiency. With this approach we examine longitudinal comparisons of adaptive interventions for improving school-based mental healthcare and contrast its performance with existing approaches for studying static end-of-study outcomes. Methods were motivated by the Adaptive School-Based Implementation of CBT (ASIC) study, a clustered SMART designed to construct an adaptive health policy to improve the adoption of evidence-based CBT by mental healthcare professionals in high schools across Michigan.
Zhichao Jiang, Eli Ben-Michael, D. James Greiner, Ryan Halen, Kosuke Imai
Ann. Appl. Stat. 20 (3), 1917-1941, (September 2026) DOI: 10.1214/26-AOAS2201
KEYWORDS: Dropout, efficient influence function, sequential ignorability, stochastic intervention, truncation by death
Dropout presents a major challenge for causal inference in longitudinal studies with time-varying treatments. We focus on selective eligibility—an important but overlooked source of nonignorable dropout that occurs when a unit’s prior treatment history affects its eligibility for future treatments. Our motivating application evaluates a pretrial risk assessment instrument in the criminal justice system, where cases were randomly assigned to receive a risk assessment report. Selective eligibility arises because individuals can be rearrested over time, and rearrest depends on prior judicial decisions, which may be influenced by earlier treatment assignments. We develop a general methodological framework for longitudinal causal inference under selective eligibility. By conditioning on subgroups of units who would remain eligible for treatment under a given treatment history, we define time-specific eligible treatment effects and the expected total outcome under a given treatment sequence. Assuming a generalized form of sequential ignorability, we derive nonparametric identification formulas and the efficient influence functions for both estimands, yielding doubly robust estimators. Applying this methodology, we find that risk assessment reports increase alignment between judicial decisions and algorithmic recommendations, particularly when reports are provided repeatedly. However, the reports have a limited impact on downstream criminal justice outcomes.
Saijun Zhao, Peter F. Thall, Yong Zang
Ann. Appl. Stat. 20 (3), 1942-1960, (September 2026) DOI: 10.1214/26-AOAS2219
KEYWORDS: Bayesian statistics, clinical trial, dose finding, phase 1–2 clinical trial, phase 3 clinical trial
A family of Bayesian generalized phase 1–2 oncology trial designs recently was proposed that uses a phase 1–2 design based on early response and toxicity to choose a set of candidate doses rather than one dose, randomizes additional patients among the candidates, and selects a best dose based on progression-free survival time over longer follow-up. This was extended to a generalized phase 1–2–3 design by including possible expansion to a phase 3 trial of the new agent, at its selected optimal dose, compared to standard of care. The decision of whether to conduct phase 3 was based on the predictive probability of a positive trial. This approach was shown to be greatly superior to the conventional three-phase paradigm. In this paper we extend the generalized phase 1–2–3 design to accommodate patient subgroups by doing subgroup-specific dose-finding and phase 3 treatment comparisons. The proposed design adaptively clusters subgroups having similar estimated dose-outcome distributions. While applying this proposed design requires substantial planning, trial conduct is very similar to running a phase 1–2 trial followed by a phase 3 trial. A simulation study shows that the proposed design has much more desirable properties than either ignoring subgroups or running separate trials within subgroups.
Jeongjin Lee, Junke Yang, Grzegorz A. Rempala, Patrick M. Schnell
Ann. Appl. Stat. 20 (3), 1961-1984, (September 2026) DOI: 10.1214/26-AOAS2220
KEYWORDS: Causal inference, infectious disease modeling, repeated testing
This paper addresses the problem of estimating infectious disease prevalence under longitudinal testing programs that include scheduled, symptomatic, and contact-tracing testing. Our study is motivated by data from The Ohio State University, where a mandatory once-per-week COVID-19 testing and isolation program was implemented during the Fall 2020 semester, supplemented by additional testing for symptomatic individuals and identified contacts. In this setting, the probability of being tested depends on symptoms or contact-tracing status, creating a complex observation process. We develop a counterfactual framework that links the observation process to a hypothetical process in which infection is prevented. This formulation enables unbiased estimation of disease prevalence by modeling the testing process, possibly nonparametrically, without requiring explicit modeling of transmission dynamics, even though the testing and infection processes are jointly dependent.
Indrabati Bhattacharya, Debajyoti Sinha, Durbadal Ghosh, George Rust
Ann. Appl. Stat. 20 (3), 1985-2004, (September 2026) DOI: 10.1214/26-AOAS2233
KEYWORDS: Bayesian inference, Causal inference, survival data, spatial data
We propose a novel Bayesian approach to estimate causal effects in spatially clustered survival data. Using Soft Bayesian Additive Regression Trees (SBART), we introduce a nonparametric regression for a log-Normal survival model that accommodates spatial associations among unknown cluster effects through a directed acyclic graph autoregressive (DAGAR) model. We employ a two-stage approach, which entails estimating the propensity score in the first step and incorporating it as a confounder of the outcome model in the second step. In our simulation study, we compare our method with existing approaches under various simulation scenarios, including both correctly specified and misspecified outcome models, to demonstrate the superior performance of our method. We then apply our method to analyze the causal effect of treatment delay (TD) on posttreatment survival of breast cancer patients from the Florida Cancer Registry (FCR). Our analysis produces the county-specific as well as state-wide assessment of the causal effects while accommodating spatial association among counties.
Yanke Song, Xihong Lin, Pragya Sur
Ann. Appl. Stat. 20 (3), 2005-2026, (September 2026) DOI: 10.1214/26-AOAS2152
KEYWORDS: Heritability estimation, fixed effects method, high-dimensional asymptotics, signal-to-noise ratio estimation, debiased estimators, ensemble selection
Estimating heritability using a large number of SNPs in Genome-Wide Association Studies (GWAS) is of substantial interest. Various approaches have emerged over the years, broadly categorized as either random effects or fixed effects heritability methods. These methods are sensitive to their assumptions on the underlying signal structure. We propose a robust fixed effect-based ensemble approach to estimate heritability or the signal-to-noise ratio in high-dimensional linear models where the sample size and dimension grow proportionally. Our method ensembles post-processed versions of the debiased lasso and debiased ridge estimators, and incorporates a data-driven strategy for hyperparameter selection that significantly boosts estimation performance. We establish rigorous consistency guarantees that hold despite adaptive tuning. Extensive simulations demonstrate our method’s superiority over state-of-the-art methods across various signal structures and genetic architectures, ranging from sparse to relatively dense and from evenly to unevenly distributed signals. Furthermore, we discuss advantages of fixed effects heritability estimation compared to random effects estimation. Our theoretical guarantees hold for realistic distributions observed in genetic studies, where genotypes typically take on discrete values and are often well modeled by sub-Gaussian random variables. We establish our theoretical results by deriving uniform bounds, built upon the convex Gaussian min-max theorem, and leveraging universality results. Finally, we showcase the efficacy of our approach in estimating height and BMI heritability using the GWAS data from the UK Biobank.
Joel Eliason, Arvind Rao, Timothy L. Frankel, Michele Peruzzi
Ann. Appl. Stat. 20 (3), 2027-2052, (September 2026) DOI: 10.1214/26-AOAS2178
KEYWORDS: hierarchical Bayesian modeling, cell-cell interactions, tumor microenvironment, spatial correlation functions, log-Gaussian Cox processes
The tumor microenvironment (TME) is a spatially heterogeneous ecosystem where cellular interactions shape tumor progression and response to therapy. Multiplexed imaging technologies enable high-resolution spatial characterization of the TME, yet statistical methods for analyzing multi-subject spatial tissue data remain limited. We propose a Bayesian hierarchical model for inferring spatial dependencies in multiplexed imaging datasets across multiple subjects. Our model represents the TME as a multivariate log-Gaussian Cox process, where spatial intensity functions of different cell types are governed by a latent multivariate Gaussian process. By pooling information across subjects, we estimate spatial correlation functions that capture within-type and cross-type dependencies, enabling interpretable inference about disease-specific cellular organization. We validate our method using simulations, demonstrating robustness to latent factor specification and spatial resolution. We apply our approach to two multiplexed imaging datasets: pancreatic cancer and colorectal cancer, revealing distinct spatial organization patterns across disease subtypes and highlighting tumor-immune interactions that differentiate immune-permissive and immune-exclusive microenvironments. These findings provide insight into mechanisms of immune evasion and may inform novel therapeutic strategies. Our approach offers a principled framework for modeling spatial dependencies in multisubject data, with broader applicability to spatially resolved omics and imaging studies. An R package, available online, implements our methods.
Bingying Dai, Yinan Lin, Kejue Jia, Zhao Ren, Wen Zhou
Ann. Appl. Stat. 20 (3), 2053-2073, (September 2026) DOI: 10.1214/26-AOAS2207
KEYWORDS: Nodewise multinomial regression, Potts model, protein mutation effects, protein structure, sparse group lasso
Quantifying the effects of amino acid mutations in proteins presents a significant challenge due to the vast combinations of residue sites and amino acid types, making experimental approaches costly and time-consuming. The Potts model has been used to address this challenge, with parameters capturing evolutionary dependency between residue sites within a protein family. However, existing methods often use the mean-field approximation to reduce computational demands, which lacks provable guarantees and overlooks critical structural information for assessing mutation effects. We propose a new framework for analyzing protein sequences using the Potts model with nodewise high-dimensional multinomial regression. Our method identifies key residue interactions and important amino acids, quantifying mutation effects through evolutionary energy derived from model parameters. It encourages sparsity in both sitewise and amino acidwise dependencies through elementwise and group sparsity. We have established, for the first time to our knowledge, the convergence rate for estimated parameters in the high-dimensional Potts model using sparse group Lasso, matching the existing minimax lower bound for high-dimensional linear models with a sparse group structure, up to a factor depending only on the multinomial nature of the Potts model. This theoretical guarantee enables accurate quantification of estimated energy changes. Additionally, we incorporate structural data into our model by applying penalty weights across site pairs. Our method outperforms others in predicting mutation fitness, as demonstrated by comparisons with high-throughput mutagenesis experiments across 12 protein families.
Hong Zhang, Ming Liu, John E. Landers, Zheyang Wu
Ann. Appl. Stat. 20 (3), 2074-2097, (September 2026) DOI: 10.1214/26-AOAS2223
KEYWORDS: data integration, p-value combination, signal detection, weighting, whole genome sequencing study
The integrative association test utilizes a weighting scheme to combine prior information and increase statistical power. In whole-genome sequencing (WGS) studies, it facilitates the integration of biological characteristics of single nucleotide variants (SNVs) to improve the detection of novel disease genes. Despite recent applied advances, determining the best weights to fully leverage relevant information remains an open question. For a broad family of weighted integrative tests, this paper proposes optimized weights that maximize the tests’ asymptotic efficiency, a dominant metric influencing statistical power. The study elucidates how weighting enhances statistical power and designs a practical approach for integrating effective information from SNV allele frequencies, annotations, and linkage disequilibrium. Extensive simulations demonstrate improved statistical power compared to existing methods. An osteoporosis case study further illustrates the method’s application and potential for detecting additional novel disease genes.
Arhit Chakrabarti, Yang Ni, Yuchao Jiang, Bani K. Mallick
Ann. Appl. Stat. 20 (3), 2098-2124, (September 2026) DOI: 10.1214/26-AOAS2224
KEYWORDS: Bayesian nonparametrics, common atoms model, nested dataset, single-cell data, variational inference
We consider the problem of clustering nested or hierarchical data, where observations are grouped and there are both group-level and observation-level variables. In our motivating OneK1K dataset, observations consist of single-cell RNA-sequencing (scRNA-seq) data from 982 individuals (groups), totaling 1.27 million cells (observations), along with individual-specific genotype data. This type of data would enable the identification of cell types and the investigation of how genetic variations among individuals influence differences in cell-type profiles. Our goal, therefore, is to jointly cluster cells and individuals to capture the heterogeneity across both levels using cell-specific gene expressions as well as individual-specific genotypes. However, existing grouped clustering methods do not incorporate group-level variables, thereby limiting their ability to capture the heterogeneity of genotypes in our motivating application. To address this, we propose the nested atoms model (NAM), a new Bayesian nonparametric approach that enables the desired two-layered clustering, accounting for both group-level and observation-level variables. To scale NAM for high-dimensional data, we develop a fast variational Bayesian inference algorithm. Simulations show that NAM outperforms existing methods that ignore group-level variables. Applied to the OneK1K dataset, NAM identifies clusters of genetically similar individuals with homogeneous cell-type profiles. The resulting cell clusters align with known immune cell types based on differential gene expression, underscoring the ability of NAM to capture nested heterogeneity and provide biologically meaningful insights.
Michelle Pistner Nixon, Kyle C. McGovern, Maxwell A. Konnaris, Jeffrey Letourneau, Lawrence A. David, Nicole A. Lazar, Sayan Mukherjee, Justin D. Silverman
Ann. Appl. Stat. 20 (3), 2125-2147, (September 2026) DOI: 10.1214/26-AOAS2230
KEYWORDS: multivariate count data, partially identified models, Compositional data analysis, microbiome data
Many scientific fields, including human gut microbiome science, collect multivariate count data where the sum of the counts is unrelated to the scale of the underlying system being measured (e.g., total microbial load in a subject’s colon). This disconnect complicates downstream analyses such as differential analysis in case-control studies. This article is motivated by a novel study of in vitro human gut microbiome models. Popular tools for analyzing these data led to dramatically elevated rates of both false positives and false negatives. To understand those failures, we provide a formal problem statement that frames these challenges of scale in terms of the classical theory of identifiability. We call this the problem of scale reliant inference (SRI). We use this formulation to prove fundamental limits on SRI in terms of criteria such as consistency and type-I error control. We show that the failures of existing methods stem from a fundamental failure to properly quantify uncertainty in the system scale. We demonstrate that a particular type of Bayesian model, called a Bayesian partially identified model (PIMs), can correctly quantify uncertainty in SRI. We introduce scale simulation random variables (SSRVs) as a flexible and efficient approach to specifying and inferring Bayesian PIMs. In the context of both real and simulated data, we find SSRVs drastically decrease type-I and type-II error rates.
Economics and Finance
Danyang Huang, Ziyi Kong, Shuyuan Wu, Hansheng Wang
Ann. Appl. Stat. 20 (3), 2148-2170, (September 2026) DOI: 10.1214/26-AOAS2205
KEYWORDS: Spatial autoregressive model, privacy protection, bias-corrected estimation, Least squares estimation
Spatial autoregressive (SAR) models and their extensions are important tools for studying network effects. However, with an increasing emphasis on data privacy, data providers often implement protection measures that render standard SAR models inapplicable. In this study, we introduce a privacy-protected SAR model that incorporates noise into both the response and covariates to meet privacy requirements. With noise present in both components, the traditional quasi-maximum likelihood estimator becomes difficult to compute because the likelihood function cannot be directly formulated. To bypass this hurdle, we begin with a pseudolikelihood approach, initially omitting the noise in the covariates. A Newton–Raphson algorithm is then applied to compute the estimator; however, the estimator is biased. To address this, we propose a bias-corrected Newton–Raphson-type algorithm that simultaneously accounts for noise in both the response and covariates. We further show, under appropriate regularity conditions, that the resulting estimator is consistent and asymptotically normal. To further enhance computational efficiency, we also develop a bias-corrected least squares estimator. Several extensions are discussed, and the finite-sample performance of the proposed methods is evaluated through extensive simulations. We apply the proposed methodology to restaurant transaction data from a third-party payment platform. Our method identifies a statistically significant competitive network effect among restaurants and further reveals meaningful restaurant-customer interaction patterns.
Niko Hauzenberger, Florian Huber, Gary Koop, James Mitchell
Ann. Appl. Stat. 20 (3), 2171-2194, (September 2026) DOI: 10.1214/26-AOAS2217
KEYWORDS: Bayesian vector autoregression, time-varying parameters, nonparametric modeling, machine learning, regression trees, business cycle shocks
In light of widespread evidence of parameter instability in macroeconomic models, many time-varying parameter (TVP) models have been proposed. This paper develops a semiparametric TVP vector autoregression (VAR) using Bayesian additive regression trees (BART) that models the TVPs as an unknown function of effect modifiers. The novelty of this model arises from the fact that the law of motion driving the parameters is treated nonparametrically. This leads to great flexibility in the nature and extent of parameter change, both in the conditional mean and in the conditional variance. Parsimony is achieved through adopting nonparametric factor structures and use of shrinkage priors. In an application to U.S. macroeconomic data, we illustrate the use of our model in tracking the evolving nature of the Phillips curve trade-off between inflation and unemployment in response to business cycle shocks.
William S. Daniels, Douglas W. Nychka, Dorit M. Hammerling
Ann. Appl. Stat. 20 (3), 2195-2213, (September 2026) DOI: 10.1214/26-AOAS2202
KEYWORDS: Climate change, Methane emissions, Bayesian inverse modeling, Gibbs sampling, spike-and-slab prior
Reducing methane emissions from the oil and gas sector is a key component of short-term climate action. Emission reduction efforts are often conducted at the individual site-level, where being able to apportion emissions between a finite number of potentially emitting equipment is necessary for leak detection and repair as well as regulatory reporting of annualized emissions. We present a hierarchical Bayesian model, referred to as the multisource detection, localization, and quantification (MDLQ) model, for performing source apportionment on oil and gas sites using methane measurements from point sensor networks. The MDLQ model accounts for autocorrelation in the sensor data and enforces sparsity in the emission rate estimates via a spike-and-slab prior, as oil and gas equipment often emit intermittently. We use the MDLQ model to apportion methane emissions on an experimental oil and gas site designed to release methane in known quantities, providing a means of model evaluation. Data from this experiment are unique in their size (i.e., the number of controlled releases) and in their close approximation of emission characteristics on real oil and gas sites. As such, this study provides a baseline level of apportionment accuracy that can be expected when using point sensor networks on operational sites.
M. Daniela Cuba, Craig Wilkie, Marian Scott, Daniela Castro-Camilo
Ann. Appl. Stat. 20 (3), 2214-2234, (September 2026) DOI: 10.1214/26-AOAS2238
KEYWORDS: Air quality monitoring, Dirac-delta generalised Pareto distribution, Extreme value theory, hierarchical Bayesian model, spatiotemporal data fusion, threshold-exceedances
Data fusion models are widely used in air quality monitoring to integrate in situ and large-scale gridded products, offering spatially complete and temporally detailed estimates. However, traditional Gaussian-based models often underestimate extreme pollution values, leading to biased risk assessments. To address this, we present a Bayesian hierarchical data fusion framework rooted in extreme value theory, using the Dirac-delta generalised Pareto distribution to jointly account for threshold and nonthreshold exceedances while preserving the timing of exceedance and nonexceedance episodes. Our model is used to describe and predict censored threshold exceedances of pollution in the Greater London region by using CAMS atmospheric composition reanalysis and in situ observation stations from the automatic urban and rural network (AURN) run by the U.K. government. Key features of our approach include combining data with varying spatiotemporal resolutions and fully accounting for parameter uncertainties. Results show that our model outperforms log-Gaussian alternatives and standalone reanalysis data in predicting threshold exceedances at the majority of observation sites and can even result in improved spatial patterns of pollution than those discernible from the background data. Moreover, our approach captures greater variability and spatial patterns, such as higher concentrations near coastal areas, which are not evident in the reanalysis data alone.
Machine Learning and Artificial Intelligence
Anastasios N. Angelopoulos, John C. Duchi, Tijana Zrnic
Ann. Appl. Stat. 20 (3), 2235-2252, (September 2026) DOI: 10.1214/26-AOAS2215
KEYWORDS: Prediction-powered inference, confidence intervals, machine learning
We present PPI++: a computationally lightweight methodology for estimation and inference based on a small labeled dataset and a typically much larger dataset of machine-learning predictions. The methods automatically adapt to the quality of available predictions, yielding easy-to-compute confidence sets—for parameters of any dimensionality—that always improve on classical intervals using only the labeled data. PPI++ builds on prediction-powered inference (PPI), which targets the same problem setting, improving its computational and statistical efficiency. Real and synthetic experiments demonstrate the benefits of the proposed adaptations.
Zhiyu Xu, Jia Liu, Yixin Wang, Yuqi Gu
Ann. Appl. Stat. 20 (3), 2253-2273, (September 2026) DOI: 10.1214/26-AOAS2236
KEYWORDS: Large language models, chain of thought, item response theory, Identifiability, stochastic approximation
The proliferation of large language models (LLMs) necessitates valid evaluation methods to provide guidance for both downstream applications and actionable future improvements. The item response theory (IRT) model with computerized adaptive testing has recently emerged as a promising framework for evaluating LLMs via their response accuracy. Beyond simple response accuracy, LLMs’ chain of thought (CoT) lengths serve as a vital indicator of their reasoning ability. To leverage the CoT length information to assist LLM evaluation, we propose the Latency-Response Theory (LaRT) model, which jointly models both the response accuracy and CoT length by introducing a key correlation parameter between the latent ability and the latent speed. We derive an efficient stochastic approximation expectation-maximization algorithm for parameter estimation. We establish rigorous identifiability results for the latent ability and latent speed parameters to ensure the statistical validity of their estimation. Through both theoretical asymptotic analyses and simulation studies, we demonstrate LaRT’s advantages over IRT in terms of superior estimation accuracy and shorter confidence intervals for latent trait estimation. To evaluate LaRT in real data, we collect responses from diverse LLMs on popular benchmark datasets. We find that LaRT yields different LLM rankings than IRT and outperforms IRT across multiple key evaluation metrics including predictive power, item efficiency, ranking validity, and LLM evaluation efficiency. Code and data used for the analyses are provided as Supplementary Material.
Medical and Health Sciences
Qiquan Wang, Anna Song, Antoniana Batsivari, Dominique Bonnet, Anthea Monod
Ann. Appl. Stat. 20 (3), 2274-2298, (September 2026) DOI: 10.1214/26-AOAS2175
KEYWORDS: Confocal imaging, classification, Gaussian mixture models, leukemia, Persistent homology, prediction
Acute myeloid leukemia (AML) is a type of blood and bone marrow cancer characterized by the proliferation of abnormal clonal hematopoietic cells in the bone marrow leading to bone marrow failure. Over the course of the disease, angiogenic factors released by leukemic cells drastically alter the bone marrow vascular niches resulting in observable structural abnormalities. We use a technique from topological data analysis—persistent homology—to quantify the images and infer on the disease through the imaged morphological features on a proprietary high-resolution image dataset. We find that persistent homology uncovers succinct dissimilarities between the control, early and late stages of AML development. We then integrate persistent homology into stage-dependent Gaussian mixture models, proposing a new class of models, which are applicable to persistent homology summaries and able to both infer patterns in morphological changes between different stages of progression as well as provide a basis for prediction.
Baichen Yu, Xuetong Li, Jing Zhou, Hansheng Wang
Ann. Appl. Stat. 20 (3), 2299-2320, (September 2026) DOI: 10.1214/26-AOAS2189
KEYWORDS: Breast cancer metastasis diagnosis, whole-slide image analysis, Gaussian mixture model, multiple instance learning, robustness analysis
Breast cancer is the most prevalent cancer in women worldwide. Histopathology image analysis serves as the gold standard for cancer diagnosis. In this regard, whole-slide imaging (WSI), a revolutionary technology in digital pathology, allows for ultrahigh-resolution tissue analysis. Despite its promise, WSI analysis faces significant computational challenges due to its massive data size and tissue heterogeneity. To address this issue, we present a Gaussian mixture based multiple instance learning (MIL) framework for WSI analysis with partially subsampled instances. Our approach models a WSI as a bag of instances (i.e., randomly cropped subimages), leveraging a bag-based maximum likelihood estimator (BMLE) to predict metastases. Furthermore, we introduce a subsampling-based maximum likelihood estimator (SMLE) to refine predictions by selectively labeling a subset of instances. Extensive evaluations of the breast carcinoma metastasis prediction demonstrate that BMLE surpasses state-of-the-art methods, while the SMLE further improves the prediction accuracy at both bag and instance levels. We find that our method is fairly robust against various plausible model misspecifications. Theoretical analyses and simulation studies validate the performance and robustness of our methods.
Jacopo Gardella, Raffaele Argiento, Alessandro Casa, Alessia Pini
Ann. Appl. Stat. 20 (3), 2321-2340, (September 2026) DOI: 10.1214/26-AOAS2192
KEYWORDS: Bayesian modelling, flexible smoothing, grouping structure, warping functions constraints, registration
In many real-world applications, functional data exhibit considerable variability in both amplitude and phase. This is especially true in biomechanical data, such as the knee flexion angle dataset motivating our work where timing differences across curves can obscure meaningful comparisons. The curves of this study also exhibit substantial variability from one another. These pronounced differences make the dataset particularly challenging to align properly without distorting or losing some of the individual curve characteristics. Our alignment model addresses these challenges by eliminating phase discrepancies while preserving the individual characteristics of each curve and avoiding distortion, thanks to its flexible smoothing component. Additionally, the model accommodates group structures through a dedicated parameter. By leveraging a Bayesian approach, the new prior on the warping parameters ensures that the resulting warping functions automatically satisfy all necessary validity conditions. We applied our model to the knee flexion dataset, demonstrating excellent performance in both smoothing and alignment, particularly in the presence of high intercurve variability and complex group structures.
Ariane Ducellier, Alexander Hsu, Parkes Kendrick, Bill Gustafson, Laura Dwyer-Lindgren, Christopher Murray, Peng Zheng, Aleksandr Aravkin
Ann. Appl. Stat. 20 (3), 2341-2363, (September 2026) DOI: 10.1214/26-AOAS2197
KEYWORDS: Raking, global health estimates, variance estimation, benchmarking
Raking adjusts observations in contingency tables to match known marginals, reconciling estimates across models with different granularities in survey inference and global health. A key use case is the Global Burden of Diseases, Injuries, and Risk Factors Study (GBD), which estimates mortality rates and other health metrics at multiple scales and can produce inconsistent tables (e.g., subregion totals that do not sum to regional totals). GBD outputs include uncertainty (standard deviations or confidence intervals), but variance estimation for raked values is typically obtained by computationally intensive Monte Carlo methods that pass thousands of samples through the full raking procedure. In this paper, we propose a variance estimation approach that, at the cost of a single solve, achieves nearly the same variance estimates as these Monte Carlo techniques. We also introduce a convex optimization framework that unifies raking extensions, including differential weights, alternative distance measures that enforce bounds, feasibility checks, raking to margins as hard constraints or aggregate observations, and handling missing data. We illustrate the general approach and uncertainty computation on a large-scale 3D mortality estimation use case that reconciles estimates across race, county, and cause.
Can Xu, Zeyu Lu, Yichen Cheng, Qifan Song, Yichuan Zhao, Xinlei Wang
Ann. Appl. Stat. 20 (3), 2364-2385, (September 2026) DOI: 10.1214/26-AOAS2210
KEYWORDS: active set, double penalty, Estimating equation, Lagrange multiplier, pathwise coordinate optimization, Sparse learning
Predicting drug sensitivity and selecting informative biomarkers are fundamental challenges in precision oncology. These tasks are complicated by the ultrahigh-dimensional nature of omics data, where the number of covariates p far exceeds the sample size n. While variable selection methods are widely used, traditional high-dimensional regression techniques often rely on the assumption of Gaussian noise, which is frequently violated in real-world omics datasets. To overcome these limitations, we propose a Bayesian empirical likelihood approach (BEL_HD) for variable selection and prediction in a semiparametric framework. BEL_HD avoids strict distributional assumptions by utilizing estimating equations and imposes joint regularization on both regression coefficients and the Lagrange multipliers to promote sparsity. We also introduce an efficient active set MCMC algorithm that enables scalable inference in ultrahigh-dimensional spaces. Through simulations and real-data applications, including a case study on leukemia drug sensitivity, BEL_HD demonstrates superior predictive performance and selects biologically interpretable features. Our method combines the flexibility of empirical likelihood with the inferential benefits of Bayesian modeling, offering a robust and extensible tool for high-dimensional biomedical research.
Shanpeng Li, Daniel S. Nuyujukian, Robyn L. McClelland, Peter D. Reaven, Jin Zhou, Hua Zhou, Gang Li
Ann. Appl. Stat. 20 (3), 2386-2411, (September 2026) DOI: 10.1214/26-AOAS2218
KEYWORDS: Competing risks data, longitudinal data, massive sample size, nonignorable missing data, WS variability
Motivated by a growing body of research emphasizing the importance of modeling within-subject (WS) variability in longitudinal biomarkers and its association with health outcomes, this paper proposes a semiparametric joint model for both the mean and WS variability of a longitudinal biomarker, jointly with competing-risk time-to-event outcomes. We derive an expectation-maximization algorithm for parameter estimation and a profile-likelihood method for standard error estimation and inference, which allows time-dependent covariates and general forms of the latent association structure. Furthermore, we optimize the implementation of our joint model when the survival submodel includes only time-independent baseline covariates and shared random effects, allowing it to scale effectively to biobank-scale data involving tens of thousands of subjects. Our method demonstrates satisfactory performance in simulations, whereas classical joint models that assume homogeneous WS variability may suffer from substantial estimation bias, invalid inference, and inferior prediction when confronted with heterogeneous WS variability. We illustrate the utility of our method using the Multi-Ethnic Study of Atherosclerosis (MESA) cohort. Our analysis demonstrates that associations between WS blood pressure variability and cardiovascular outcomes, previously observed in clinical trials involving relatively homogeneous populations, extend to a more ethnically diverse and generally healthier cohort, and that explicitly modeling heterogeneous WS variability substantially enhances risk discrimination. A user-friendly R package, JMH, has been developed for the proposed shared random effects model with efficient implementation and is publicly available on the Comprehensive R Archive Network https://CRAN.R-project.org/package=JMH.
Camilo A. Cárdenas-Hurtado, Sze Ming Lee, Yunxiao Chen, Irini Moustaki
Ann. Appl. Stat. 20 (3), 2412-2432, (September 2026) DOI: 10.1214/26-AOAS2222
KEYWORDS: Semiparametric model, latent variable model, monotone function, nonparametric item response theory, exploratory data analysis
Cognitive diagnosis models (CDMs) are restricted latent class models widely used to measure attributes of interest in diagnostic assessments across education, psychology, biomedical sciences, and related fields. Partial-mastery CDMs (PM-CDMs) are an important extension of CDMs. They model individuals’ status for each attribute as continuous to measure partial mastery levels, thereby relaxing the restrictive discrete-attribute assumption of classical CDMs. As a result, PM-CDMs often yield better fits to real-world data and more refined measurements of the substantive attributes of interest. However, these models inherit strong parametric assumptions from traditional CDMs about item response functions and thus still face a significant risk of model misspecification. This paper proposes a generalized additive PM-CDM (GaPM-CDM) that substantially relaxes the parametric assumptions of PM-CDMs. This proposal leverages model parsimony and interpretability by modeling each item response function as a mixture of nonparametric monotone functions of attributes. A method for estimating GaPM-CDM is developed that combines the marginal maximum likelihood estimator with a sieve approximation of the nonparametric functions. The new model is applicable in both confirmatory and exploratory settings, depending on whether prior knowledge of the relationship between observed variables and attributes is available. The proposed method is evaluated and compared with PM-CDMs through extensive simulation studies and further applied to two measurement problems from educational testing and healthcare research, respectively.
Can Xie, Ruosha Li, Jeffrey Morris, Nicholas Short, Hagop Kantarjian, Xuelin Huang
Ann. Appl. Stat. 20 (3), 2433-2452, (September 2026) DOI: 10.1214/26-AOAS2232
KEYWORDS: Landmarking, cure, nonproportional hazards, functional principal component analysis, longitudinal data, Survival analysis
Important biomarkers are routinely measured during patients’ follow-up visits to monitor their response to treatment and track disease progression. Accurate, individualized dynamic prediction based on data available at these visits is essential for making informed decisions regarding the next stage of treatment. We propose a versatile landmark model designed to address the complexities of real-world clinical scenarios, including (1) complex longitudinal biomarker trajectories not following a prespecified parametric function over time, (2) the potential for a fraction of patients to be cured, and (3) the violation of the proportional hazards assumption in the relationship between predictors and cure or survival outcomes. The proposed landmark model leverages functional principal component analysis (FPCA) to extract predictive features from longitudinal biomarker data, using the resulting functional principal component scores as predictors. Simulation studies demonstrate that the proposed model outperforms commonly used prediction models in terms of AUC values and Brier scores across a range of complex scenarios. We illustrate the practical utility of our model through its application to a study of chronic myeloid leukemia, showcasing its effectiveness in dynamically predicting individual conditional cure and survival probabilities.
Medical and Health Sciences: Deep Learning
Stephen Salerno, Zhilin Zhang, Yi Li
Ann. Appl. Stat. 20 (3), 2453-2473, (September 2026) DOI: 10.1214/26-AOAS2190
KEYWORDS: deep learning, semicompeting risks, EM algorithm, Brier score, lung cancer
Prognostic modeling for lung cancer, a leading cause of mortality, remains a complex task, as it needs to quantify the associations of risk factors and health events spanning a patient’s entire life. One challenge is that an individual’s disease course involves nonterminal (e.g., disease progression) and terminal (e.g., death) events, which form semicompeting relationships. Our motivation comes from the Boston Lung Cancer Study, a large lung cancer survival cohort, which investigates how risk factors influence a patient’s disease trajectory. Following developments in the prediction of time-to-event outcomes with neural networks, deep learning has become a focal area for the development of risk prediction methods in survival analysis. However, limited work has been done to predict multistate or semicompeting risk outcomes, in which a patient may experience adverse events such as disease progression prior to death. We propose a neural expectation-maximization algorithm for semi-competing risks to bridge the gap between classical semicompeting survival models and deep learning. Our algorithm enables estimation of the nonparametric baseline hazards of each state transition, predictor risk functions, and the degree of dependence among different transitions, via a multitask deep neural network with transition-specific subarchitectures. We apply our method to the Boston Lung Cancer Study and investigate the impact of clinical and genetic predictors on disease progression and mortality.
Haifeng Wang, Hao Xu, Jun Wang, Jian Zhou, Ke Deng
Ann. Appl. Stat. 20 (3), 2474-2496, (September 2026) DOI: 10.1214/26-AOAS2211
KEYWORDS: Surgical video analysis, surgical phase recognition, object detection, Hidden Markov model
Recognizing various surgical tools, actions and phases from surgery videos is an important problem in computer vision with exciting clinical applications. Existing deep-learning-based methods for this problem either process each surgical video as a series of independent images without considering their dependence or rely on complicated deep-learning models to count for dependence of video frames. In this study we revealed from exploratory data analysis that surgical videos enjoy relatively simple semantic structure, where the presence of surgical phases and tools can be well modeled by a compact hidden Markov model (HMM). Based on this observation, we propose an HMM-stabilized deep-learning method for tool presence detection. A wide range of experiments confirm that the proposed approaches achieve better performance with lower training and running costs, and support more flexible ways to construct and utilize training data in scenarios where not all surgery videos of interest are extensively labelled. These results suggest that popular deep-learning approaches with overcomplicated model structures may suffer from inefficient utilization of data, and integrating ingredients of deep learning and statistical learning wisely may lead to more powerful algorithms that enjoy competitive performance, transparent interpretation and convenient model training simultaneously.
Medical and Health Sciences: Electronic Health Records
Yang Wang, Qingning Zhou, Tianxi Cai, Xuan Wang
Ann. Appl. Stat. 20 (3), 2497-2517, (September 2026) DOI: 10.1214/26-AOAS2212
KEYWORDS: Double censoring, electronic health record, optimal combination, semisupervised estimation, cross-fitting
Electronic health records (EHRs) provide a rich data source for translational research, particularly for time-to-event outcomes like disease onset in clinical decision support. However, events often occur outside the capturing health system, leading to double censoring (left censoring and right censoring). Precise event times, embedded in unstructured clinical notes, require costly manual chart reviews. Proxies, such as time to the first diagnostic code or the first mention of the disease, are readily available but error-prone, causing biased estimates if used directly. Relying solely on small labeled sets from reviews yields high variability, emphasizing the need for methods leveraging both labeled and large-scale unlabeled data with surrogate outcomes. While semisupervised methods exist for binary and right-censored outcomes, none address double censoring. The paper fills the gap by developing a robust and efficient Semisupervised Estimation of survival ratE with Doubly-censored Survival data (SEEDS) by leveraging a small set of gold standard labels and a large set of surrogate features. Under mild conditions, SEEDS is consistent and asymptotically normal. Simulations demonstrate SEEDS outperforms supervised methods. We apply the SEEDS procedure to estimate the sex-specific survival rate of type 2 diabetes (T2D) using EHR data from Mass General Brigham (MGB).
Medical and Health Sciences: Epidemiology
Jeremy Goldwasser, Addison J. Hu, Alyssa Bilinski, Daniel J. McDonald, Ryan J. Tibshirani
Ann. Appl. Stat. 20 (3), 2518-2544, (September 2026) DOI: 10.1214/26-AOAS2216
KEYWORDS: Severity rates, case-fatality rate, Deconvolution, Trend filtering, Covid-19, epidemiology
Several key metrics in public health convey the probability that a primary event will lead to a more serious secondary event in the future. These “severity rates” can change over the course of an epidemic in response to shifting conditions, like new therapeutics, variants, or public health interventions. In practice, time-varying parameters, such as the case-fatality rate, are typically estimated from aggregate count data. Prior work has demonstrated that commonly-used ratio-based estimators can be highly biased, motivating the development of new methods. In this paper we develop an adaptive deconvolution approach based on approximating a Poisson-binomial model for secondary events, and we regularize the maximum likelihood solution in this model with a trend filtering penalty to produce smooth but locally adaptive estimates of severity rates over time. This enables us to compute severity rates both retrospectively and in real time. Experiments based on COVID-19 death and hospitalization data show that our deconvolution estimator is generally more accurate than the standard ratio-based methods and displays reasonable robustness to model misspecification.
Haoming Shi, Shan Yu, Eric C. Chi
Ann. Appl. Stat. 20 (3), 2545-2564, (September 2026) DOI: 10.1214/26-AOAS2231
KEYWORDS: Epidemic model, spatiotemporal data, robust estimation, outlier
In epidemic modeling, outliers can distort parameter estimation and ultimately lead to misguided public health decisions. Although there are existing robust methods that can mitigate this distortion, the ability to simultaneously detect outliers is equally vital for identifying potential disease hotspots. In this work we introduce a robust spatiotemporal generalized additive model () to address this need. We accomplish this with a mean-shift parameter to quantify and adjust for the effects of outliers and rely on adaptive Lasso regularization to model the sparsity of outlying observations. We use univariate polynomial splines and bivariate penalized splines over triangulations to estimate the functional forms and a data-thinning approach for data-adaptive weight construction. We derive a scalable proximal algorithm to estimate model parameters by minimizing a convex negative log-quasi-likelihood function. Our algorithm uses adaptive step-sizes to ensure global convergence of the resulting iterate sequence. We establish error bounds and selection consistency for the estimated parameters and demonstrate our model’s effectiveness through numerical studies under various outlier scenarios. We demonstrate the practical utility of by analyzing county-level COVID-19 infection data in the United States, highlighting its potential to inform public health decision-making.
Medical and Health Sciences: Mobile Health
Beniamino Hadj-Amar, Vaishnav Krishnan, Marina Vannucci
Ann. Appl. Stat. 20 (3), 2565-2585, (September 2026) DOI: 10.1214/26-AOAS2235
KEYWORDS: Frequency selection, Fourier, multivariate data, rest-activity rhythms, spike-and-slab, stochastic search MCMC
This paper introduces a Bayesian spike-and-slab framework for spectral analysis of time series data. The proposed method combines frequency selection and dimensionality reduction with a refined grid of candidate frequencies, enabling high-resolution recovery of oscillatory components while promoting sparsity through a structured spike-and-slab prior. A stochastic search algorithm efficiently explores the posterior space, yielding posterior inclusion probabilities that quantify the relevance of each frequency. We extend the framework to multivariate signals via a hierarchical prior on frequency inclusion patterns, allowing the model to capture both shared and component-specific rhythms across multiple time series. Extensive simulation studies demonstrate the method’s robustness and superior performance in frequency estimation and spectral power reconstruction compared to existing approaches. Applied to actigraphy data from individuals with partial-onset seizures, the univariate model identifies clinically relevant circadian and ultradian rhythms. In a second application, for the joint analysis of physical activity and skin temperature from a healthy individual, the multivariate model reveals partially overlapping rhythmic components consistent with known physiological coupling. This work establishes a powerful and interpretable approach to spectral analysis, with broad applicability to wearable data, chronobiology, and personalized health monitoring.
Jian Kang, Thomas Nichols, Lexin Li, Martin A. Lindquist, Hongtu Zhu
Ann. Appl. Stat. 20 (3), 2586-2612, (September 2026) DOI: 10.1214/26-AOAS2204
KEYWORDS: challenges, neuroimaging, statistics
Neuroimaging has profoundly enhanced our understanding of the human brain by characterizing its structure, function, and connectivity through modalities like MRI, fMRI, EEG, and PET. These technologies have enabled major breakthroughs across the lifespan, from early brain development to neurodegenerative and neuropsychiatric disorders. Despite these advances, the brain is a complex, multiscale system, and neuroimaging measurements are correspondingly high-dimensional. This creates major statistical challenges, including measurement noise, motion-related artifacts, substantial intersubject and site/scanner variability, and the sheer scale of modern studies. This paper explores statistical opportunities and challenges in neuroimaging across four key areas: (i) brain development from birth to age 20, (ii) the adult and aging brain, (iii) neurodegeneration and neuropsychiatric disorders, and (iv) brain encoding and decoding. After a quick tutorial on major imaging technologies, we review cutting-edge studies, underscore data and modeling challenges, and highlight research opportunities for statisticians. We conclude by emphasizing that close collaboration among statisticians, neuroscientists, and clinicians is essential for translating neuroimaging advances into improved diagnostics, deeper mechanistic insight, and more personalized treatments.
Hongyang Lu, Yunpeng Zhao, Dong Song, Haonan Wang
Ann. Appl. Stat. 20 (3), 2613-2638, (September 2026) DOI: 10.1214/26-AOAS2209
KEYWORDS: neural spike trains, Point processes, generalized linear models, expectation-maximization algorithm
This paper considers the problem of understanding how neurons interact in the presence of external stimuli. Many existing methods for analyzing neuron spiking dynamics overlook a critical feature of spike trains: their variability across trials under identical conditions. To address this, we adopt the point process generalized linear model and propose to characterize between-trial variation in functional connections between neurons through functional random effects. An expectation-maximization algorithm with expectation propagation (EMEP) is proposed for estimating the neural interaction kernel functions. Our proposed model, along with the estimation procedure, is evaluated through simulation studies. The new method significantly outperforms a GLM-based approach that omits random effects, which aggregates data from all trials, treating the kernel coefficients as fixed parameters by providing more precise estimates. We further apply the proposed method to a case study on neural connectivity in a visual cortex dataset.
Physical Science and Engineering
Rashmi Ranjan Bhuyan, Wreetabrata Kar, Gourab Mukherjee
Ann. Appl. Stat. 20 (3), 2639-2660, (September 2026) DOI: 10.1214/26-AOAS2184
KEYWORDS: Digital communications, joint modeling, logistic regression, Polya gamma, generalized clusterwise linear regression, path algorithm
We develop a novel statistical procedure for analyzing the effects of various messaging components in digital communications. A key feature of the proposed method is that it segments consumers based on their historical characteristics and provides fine-grained estimates of the heterogeneous effects that messaging components have on the different user segments. The proposed method CURM, provides clusterwise analysis of users’ responses to digital messages. While CURM uses historical characteristics to segment users via generalized clusterwise linear regression, and through dynamic indices, it incorporates (a) short-term changes in user engagement, and (b) the impact of recent messages on users’ responses.
We conduct Bayesian estimation of the model parameters by using a novel Gibbs algorithm, which is highly scalable due to the usage of Polya gamma distributions based data-augmentation strategy in handling Binomial likelihoods of users’ responses to messages. Finally, through a path-algorithm (PA), we provide an integrated framework for providing fine-grained analysis of the messaging component effects at various levels of heterogeneity. The entire methodology is implemented by the associated R package . We establish large-sample properties on the operational characteristics of the developed algorithm and also describe its computational complexity.
We apply PACURM on consumer responses to digital coupons from the apparel industry. We demonstrate the importance of user segmentation, the frequency of communication messages, and the utility of messaging components for effective digital couponing.
Social and Political Sciences
Lei Shi, Jie Hu, Huaiyu Tan, Libin Jin, Wei Zhong, Chen Shen
Ann. Appl. Stat. 20 (3), 2661-2683, (September 2026) DOI: 10.1214/26-AOAS2206
KEYWORDS: Signal parameters, signal lasso, Linear regression, additive penalty, nonconvex penalty, network reconstruction
Inferring network structures remains an interesting question due to its importance in understanding and controlling the collective dynamics of complex systems. Most real networks exhibit sparsely connected properties, and the connection parameter is a signal connected (0 or 1). Existing shrinking methods such as lasso-type estimation cannot suitably reveal such property. An alternative method, called signal lasso (or its updating version: adaptive signal lasso), was recently proposed to solve the network reconstruction, where the signal parameter can be shrunk to either 0 or 1 in two different directions, and they found that signal lasso outperformed the lasso-type method. The signal lasso or adaptive signal lasso employed the additive penalty of signal and nonsignal terms, which is a convex function and easily to complementation in computation. However, their methods need tuning the one or two parameters to find an optimal solution, which is time cost for large size network. In this paper, we propose a new signal lasso method based on two penalty functions to estimate the signal parameter and uncover the network topology in a complex network with a small number of observations. The penalty functions we introduced are nonconvex functions, thus coordinate descent algorithms are suggested. We find in this method that the tuning parameter can be set to a large enough value such that the signal parameter can be completely shrunk either 0 or 1. The extensive simulations are conducted in linear regression models with different assumptions, the evolutionary-game-based dynamic model, and Kuramoto model of the synchronization problem. The advantages and disadvantages of each method are fully discussed under various conditions. Finally, a real example comes from a behavioral experiment is detailed analyzed. Our results show that signal lasso with nonconvex penalties is effective and fast in estimating signal parameters in the linear regression model.
Ann. Appl. Stat. 20 (3), 2684-2701, (September 2026) DOI: 10.1214/26-AOAS2234
KEYWORDS: Misinformation detection, maximum entropy bootstrap, extreme learning machines, sentiment dynamics
This paper proposes a two-tier hybrid framework for the early detection of fake news in fast-moving, time-dependent online environments. At the first level, lexicon-based sentiment analysis is used to monitor the emotional and sentiment structure of news articles, capturing within-text trajectories and identifying patterns of emotional dissonance that distinguish manually fabricated from genuine reports. At the second level, the maximum entropy bootstrap (MEB) is combined with an ensemble of forecasting models—exponential smoothing (ETS), autoregressive integrated moving average (ARIMA), and extreme learning machines (ELM)—to model keyword-based time-series data from Google Trends and to construct one-sided predictive intervals for public-attention dynamics. Anomalies in keyword attention are then identified via bootstrap-based confidence bounds, without relying on Gaussian error assumptions, and interpreted as potential early-warning signals of misinformation-related shocks. The framework is evaluated on a curated corpus of 25 news articles (20 authentic and five manually fabricated), on synthetic stress tests with engineered attention spikes, and on independently documented misinformation episodes involving keywords such as “Obama birth certificate,” “covid 5g,” and “pizzagate.” Empirical results show that the hybrid system reliably separates fake from genuine texts in the controlled corpus and consistently flags crisis-like surges in keyword attention, providing an interpretable, statistically grounded decision-support tool for early-stage misinformation monitoring that integrates semantic analysis with robust time-series modeling.