Statistical Advancements in Causal Decomposition and Causal Discovery for Real-World Applications

relationships.isAuthorOf

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

Given recent advances in causal inference, what challenges still prevent these methods from being widely or reliably used in real-world applications? This dissertation considers two tasks in causal inference - causal decomposition and causal discovery - and develops statistical methodology in three separate projects to support their use in practice. In the first project, we study Black-white disparities in NIH R01 grant peer review. Prior work has identified the step of selection into panel discussion as an important contributor to Black-white funding disparities at the NIH. In this project, our objective is to use a causal decomposition framework to understand what factors could causally contribute to reducing Black-white disparities in selection into discussion, and to assess how recent procedural changes may affect Black-white disparities in selection into discussion. To achieve these objectives, we extend the causal decomposition framework to accommodate the complexities of the NIH peer review process. For example, we account for interference between applications, since proposals are evaluated relative to others and interventions on one proposal’s score may affect the discussion outcomes of others. We also develop a Bayesian estimation procedure for causal decomposition that incorporates group-level structure. Using our causal decomposition framework, we first explore whether hypothetical interventions eliminating Black-white disparities and/or differences in attributes would reduce Black-white disparities in selection of proposals for panel discussion. The attributes we consider include those associated with applicants (degree, career stage, their institution’s funding bin) and their submitted proposals (application type, amended status, Preliminary Overall Impact Score, and NIH assigned Administering Organization and Integrated Review Group). Under reasonable assumptions, we find that among the attributes considered, Preliminary Overall Impact Score was the only one that, after equalizing, could dissolve Black-white disparities in selection into discussion. In addition, we conduct a thought experiment to explore potential impacts of the NIH's move to binary scoring for the Investigator and Environment criteria in the new Simplified Peer Review Framework on Black-white disparities in selection into discussion. To carry out this thought experiment, we introduce a new estimand within our causal decomposition framework, the Thresholded Disparity, which captures the impact of discretizing evaluation criteria while accounting for the same complexities of the NIH peer review process (e.g., interference), and estimate it using a similar Bayesian procedure that incorporates group-level structure. Analyses from our thought experiment suggest that changing to binary scoring of the Investigator-and-Environment factor may not have any effect on disparities in selection into discussion. Overall, under the assumptions of the causal decomposition framework, these results suggest that efforts aimed at reducing disparities in selection into discussion should focus on identifying upstream, policy-relevant, and practically feasible interventions for eliminating racial disparities in Preliminary Overall Impact Scores. In contrast, the transition to the Simplified Peer Review Framework is unlikely to substantially reduce disparities in selection into discussion. In the second and third projects, we develop statistical inference tools for causal discovery to address the lack of formal inference and sensitivity to assumption violations in existing methods. Existing functional causal discovery approaches use structural asymmetries to identify causal directionality but rely on strong modeling assumptions and provide limited support for uncertainty quantification. To address these limitations, in the second project, we develop a foundational hypothesis testing-based framework that provides formal inference on causal direction and diagnostic information about potential assumption violations for bivariate functional causal discovery. We demonstrate these inferential and diagnostic capabilities through simulations with varying degrees of assumption violation and through case studies of cause-effect datasets. Building on this foundational hypothesis testing-based framework, in the third project, we introduce Causal Discovery via Statistical Power (CDSP), a method that connects causal direction estimation - a.k.a. causal discovery - with statistical power. Within CDSP, we define notions of statistical power and effect size in causal discovery and characterize, through the effect-size asymmetry assumption, when the observational data contain sufficient information to favor one causal direction over the other. We establish theoretical results linking effect-size asymmetry to direction detection, showing that when this assumption holds, the probability of correctly detecting the causal direction (i.e., the power of causal discovery) exceeds that of favoring the reverse direction. Leveraging these results, we develop a procedure for causal direction estimation via effect-size asymmetry and provide statistical inference for the estimated direction through a measure of directional support. Simulations show that CDSP direction estimation is robust to mild and moderate model misspecification. Real data analyses on 100 cause-effect benchmark pairs further demonstrate that CDSP reduces false discovery rates by approximately 18% relative to a commonly used existing method. Together, the contributions of this dissertation support the use of statistically grounded causal inference methodologies in complex real-world settings.

Description

Thesis (Ph.D.)--University of Washington, 2026

Citation

DOI

Collections