Essays on Causal Inference and Policy Learning
| dc.contributor.advisor | Fan, Yanqin | |
| dc.contributor.author | Qi, Yuan | |
| dc.date.accessioned | 2026-08-11T19:27:39Z | |
| dc.date.issued | 2026-08-11 | |
| dc.date.submitted | 2026 | |
| dc.description | Thesis (Ph.D.)--University of Washington, 2026 | |
| dc.description.abstract | This dissertation consists of three chapters. The first two develop new methods for causal inference, focusing on settings with limited identification or complex data structures, while the third develops a framework for policy learning that incorporates distributional and welfare considerations. In the first chapter, coauthored with Yanqin Fan, Carlos A. Manzanares, and Hyeonseok Park, we develop a sensitivity analysis of the surrogacy assumption for the surrogate index approach in Athey et al. [2025b]. We introduce "Weighted Surrogate Indices (WSIs)," the analog of the surrogate index under the surrogacy assumption. We show that under comparability, the average treatment effect (ATE) on WSI identifies the ATE on the long-term outcome when a copula of the treatment and the long-term outcome conditional on baseline covariates and surrogates is known. When the copula is unknown, we establish the identified set of the ATE on the long-term outcome. Furthermore, we construct debiased estimators of the ATE for any given copula and develop asymptotically valid inference in both point-identified and partially identified cases. Using data from a poverty alleviation program in Pakistan, we demonstrate the importance of sensitivity checks as well as the usefulness of our approach. In the second chapter, coauthored with Yiqi Liu, we discuss estimation and inference of conditional treatment effects in regression discontinuity (RD) designs with multiple scores. In addition to local linear regressions and the minimax-optimal estimator proposed by Imbens and Wager [2019], we argue that two variants of random forests, honest regression forests and local linear forests, should be added to the toolkit of applied researchers working with multivariate RD designs; their validity follows from results in Wager and Athey [2018] and Friedberg et al. [2020]. We design a systematic Monte Carlo study with data generating processes built both from functional forms that we specify and from Wasserstein Generative Adversarial Networks that closely mimic the observed data. We find no single estimator dominates across all specifications: (i) local linear regressions perform well in univariate settings, but the common practice of reducing multivariate scores to a univariate one can incur undercoverage, possibly due to vanishing density at the transformed cutoff; (ii) good performance of the minimax-optimal estimator depends on accurate estimation of a nuisance parameter and its current implementation only accepts up to two scores; (iii) forest-based estimators are not designed for estimation at boundary points and are susceptible to finite-sample bias, but their flexibility in modeling multivariate scores opens the door to a wide range of empirical applications, as illustrated by an empirical study of COVID-19 hospital funding with three eligibility criteria. In the third chapter, coauthored with Yanqin Fan and Gaoqian Xu, we propose an optimal policy that targets the average welfare of the worst-off $\alpha$-fraction of the post-treatment outcome distribution. We refer to this policy as the $\alpha$-Expected Welfare Maximization ($\alpha$-EWM) rule, where $\alpha \in (0,1]$ denotes the size of the subpopulation of interest. The $\alpha$-EWM rule interpolates between the expected welfare ($\alpha=1$) and the Rawlsian welfare ($\alpha\rightarrow 0$). For $\alpha\in (0,1)$, an $\alpha$-EWM rule can be interpreted as a distributionally robust EWM rule that allows the target population to have a different distribution than the study population. Using the dual formulation of our $\alpha$-expected welfare function, we propose a debiased estimator for the optimal policy and establish its asymptotic upper regret bounds. In addition, we develop asymptotically valid inference for the optimal welfare based on the proposed debiased estimator. We examine the finite sample performance of the debiased estimator and inference via both real and synthetic data. | |
| dc.embargo.terms | Open Access | |
| dc.format.mimetype | application/pdf | |
| dc.identifier.other | Qi_washington_0250E_29479.pdf | |
| dc.identifier.uri | https://hdl.handle.net/1773/57273 | |
| dc.language.iso | en_US | |
| dc.rights | none | |
| dc.subject | Causal Inference | |
| dc.subject | Debiased Estimation | |
| dc.subject | Machine Learning | |
| dc.subject | Partial Identification | |
| dc.subject | Policy Learning | |
| dc.subject | Regression Discontinuity | |
| dc.subject | Economics | |
| dc.subject.other | Economics | |
| dc.title | Essays on Causal Inference and Policy Learning | |
| dc.type | Thesis |
Files
Original bundle
1 - 1 of 1
Loading...
- Name:
- Qi_washington_0250E_29479.pdf
- Size:
- 11.03 MB
- Format:
- Adobe Portable Document Format
