Essays on Policy Learning and Network Econometrics
Date
relationships.isAuthorOf
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
This dissertation consists of three chapters: the first two develop methods for policylearning, and the third studies a framework for incorporating network data as proxy
controls for unobserved heterogeneity in economic models.
In the first chapter, coauthored with Yanqin Fan and Yuan Qi, we propose an
optimal policy that targets the average welfare of the worst-off α-fraction of the post-
treatment outcome distribution. We refer to this policy as the α-Expected Welfare
Maximization (α-EWM) rule, where α∈(0,1] denotes the size of the subpopulation
of interest. The α-EWM rule interpolates between the expected welfare (α= 1) and
the Rawlsian welfare (α →0). For α ∈(0,1), an α-EWM rule can be interpreted
as a distributionally robust EWM rule that allows the target population to have
a different distribution than the study population. Using the dual formulation of
our α-expected welfare function, we propose a debiased estimator for the optimal
policy and establish its asymptotic upper regret bounds. In addition, we develop
asymptotically valid inference for the optimal welfare based on the proposed debi-
ased estimator. We examine the finite sample performance of the debiased estimator
and inference via both real and synthetic data. In the second chapter, coauthored with Zequn Jin, Xi Zheng and Yahong Zhou,we develop a robust and efficient method for policy learning from observational data
in the presence of unobserved confounding, complementing existing instrumental
variable based approaches. We employ the marginal sensitivity model (MSM) to re-
lax the commonly used yet restrictive unconfoundedness assumption by introducing
a sensitivity parameter that captures the extent of selection bias induced by unob-
served confounders. Building on this framework, we consider two distributionally
robust welfare criteria, defined as the worst-case welfare and policy improvement
functions, evaluated over an uncertainty set of counterfactual distributions charac-
terized by the MSM. Closed-form expressions for both welfare criteria are derived.
Leveraging these identification results, we construct doubly robust scores and es-
timate the robust policies by maximizing the proposed criteria. Our approach ac-
commodates flexible machine learning methods for estimating nuisance components,
even when these converge at a moderately slow rate. We establish asymptotic re-
gret bounds for the resulting policies, providing a robust guarantee against the most
adversarial confounding scenario. The proposed method is evaluated through ex-
tensive simulation studies and empirical applications to the JTPA study and Head
Start program. In the third chapter, I study a nonparametric model where a latent variable cre-ates endogeneity by affecting both network formation and an outcome of interest. I
generalize the existing network control function approach to nonparametric outcome
models, using individuals’ link functions to account for the unobserved heterogene-
ity. My identification is a form of matching on unobservables: I conceptually match
individuals based on their latent link functions. To implement this strategy, I first
estimate the distances or dissimilarities between the latent link functions using net-
work data. Second, I apply a functional kernel smoothing over these distances to
estimate the structural parameter. My asymptotic analysis reveals a fundamental
trade-off: the robustness gained from this approach comes at the unavoidable cost
of a slow convergence rate, driven by the difficulty of matching on latent objects. I
characterize this statistical cost by deriving a minimax lower bound.
Description
Thesis (Ph.D.)--University of Washington, 2026
