Inference after Selection and Prediction

relationships.isAuthorOf

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

Modern data analysis is often structured sequentially; analysts first conduct exploratory analysis or modelling, then aim to conduct statistical inference based on their preliminary findings. This dissertation develops a suite of tools for conducting valid inference after selection and prediction. The first part of the dissertation discusses randomization procedures in which carefully constructed external noise is used to partition data into two or more folds, each of which can be used for a different stage of a data analysis pipeline. We begin by developing the theory behind randomization procedures that decompose each observation in a dataset into independent components. We then discuss randomization for multivariate Gaussian random vectors. We prove that a single multivariate Gaussian observation with unknown covariance cannot be decomposed into independent parts, then propose an algorithm that decomposes a single observation into dependent parts. The final chapter of the first part proposes a new hypothesis testing framework that exploits structure in the joint distribution of randomized folds to conduct hypothesis tests of data-driven hypotheses in challenging data contexts. The final two chapters consider conducting valid predictive inference with complex survey data. We first study variance estimation for model-assisted estimators through the lens of U- and V-statistics. The final chapter proposes a design-based conformal prediction procedure for predictive inference in fine-scale mapping studies in low- and middle-income countries.

Description

Thesis (Ph.D.)--University of Washington, 2026

Citation

DOI

Collections