On the Interaction of Learning Stages: Compression, Representation, and Transfer in Deepß Learning
Date
relationships.isAuthorOf
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
Machine learning systems are rarely monolithic. They are assembled from sequences of distinct learning stages, each transforming representations for the next. These stages include initialization and optimization, pre-training and fine-tuning, and compression and generation. In practice, however, each stage is typically designed and optimized in isolation. We argue that this neglect of cross-stage structure is a source of inefficiency and brittleness in modern deep learning: much of what determines a pipeline's behavior lies not within any single stage, but in the interactions between them. When a stage is designed with knowledge of those that follow, information and inductive biases can be carried forward in ways that improve the system as a whole. In this dissertation, we investigate how to make these cross-stage interactions explicit, and how doing so can improve the efficiency, robustness, and scaling of learning systems across both image understanding and image generation. We organize this study along three axes: representation, transfer, and compression. We first study representation, examining how a model's features can be deliberately trained so that future, independently trained models remain compatible with them, easing model updates in large-scale retrieval systems. We then turn to transfer, characterizing how properties of the pre-training stage relate to the robustness of fine-tuned models under distribution shift, and developing neural priming, which adapts vision-language models to new tasks at inference time using their own pre-training data. Finally, in generation, we examine how the rate-distortion trade-off of the tokenization stage affects how generation scales with compute, and introduce causally regularized tokenization, which incorporates inductive biases from the generation stage into the tokenizer to improve compute-optimal scaling. Taken together, these results suggest that deep learning pipelines yield more efficient and robust models when their stages are co-designed as interacting parts rather than optimized as independent components.
Description
Thesis (Ph.D.)--University of Washington, 2026
