On the Interaction of Learning Stages: Compression, Representation, and Transfer in Deepß Learning

dc.contributor.advisorFarhadi, Ali
dc.contributor.authorRamanujan, Vivek Kelan
dc.date.accessioned2026-08-11T19:26:53Z
dc.date.issued2026-08-11
dc.date.submitted2026
dc.descriptionThesis (Ph.D.)--University of Washington, 2026
dc.description.abstractMachine learning systems are rarely monolithic. They are assembled from sequences of distinct learning stages, each transforming representations for the next. These stages include initialization and optimization, pre-training and fine-tuning, and compression and generation. In practice, however, each stage is typically designed and optimized in isolation. We argue that this neglect of cross-stage structure is a source of inefficiency and brittleness in modern deep learning: much of what determines a pipeline's behavior lies not within any single stage, but in the interactions between them. When a stage is designed with knowledge of those that follow, information and inductive biases can be carried forward in ways that improve the system as a whole. In this dissertation, we investigate how to make these cross-stage interactions explicit, and how doing so can improve the efficiency, robustness, and scaling of learning systems across both image understanding and image generation. We organize this study along three axes: representation, transfer, and compression. We first study representation, examining how a model's features can be deliberately trained so that future, independently trained models remain compatible with them, easing model updates in large-scale retrieval systems. We then turn to transfer, characterizing how properties of the pre-training stage relate to the robustness of fine-tuned models under distribution shift, and developing neural priming, which adapts vision-language models to new tasks at inference time using their own pre-training data. Finally, in generation, we examine how the rate-distortion trade-off of the tokenization stage affects how generation scales with compute, and introduce causally regularized tokenization, which incorporates inductive biases from the generation stage into the tokenizer to improve compute-optimal scaling. Taken together, these results suggest that deep learning pipelines yield more efficient and robust models when their stages are co-designed as interacting parts rather than optimized as independent components.
dc.embargo.lift2027-08-11T19:26:53Z
dc.embargo.termsDelay release for 1 year -- then make Open Access
dc.format.mimetypeapplication/pdf
dc.identifier.otherRamanujan_washington_0250E_29907.pdf
dc.identifier.urihttps://hdl.handle.net/1773/57253
dc.language.isoen_US
dc.rightsCC BY
dc.subjectArtificial Intelligence
dc.subjectComputer Vision
dc.subjectMachine Learning
dc.subjectComputer science
dc.subject.otherComputer science and engineering
dc.titleOn the Interaction of Learning Stages: Compression, Representation, and Transfer in Deepß Learning
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Ramanujan_washington_0250E_29907.pdf
Size:
57.75 MB
Format:
Adobe Portable Document Format