The Science of Synthetic Data for Language Models: Search, Verification, and Scaling

dc.contributor.advisorChoi, Yejin
dc.contributor.authorJung, Jaehun
dc.date.accessioned2026-08-11T19:26:48Z
dc.date.issued2026-08-11
dc.date.submitted2026
dc.descriptionThesis (Ph.D.)--University of Washington, 2026
dc.description.abstractThe capabilities of large language models are fundamentally tied to the data on which they are trained, and as the demand for high-quality data has outpaced the supply of naturally occurring human text, the field has increasingly relied on synthetic data—text generated by models themselves—as the scalable alternative. This dissertation argues that meaningfully scaling synthetic data requires three things in concert: generating candidate data through principled search, verifying its quality with provable guarantees, and amplifying the latent quality dimensions that gate generalization. I present a research program organized around these three pillars. (1) Search: I develop iterative synthesis methods that decouple what constitutes good data from who generates it, enabling weak models to bootstrap themselves into competitive specialists for paraphrasing (Impossible Distillation), summarization (InfoSumm), and reasoning (Retro-Search). (2) Verification: I develop Cascaded Selective Evaluation, a framework for LLM-based evaluation with provable guarantees of human agreement, calibrated via a cascade of judge models to a user-specified risk tolerance. (3) Scaling: I develop two instances of an identify-measure-amplify recipe—Prismatic Synthesis amplifies gradient-space diversity through G-Vendi, and DeltaPrompts amplifies capability-gap divergence through answer divergence ∆—demonstrating that targeted amplification along a latent quality dimension restores log-linear scaling where naive scaling saturates. I conclude by foregrounding a structural limitation common to these methods, including my own: each requires a researcher to propose the relevant quality dimension. I outline a forward research program in which models themselves propose, validate, and refine the quality axes for their own training data—an autonomous data-centric analog to the AI-Scientist and algorithm-discovery agents now emerging in adjacent fields—and argue that the metrics and verification machinery developed in this dissertation are designed to be the foundation on which such an agent operates.
dc.embargo.lift2031-07-16T19:26:48Z
dc.embargo.termsRestrict to UW for 5 years -- then make Open Access
dc.format.mimetypeapplication/pdf
dc.identifier.otherJung_washington_0250E_29659.pdf
dc.identifier.urihttps://hdl.handle.net/1773/57243
dc.language.isoen_US
dc.rightsCC BY
dc.subjectArtificial intelligence
dc.subject.otherComputer science and engineering
dc.titleThe Science of Synthetic Data for Language Models: Search, Verification, and Scaling
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Jung_washington_0250E_29659.pdf
Size:
31.45 MB
Format:
Adobe Portable Document Format