Fixing the Foundation: Principled Approaches to Data Curation and Scaling

dc.contributor.advisorZettlemoyer, Luke
dc.contributor.authorKudugunta, Sneha
dc.date.accessioned2026-08-11T19:26:50Z
dc.date.issued2026-08-11
dc.date.submitted2026
dc.descriptionThesis (Ph.D.)--University of Washington, 2026
dc.description.abstractTraining models at the scales used today is dependent on (1) curating large-scale datasets and (2) perfecting how we make scaling decisions. The datasets collected for pretraining contain trillions of tokens across different modalities, domains, and languages, often from messy sources such as crawled internet content. This necessitates extensive curation, filtering, and cleaning. Then, practitioners must determine the optimal training setting for their desired compute budget: typically by extrapolating from smaller training runs using scaling laws. Yet these two pillars are often studied separately: modern models are trained on heterogeneous and actively curated datasets like \data, but most scaling-law research assumes a fixed data distribution and does not account for how changing dataset composition affects scaling behavior. In this dissertation, I present principled approaches to these interconnected challenges. First, I discuss MADLAS-400, a manually audited, general-domain 3T token monolingual dataset based on CommonCrawl, spanning 419 languages. By documenting the curation process, I demonstrate how iterative audits surface failure modes that automatic filters miss. Second, I conduct a meta-analysis of more than 50 recent papers on scaling laws in deep learning, and identify key methodological decisions involved in scaling law research. I use this to explain discrepancies in the conclusions that several prior works reach, and empirically explore how significantly conclusions can vary depending on experimental design choices. Third, I present ATLAS, the largest multilingual scaling laws study to date --- 774 training experiments spanning 10M--8B parameters and 400+ languages. I derive a language-agnostic scaling law for how specific languages in MADLAD-400 scale with added compute that improves out-of-sample R^2 by more than 0.3 over prior work, and identify compute regimes in which practitioners should pretrain from scratch versus finetune from multilingual checkpoints. Through these contributions I show that progress in foundation models depends not only on scaling up but on making the ingredients and assumptions of scaling explicit. Together, this work offers tools and empirical guidance for building more inclusive and compute-efficient foundation models.
dc.embargo.termsOpen Access
dc.format.mimetypeapplication/pdf
dc.identifier.otherKudugunta_washington_0250E_29767.pdf
dc.identifier.urihttps://hdl.handle.net/1773/57249
dc.language.isoen_US
dc.rightsCC BY
dc.subjectComputer science
dc.subject.otherComputer science and engineering
dc.titleFixing the Foundation: Principled Approaches to Data Curation and Scaling
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Kudugunta_washington_0250E_29767.pdf
Size:
2.6 MB
Format:
Adobe Portable Document Format