Fixing the Foundation: Principled Approaches to Data Curation and Scaling
| dc.contributor.advisor | Zettlemoyer, Luke | |
| dc.contributor.author | Kudugunta, Sneha | |
| dc.date.accessioned | 2026-08-11T19:26:50Z | |
| dc.date.issued | 2026-08-11 | |
| dc.date.submitted | 2026 | |
| dc.description | Thesis (Ph.D.)--University of Washington, 2026 | |
| dc.description.abstract | Training models at the scales used today is dependent on (1) curating large-scale datasets and (2) perfecting how we make scaling decisions. The datasets collected for pretraining contain trillions of tokens across different modalities, domains, and languages, often from messy sources such as crawled internet content. This necessitates extensive curation, filtering, and cleaning. Then, practitioners must determine the optimal training setting for their desired compute budget: typically by extrapolating from smaller training runs using scaling laws. Yet these two pillars are often studied separately: modern models are trained on heterogeneous and actively curated datasets like \data, but most scaling-law research assumes a fixed data distribution and does not account for how changing dataset composition affects scaling behavior. In this dissertation, I present principled approaches to these interconnected challenges. First, I discuss MADLAS-400, a manually audited, general-domain 3T token monolingual dataset based on CommonCrawl, spanning 419 languages. By documenting the curation process, I demonstrate how iterative audits surface failure modes that automatic filters miss. Second, I conduct a meta-analysis of more than 50 recent papers on scaling laws in deep learning, and identify key methodological decisions involved in scaling law research. I use this to explain discrepancies in the conclusions that several prior works reach, and empirically explore how significantly conclusions can vary depending on experimental design choices. Third, I present ATLAS, the largest multilingual scaling laws study to date --- 774 training experiments spanning 10M--8B parameters and 400+ languages. I derive a language-agnostic scaling law for how specific languages in MADLAD-400 scale with added compute that improves out-of-sample R^2 by more than 0.3 over prior work, and identify compute regimes in which practitioners should pretrain from scratch versus finetune from multilingual checkpoints. Through these contributions I show that progress in foundation models depends not only on scaling up but on making the ingredients and assumptions of scaling explicit. Together, this work offers tools and empirical guidance for building more inclusive and compute-efficient foundation models. | |
| dc.embargo.terms | Open Access | |
| dc.format.mimetype | application/pdf | |
| dc.identifier.other | Kudugunta_washington_0250E_29767.pdf | |
| dc.identifier.uri | https://hdl.handle.net/1773/57249 | |
| dc.language.iso | en_US | |
| dc.rights | CC BY | |
| dc.subject | Computer science | |
| dc.subject.other | Computer science and engineering | |
| dc.title | Fixing the Foundation: Principled Approaches to Data Curation and Scaling | |
| dc.type | Thesis |
Files
Original bundle
1 - 1 of 1
Loading...
- Name:
- Kudugunta_washington_0250E_29767.pdf
- Size:
- 2.6 MB
- Format:
- Adobe Portable Document Format
