Breaking the language model monolith

dc.contributor.advisorZettlemoyer, Luke
dc.contributor.advisorSmith, Noah
dc.contributor.authorShi, Weijia
dc.date.accessioned2026-08-11T19:26:50Z
dc.date.issued2026-08-11
dc.date.submitted2026
dc.descriptionThesis (Ph.D.)--University of Washington, 2026
dc.description.abstractLanguage models (LMs) are typically monolithic: a single model storing all knowledge and serving every use case. This design presents significant challenges; they often generate factually incorrect statements, require costly retraining to add or remove information, and face serious privacy and copyright issues. In this talk, I will discuss how to break this monolith by introducing modular architectures and training algorithms that separate capabilities across composable components. I’ll cover two forms of modularity: (1) External modularity, which augments LMs with external tools like retrievers to improve factuality and reasoning; and (2) internal modularity, which builds inherently modular LMs from decentrally trained components to enable flexible composition and an unprecedented level of control.
dc.embargo.termsOpen Access
dc.format.mimetypeapplication/pdf
dc.identifier.otherShi_washington_0250E_29746.pdf
dc.identifier.urihttps://hdl.handle.net/1773/57248
dc.language.isoen_US
dc.rightsCC BY
dc.subjectComputer science
dc.subject.otherComputer science and engineering
dc.titleBreaking the language model monolith
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Shi_washington_0250E_29746.pdf
Size:
5.76 MB
Format:
Adobe Portable Document Format