Optimization in the Era of Large Language Models: Theory and Applications

dc.contributor.advisorKarlin, Anna R
dc.contributor.authorZhang, Xinzhi
dc.date.accessioned2026-08-11T19:26:34Z
dc.date.issued2026-08-11
dc.date.submitted2026
dc.descriptionThesis (Ph.D.)--University of Washington, 2026
dc.description.abstractOptimization serves as the unifying language for intelligent decision-making, driving both classical algorithmic solvers and modern machine learning systems. This thesis explores the dynamic intersection of large language models (LLMs), practical optimization applications, and foundational optimization theory. The dissertation is organized into two primary parts. In Part I (Chapters 2-5), we focus on the systems and methodologies that enable efficient training, advanced reasoning, and reliable optimization tasks. In Chapte 3, we tackle communication bottlenecks in distributed training by introducing Pseudo-Asynchronous Local SGD (PALSGD). This method elegantly regularizes local workers through pseudo-synchronization, allowing for larger synchronization intervals while maintaining strong empirical performance. Next, Chapter 4 enhances the reasoning capabilities of LLMs with Reinforcement Learning via Self-Play (RLSP), a post-training framework that models reasoning as an active search process, encouraging complex behaviors such as exploration, backtracking, and self-verification. Finally, Chapter 5 addresses the domain-expertise bottleneck in automating natural-language-to-optimization translation. We present OptiMind, a framework that systematically distills structural expertise into LLMs through targeted data curation, class-based error analysis, and test-time scaling, significantly boosting formulation robustness on realistic benchmarks. In Part II (Chapters 6-7), we shift our focus to the rigorous theoretical analysis of optimization algorithms. Chapter 6 investigates the smoothed complexity of the shadow-vertex simplex method. By analyzing the algorithm under random input perturbations, we derive tighter upper bounds and establish the first non-trivial lower bounds within this framework. This analysis helps narrow the persistent gap between the algorithm's renowned practical efficiency and its exponential worst-case guarantees. In Chapter 7, we broaden the scope of spectral analysis by extending Kadison--Singer-type discrepancy results to hyperbolic polynomial settings, providing novel sub-exponential-time algorithms under generalized assumptions. Together, these works bridge the gap between foundational optimization theory and scalable, LLM-driven applications. By advancing both rigorous mathematical guarantees and practical system designs, this thesis demonstrates how to make complex decision-making pipelines more efficient, reliable, and accessible.
dc.embargo.lift2027-08-11T19:26:34Z
dc.embargo.termsRestrict to UW for 1 year -- then make Open Access
dc.format.mimetypeapplication/pdf
dc.identifier.otherZhang_washington_0250E_29225.pdf
dc.identifier.urihttps://hdl.handle.net/1773/57234
dc.language.isoen_US
dc.rightsCC BY-NC-SA
dc.subjectDistributed Training
dc.subjectLarge Language Models
dc.subjectOptimization
dc.subjectReinforcement Learning
dc.subjectSmoothed Analysis
dc.subjectComputer science
dc.subject.otherComputer science and engineering
dc.titleOptimization in the Era of Large Language Models: Theory and Applications
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Zhang_washington_0250E_29225.pdf
Size:
3.83 MB
Format:
Adobe Portable Document Format