Across the Reinforcement Learning and Large Language Model Spectrum

dc.contributor.advisorDu, Simon
dc.contributor.authorZhou, Runlong
dc.date.accessioned2026-08-11T19:26:51Z
dc.date.issued2026-08-11
dc.date.submitted2026
dc.descriptionThesis (Ph.D.)--University of Washington, 2026
dc.description.abstractThis dissertation studies the spectrum between reinforcement learning (RL) and large language models (LLMs). The first part develops rigorous theoretical guarantees for RL on tabular Markov decision processes (MDPs), specifically, variance-dependent regrets for MDPs and latent MDPs, and a fundamental progress on completely horizon-free regrets. The second part studies how language models retrieve knowledge across modes, showing how pretraining data organization can determine whether memorized information transfers across query formats. The third part analyzes reinforcement learning from human feedback (RLHF), including sampler-dependent convergence for online direct preference optimization (DPO) and last-iterate convergence guarantees for Nash learning under non-transitive preferences. The final part turns to RL for LLMs, presenting reflection-assisted online fine-tuning for agentic tasks, and a large-scale distributed RL post-training framework. Together, these results connect RL to LLMs from various perspectives.
dc.embargo.termsOpen Access
dc.format.mimetypeapplication/pdf
dc.identifier.otherZhou_washington_0250E_29851.pdf
dc.identifier.urihttps://hdl.handle.net/1773/57250
dc.language.isoen_US
dc.rightsnone
dc.subjectComputer science
dc.subject.otherComputer science and engineering
dc.titleAcross the Reinforcement Learning and Large Language Model Spectrum
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Zhou_washington_0250E_29851.pdf
Size:
10.95 MB
Format:
Adobe Portable Document Format