Across the Reinforcement Learning and Large Language Model Spectrum
Date
relationships.isAuthorOf
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
This dissertation studies the spectrum between reinforcement learning (RL) and large language models (LLMs). The first part develops rigorous theoretical guarantees for RL on tabular Markov decision processes (MDPs), specifically, variance-dependent regrets for MDPs and latent MDPs, and a fundamental progress on completely horizon-free regrets. The second part studies how language models retrieve knowledge across modes, showing how pretraining data organization can determine whether memorized information transfers across query formats. The third part analyzes reinforcement learning from human feedback (RLHF), including sampler-dependent convergence for online direct preference optimization (DPO) and last-iterate convergence guarantees for Nash learning under non-transitive preferences. The final part turns to RL for LLMs, presenting reflection-assisted online fine-tuning for agentic tasks, and a large-scale distributed RL post-training framework. Together, these results connect RL to LLMs from various perspectives.
Description
Thesis (Ph.D.)--University of Washington, 2026
