Across the Reinforcement Learning and Large Language Model Spectrum

relationships.isAuthorOf

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

This dissertation studies the spectrum between reinforcement learning (RL) and large language models (LLMs). The first part develops rigorous theoretical guarantees for RL on tabular Markov decision processes (MDPs), specifically, variance-dependent regrets for MDPs and latent MDPs, and a fundamental progress on completely horizon-free regrets. The second part studies how language models retrieve knowledge across modes, showing how pretraining data organization can determine whether memorized information transfers across query formats. The third part analyzes reinforcement learning from human feedback (RLHF), including sampler-dependent convergence for online direct preference optimization (DPO) and last-iterate convergence guarantees for Nash learning under non-transitive preferences. The final part turns to RL for LLMs, presenting reflection-assisted online fine-tuning for agentic tasks, and a large-scale distributed RL post-training framework. Together, these results connect RL to LLMs from various perspectives.

Description

Thesis (Ph.D.)--University of Washington, 2026

Citation

DOI