High-dimensional Limit of SGD for Adaptive Stepsize Algorithms and Diagonal Linear Networks
Date
relationships.isAuthorOf
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
Although stochastic gradient methods are widely used in practice, a complete theoretical understanding of their dynamics is still lacking. Classical analyses typically rely on asymptotic regimes with vanishing stepsizes, leading to continuous-time approximations such as deterministic gradient flow or small-noise diffusion models, which often fail to capture the rich behavior observed in practice. This thesis develops a unified framework for describing stochastic gradient descent (SGD) at finite stepsizes through high-dimensional limits, where algorithmic randomness gives rise to deterministic evolution laws for key quantities of interest. We begin by studying adaptive stepsize methods in high-dimensional optimization problems, which we refer to as the high line. In this setting, we show that both the risk and the stepsize dynamics of one-pass SGD admit exact deterministic descriptions via a system of ordinary differential equations. This perspective enables a precise analysis of commonly used adaptive strategies. In particular, we demonstrate that idealized line search procedures can exhibit arbitrarily slow convergence compared to optimally tuned fixed stepsizes, even in simple least squares problems. We further characterize the long-time behavior of adaptive methods such as AdaGrad-Norm, showing that their stepsizes converge to explicit deterministic limits governed by the spectral properties of the data covariance, and uncover phase transitions under power-law eigenvalue distributions. We then turn to diagonal linear networks as a canonical model for understanding neural optimization. In the high-dimensional setting, the trajectory of stochastic gradient descent admits a continuous-time stochastic representation, formalized through a stochastic differential equation, in which the deterministic and stochastic components of the dynamics are explicitly disentangled. This representation induces a closed deterministic evolution for a collection of low-dimensional summary statistics, including measures of risk and curvature. Leveraging this structure, we obtain a sharp description of the global behavior of the dynamics, establishing well-posedness and exponential convergence to zero risk with high probability. Together, these results demonstrate that, in high-dimensional settings, stochastic gradient methods admit precise deterministic descriptions that capture both optimization performance and stepsize behavior. This perspective offers a new lens on algorithmic design and analysis, bridging stochastic optimization, high-dimensional probability, and dynamical systems.
Description
Thesis (Ph.D.)--University of Washington, 2026
