Interpretable machine learning and governing law discovery

relationships.isAuthorOf

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

Modern machine learning has greatly expanded our ability to model complex scientific data, but prediction alone is not sufficient for scientific discovery. In many real-world systems, the goal is not only to approximate observations, but to recover interpretable coordinates, identify governing equations, quantify uncertainty, and make reliable predictions from sparse, noisy, and partially observed measurements. In this thesis, I present my work toward developing machine learning methods for interpretable physical law discovery in realistic data regimes. In the first work, I developed Bayesian SINDy autoencoders for jointly discovering latent coordinates, governing equations, and physical constants from high-dimensional observations. By combining representation learning, sparse regression, and Bayesian uncertainty quantification, this method enables more reliable discovery when data are limited or noisy, including from real video data where the intrinsic physical variables are not directly observed. Next, I introduced SINDy-SHRED and Koopman-SHRED, which extend governing law discovery to high-dimensional spatiotemporal systems measured only through sparse sensors. These models use sensor histories to learn informative latent states, reconstruct full fields, and impose parsimonious nonlinear or linear dynamical structure on the latent space, enabling interpretable long-term prediction for systems such as fluids, sea-surface temperature, and video dynamics. I then studied the reliability of sparse model discovery through ensemble and Bayesian uncertainty estimation, showing how uncertainty estimates can support stable variable selection and distinguish reliable terms from artifacts of noise or limited data. Moving beyond structured-grid assumptions, I proposed mesh-free sparse identification of nonlinear dynamics, using neural networks and automatic differentiation to discover PDEs from scarce, noisy, and irregularly sampled measurements. Finally, I developed UQ-SHRED, a distributional sparse-sensing framework that learns conditional uncertainty over full spatiotemporal fields from limited sensor histories. Collectively, this thesis advances a unified view of scientific machine learning in which models do not merely fit data, but discover interpretable structure, quantify uncertainty, and adapt physical law discovery to the imperfect data conditions encountered in real scientific systems.

Description

Thesis (Ph.D.)--University of Washington, 2026

Citation

DOI