Inferring Spatiotemporal Structure through Self-Supervised Learning: From Predictive Coding to Neuro Foundation Models
Date
relationships.isAuthorOf
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
A hallmark of intelligence is the ability to infer spatiotemporal structure from streams of external input. In this thesis, I investigate prediction-based self-supervised learning as a principle for learning such structure in both biological and artificial systems. I first study sparse coding variational autoencoders as a bridge between classical predictive coding models and modern self-supervised generative models, showing that weight normalization constraints substantially improve decoder filter quality while preserving sparse, V1-like representations. I then introduce dynamic predictive coding (DPC), a hierarchical model of spatiotemporal prediction and sequence learning in the neocortex. The model assumes that higher cortical levels modulate the temporal dynamics of lower levels, correcting their predictions of dynamics using prediction errors. When trained by minimizing spatiotemporal prediction errors, DPC accounts for a wide range of neuronal and cognitive phenomena, from V1-like space-time receptive fields and cortical-like temporal response hierarchies to postdictive motion perception, cue-triggered sequence recall, and higher-order temporal abstractions. Finally, inspired by the recent success of self-supervised pretraining in large language models, I investigate whether similar approaches can learn generalizable structure from large-scale neural recordings. Using neural data transformers trained on multi-session mouse electrophysiology datasets, I show that pretraining improves neural activity prediction but scales modestly, with brain-region mismatch and session-to-session variability strongly limiting transfer. Together, these results suggest that prediction-based self-supervised learning can be a powerful principle for discovering spatiotemporal structure in data. At the same time, scaling this principle to large-scale neural recordings will require methods that account for the heterogeneity of neural dynamics across brain regions, recording sessions, and individuals.
Description
Thesis (Ph.D.)--University of Washington, 2026
