Inferring Spatiotemporal Structure through Self-Supervised Learning: From Predictive Coding to Neuro Foundation Models

dc.contributor.advisorRao, Rajesh P. N.
dc.contributor.authorJiang, Linxing
dc.date.accessioned2026-08-11T19:26:52Z
dc.date.issued2026-08-11
dc.date.submitted2026
dc.descriptionThesis (Ph.D.)--University of Washington, 2026
dc.description.abstractA hallmark of intelligence is the ability to infer spatiotemporal structure from streams of external input. In this thesis, I investigate prediction-based self-supervised learning as a principle for learning such structure in both biological and artificial systems. I first study sparse coding variational autoencoders as a bridge between classical predictive coding models and modern self-supervised generative models, showing that weight normalization constraints substantially improve decoder filter quality while preserving sparse, V1-like representations. I then introduce dynamic predictive coding (DPC), a hierarchical model of spatiotemporal prediction and sequence learning in the neocortex. The model assumes that higher cortical levels modulate the temporal dynamics of lower levels, correcting their predictions of dynamics using prediction errors. When trained by minimizing spatiotemporal prediction errors, DPC accounts for a wide range of neuronal and cognitive phenomena, from V1-like space-time receptive fields and cortical-like temporal response hierarchies to postdictive motion perception, cue-triggered sequence recall, and higher-order temporal abstractions. Finally, inspired by the recent success of self-supervised pretraining in large language models, I investigate whether similar approaches can learn generalizable structure from large-scale neural recordings. Using neural data transformers trained on multi-session mouse electrophysiology datasets, I show that pretraining improves neural activity prediction but scales modestly, with brain-region mismatch and session-to-session variability strongly limiting transfer. Together, these results suggest that prediction-based self-supervised learning can be a powerful principle for discovering spatiotemporal structure in data. At the same time, scaling this principle to large-scale neural recordings will require methods that account for the heterogeneity of neural dynamics across brain regions, recording sessions, and individuals.
dc.embargo.termsOpen Access
dc.format.mimetypeapplication/pdf
dc.identifier.otherJiang_washington_0250E_29883.pdf
dc.identifier.urihttps://hdl.handle.net/1773/57252
dc.language.isoen_US
dc.rightsnone
dc.subjectneuro foundation models
dc.subjectpredictive coding
dc.subjectpretraining
dc.subjectsparse coding
dc.subjectComputer science
dc.subjectNeurosciences
dc.subject.otherComputer science and engineering
dc.titleInferring Spatiotemporal Structure through Self-Supervised Learning: From Predictive Coding to Neuro Foundation Models
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Jiang_washington_0250E_29883.pdf
Size:
16.66 MB
Format:
Adobe Portable Document Format