Toward Generalist Ophthalmic Intelligence: From Scalable Multimodal Representation Learning to a Clinical World Model
Date
relationships.isAuthorOf
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
The eye is among the very few organs that can be imaged non-invasively, in vivo, and repeatedly. Through the retina, a single examination can capture many two- and three-dimensional modalities that together reflect ocular, systemic, and neurological health. The obstacle to using this access computationally is the form of the data rather than its quantity: retinal images arrive from different devices, sites, modalities, dimensionalities, and time points, and models built for one configuration degrade on the next. Prior work has largely responded by narrowing scope to one modality, one task and one time point, which makes each problem tractable and their union intractable. This dissertation argues that the structure required to reconcile heterogeneous retinal imaging can be moved from hand-specified rules into the learned representation; that doing so increases the range of clinical questions a single model can answer without making the architecture grow with the number of modalities; and that the resulting gain concentrates on tasks requiring integration across modality, dimensionality or time. The argument proceeds in three steps. First, the reconciling structure sits outside the model. GrInAdapt performs source-free multi-target domain adaptation for retinal vessel segmentation by grounding multiple views of a subject into a common anchor space, integrating their predictions into region-wise consensus labels, and adapting the source model to the resulting pseudo-supervision. Reconciliation here is explicit and hand-built, since a person states which modality is trustworthy in which anatomical region. The method improves the source model by 4.3 per cent Dice on average, and its limits are informative, since the hand-written rule covers one pair of modalities and operates on predictions rather than on representations. Second, the structure moves partly inside the model. OCTCube-M is a 3D multimodal foundation model for optical coherence tomography that treats the OCT volume as a native three-dimensional object rather than a stack of slices, and admits other retinal modalities through contrastive OCT volume–en face image pre-training. Pre-trained on 26,605 OCT volumes comprising 1.62 million slices, it predicts eight retinal diseases at state-of-the-art accuracy and generalises across cohorts, devices, modalities, and organs, extending to systemic disease prediction and to geographic atrophy prognosis across multi-centre clinical trials. Third, the structure becomes the representation. GigaSight is a clinical world model for ophthalmology that learns a unified representation from approximately 18.9 million images across 13 modalities, organising the encoding by dimensionality rather than by modality, so that thirteen modalities share one frame encoder and only two sequence encoders, corresponding to their two dimensionalities. Across 20 task categories and 101 prediction targets, its margin over prior foundation models is smallest on screening and grows with the integration a task demands; it extends the read-out from single images to longitudinal sequences, forecasting judgements that no single image contains; and it reaches beyond the eye to systemic and neurocognitive signal. Taken together, the three systems form a progression. Structure that a person had to supply is replaced at each step by structure the model learns, and the clinical read-out widens with it, from a single segmentation task, to retinal and systemic disease, to a representation that spans modality, task, and time. A control that keeps the dimensionality-organised sequence encoder but removes the frame encoder reaches macro AUROC 0.742 across twelve systemic conditions, against 0.768 for the full model and 0.599 for a single-image baseline: the frame-level representation accounts for 0.026, and most of the distance from a single-image model is covered without it. Routine retinal imaging can accordingly be read for evidence of health beyond the eye, reaching the body and the brain.
Description
Thesis (Ph.D.)--University of Washington, 2026
