On Unifying Generative Methods for Protein Design
Date
relationships.isAuthorOf
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
Biology has for a long time existed as a discipline of discovery.Progress in generative AI has made it possible to turn a functional specification into directly testable biological structures,
shifting the paradigm of biology towards a process of engineering.
Significant limitations remain, however, both in models themselves and the practice that surrounds them.
This thesis presents three contributions towards addressing these limitations. We first develop RFdiffusion3 (RFD3), a generative all-atom diffusion model fordesigning proteins, the broadest range of biomolecular interactions yet
addressed by a single method. We train and deploy RFD3 across unified biomolecular tasks including protein--small molecule
interactions, protein-protein interactions, symmetric assemblies,
protein-nucleic acid binding, as well as the design of nucleic acids themselves.
We find the key to successful implementation requires careful choices of architecture and inference context.
One application arising from these advancements in capabilities is enzyme design, where the
quality and speed of RFD3 enables unprecedented design complexity compared to previous state-of-the-art methods. Like most generative models, RFD3 is one component in a suite of tools, and tailoringits outputs to a specific function in practice requires orchestrating additional tools.
This motivates our second contribution, Chaperone: an HPC
orchestration engine unifying the range of biomolecular tools behind a common interface.
We specifically design the framework as a toolset harness for LLM agents. In
Chaperone-1, the harness lowers API token cost per passing
design and improves yield at higher reasoning effort; an autonomous campaign
then produces experimentally validated de novo binders. Finally, we reflect on the lessons of these two contributions under a unifying principle:the careful management of context to generative models, both LLMs and diffusion models, is crucial for effective training and practical use for protein design.
We therefore briefly explore a unified architectural approach that mirrors multi-modal LLM architectures,
motivated by the need to significantly scale training, support longer context windows, and to unify cross-domain modalities.
Evaluated on prediction tasks, our approach performs competitively against previous efforts in this direction, providing
useful groundwork for future research towards unified biological reasoning models.
Description
Thesis (Ph.D.)--University of Washington, 2026
