On Unifying Generative Methods for Protein Design

relationships.isAuthorOf

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

Biology has for a long time existed as a discipline of discovery.Progress in generative AI has made it possible to turn a functional specification into directly testable biological structures, shifting the paradigm of biology towards a process of engineering. Significant limitations remain, however, both in models themselves and the practice that surrounds them. This thesis presents three contributions towards addressing these limitations. We first develop RFdiffusion3 (RFD3), a generative all-atom diffusion model fordesigning proteins, the broadest range of biomolecular interactions yet addressed by a single method. We train and deploy RFD3 across unified biomolecular tasks including protein--small molecule interactions, protein-protein interactions, symmetric assemblies, protein-nucleic acid binding, as well as the design of nucleic acids themselves. We find the key to successful implementation requires careful choices of architecture and inference context. One application arising from these advancements in capabilities is enzyme design, where the quality and speed of RFD3 enables unprecedented design complexity compared to previous state-of-the-art methods. Like most generative models, RFD3 is one component in a suite of tools, and tailoringits outputs to a specific function in practice requires orchestrating additional tools. This motivates our second contribution, Chaperone: an HPC orchestration engine unifying the range of biomolecular tools behind a common interface. We specifically design the framework as a toolset harness for LLM agents. In Chaperone-1, the harness lowers API token cost per passing design and improves yield at higher reasoning effort; an autonomous campaign then produces experimentally validated de novo binders. Finally, we reflect on the lessons of these two contributions under a unifying principle:the careful management of context to generative models, both LLMs and diffusion models, is crucial for effective training and practical use for protein design. We therefore briefly explore a unified architectural approach that mirrors multi-modal LLM architectures, motivated by the need to significantly scale training, support longer context windows, and to unify cross-domain modalities. Evaluated on prediction tasks, our approach performs competitively against previous efforts in this direction, providing useful groundwork for future research towards unified biological reasoning models.

Description

Thesis (Ph.D.)--University of Washington, 2026

Citation

DOI