Physical and Computational Approaches Towards More Scalable Molecular Systems

relationships.isAuthorOf

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

The ability to understand and engineer biological systems depends on efficiently manipulating and interpreting molecular information. Design-Build-Test-Learn (DBTL) cycles remain limited by bottlenecks in both experimental workflows and computational analysis. This thesis addresses these complementary challenges through both physical laboratory automation and machine learning methods for interpreting nanopore sequencing signals from chemically diverse biopolymers. First, I present work on digital microfluidic automation using the PurpleDrop platform, enabling programmable droplet manipulation for applications including DNA data storage and aptamer discovery. This work demonstrates the potential of integrated hardware-software systems for automated experimentation while highlighting challenges in reliability and system integration. The primary focus of this thesis develops structure-informed machine learning models for nanopore sequencing of expanded molecular alphabets. Xenonucleic acids (XNAs), synthetic nucleotides with modified bases, sugars, or backbones, hold promise for therapeutics, diagnostics, and information storage but lack robust sequencing methods. Existing nanopore approaches require extensive experimental calibration across many sequence contexts. I show that models incorporating molecular structure and chemical features can predict nanopore ionic current signatures for unseen XNA bases, substantially reducing experimental requirements. Structure-informed models outperform sequence-only approaches in few-shot and zero-shot settings and enable practical XNA sequencing with limited calibration data. Finally, I extend this framework to nanopore protein sequencing for detection of post-translational modifications such as phosphorylation. Despite the greater complexity of proteins, similar modeling approaches enable phosphorylation detection, suggesting broader applicability of structure-based methods for complex biopolymer sequencing. Together, these contributions reduce experimental burden through both physical automation and computational prediction, enabling more scalable DBTL cycles for increasingly complex molecular systems.

Description

Thesis (Ph.D.)--University of Washington, 2026

Citation

DOI