The Evolutionary Potential of New Protein Domains
| dc.contributor.advisor | Campbell, Melody | |
| dc.contributor.author | Hollis, Jeremy | |
| dc.date.accessioned | 2026-09-16T18:32:21Z | |
| dc.date.issued | 2026-09-16 | |
| dc.date.submitted | 2026 | |
| dc.description | Thesis (Ph.D.)--University of Washington, 2026 | |
| dc.description.abstract | Proteins are the molecular working units that build, control, maintain and/or innovate in the biological realm. Proteins power everything from the most fundamental, shared processes across all of life to the highly specialized, niche interactions that generate specificity and distinction between even the closest living relatives. In keeping with their functional breadth, proteins are equally complex in how they are encoded and assembled, and how they divide labor across their physical surface. To understand how proteins function, we often take a reductionist approach, fractionating them from their entirety into subunits that share physical or functional relationships with other protein subunits into what we term “domains.” Often, these subunit boundary delineations are guided by nature: protein domains are typically shared across homologous proteins and have common sequential and structural ancestral features, which help us both understand their evolutionary relationships and their functional similarities.Work across all scales of biological research has contributed greatly to our understanding of how protein domains contribute to biological complexity. Structural and biochemical work has given us molecular insight into the atomic mechanisms by which protein domains are built and operate, while work at the cellular, organismal, and population levels has helped us understand how these simple building blocks assemble to form complex machinery that fuels life. Although some proteins are comprised of only a single functional domain, many contain multiple domains which collectively allow complex work by a single protein by uniting multiple specific processes to create a multifunctional protein. For example, the combination of a LIM domain and kinase domain may allow phosphorylation to happen specifically at curved cell surfaces, offering a mechanism for cells to sense and respond to exterior confinement via intracellular signaling. This complexity is widespread; 60% of proteins across the eukaryotic proteome contain multiple functional domains. I am fascinated by how these discrete subunits can build, stack, and recombine to generate the beautiful diversity that exists in the natural world. I am interested in both how new domains move and diversify across evolutionary timescales, as well as how domains can retain function despite strong evolutionary pressures to innovate. To tackle these large questions, I have primarily focused my studies on the integrin alphaI domain as a case study. This domain fundamentally changed how integrins interact both with the exterior world and the interior of the cell and shows a remarkable diversity across the chordate tree of life. I have secondarily traced the evolutionary retention of the scm3 domain, a small protein domain responsible for delineating the centromeric region of chromosomes necessary for faithful cell division that has been functionally retained despite near complete sequence turnover in Metazoa. Integrins are heterodimeric cell surface receptor proteins comprised of an alpha and beta subunit which transmit signals from ligand binding across the cytoplasmic membrane to trigger cellular responses via signaling cascades. Integrins are largely structurally homologous across all of Metazoa, but some alpha subunits contain an additional domain called the alphaI or “inserted” domain, a von Willebrand Factor A-family domain which fundamentally changed how integrins interact with their ligands. When I began my thesis work, we knew structurally how these domains in isolation could change structural conformation, as well as separately how they could bind to proteinaceous ligands. We did not know, however, where this domain came from and how it was poised for evolutionary success despite its insertion into an exceptionally risky area of the integrin protein. We also lacked a structural sense of how ligand binding was allosterically communicated through this domain to the rest of the integrin molecule and ultimately into the cell. To understand how this domain mechanistically transmits signal, in my first thesis chapter, I employed single particle cryogenic electron microscopy (cryoEM) to compare two integrins: alphaEbeta7, which contains an I domain, and alpha4beta7 which does not. I generated high-resolution structures of the integrins in both apo and ligand-bound states to directly observe how ligand binding changed upon acquisition of the I domain, as well as how it uses a structurally labile C-terminal helix to link ion coordination to integrin activation. These studies generated the first high-resolution structures of these integrin ectodomains and further revealed how ancestral integrin signaling could broadly be retained despite the steric occlusion of the ancestral ligand binding site by the I domain. To understand the evolutionary origin of the I domain, I combined phylogenetic analyses with structural modeling, ancestral reconstruction, and flow cytometry which allowed me to pinpoint the serendipitous features of its acquisition that permitted this complex ion coordination behavior from its inception. I also found that this domain has a collagen origin, hinting at a possible ancestral function in collagen binding that is still present in extant integrin I domains across chordates. Integrin I domains vary widely in their ligand binding capacities, with some showing extreme specificity for a single cognate ligand and others having exceptional promiscuity. Thus, the I domain is also a great model for understanding how specificity is generated in the biological world. To dissect the biochemical basis of this dichotomy, in my second chapter I focused on how two closely related integrins – the highly specific alphaLbeta2 and the highly promiscuous alphaMbeta2 – distinguish their ligand binding profiles at the amino acid level. I used deep mutational scanning to probe how individual mutations affect ligand affinities at a high-throughput scale, cryoEM to mechanistically dissect division-of-labor across the integrin surface, and phylogenetics to link these dual behaviors to broad evolutionary trajectories. These studies not only answered longstanding questions in the field as to how alphaMbeta2 can be so promiscuous, but how purifying pressures are equally important in both protein-protein interaction and allosteric regulation as well as how avidity can relax constraint on ligand recognition. Finally, in my third chapter I explore how protein domain structure and function can be preserved despite extreme divergence at the sequence level using a different model domain. In this case, I used a combination of remote homology detection approaches to trace the scm3 domain, which chaperones the centromeric histone variant CENPA to chromosomes, across the animal tree. Due to rapid sequence divergence, it was thought this domain was missing in most animals, including key model species. By uncovering this domain across these “missing lineages,” my work supports the model that structure is more conserved that sequence, function can be maintained across long evolutionary timescales despite a near complete lack of sequence homology, thus changing our perspective on protein constraint as it relates to function. Altogether these studies shape our perspective on how evolutionary potential drives protein domain innovation. Tracing domain recombination and retention helps inform the mechanisms by which proteins build complexity and the subsequent limitations of that complexity, or lack thereof. Domain structural studies in the context of near complete protein provide insight into how novelty is united with ancestry to form functional machinery despite the high risks that domain insertions impose. Finally, integrating mutational data informs us where the limits of protein domains lie, and if we look close enough might even give us a clue as to where they’ll go next. | |
| dc.embargo.terms | Open Access | |
| dc.format.mimetype | application/pdf | |
| dc.identifier.other | Hollis_washington_0250E_30159.pdf | |
| dc.identifier.uri | https://hdl.handle.net/1773/57844 | |
| dc.language.iso | en_US | |
| dc.rights | CC BY | |
| dc.subject | adaptation | |
| dc.subject | cryoEM | |
| dc.subject | deep mutational scanning | |
| dc.subject | HJURP | |
| dc.subject | integrin | |
| dc.subject | phylogenetics | |
| dc.subject | Molecular biology | |
| dc.subject | Evolution & development | |
| dc.subject | Immunology | |
| dc.subject.other | Molecular and cellular biology | |
| dc.title | The Evolutionary Potential of New Protein Domains | |
| dc.type | Thesis |
Files
Original bundle
1 - 1 of 1
Loading...
- Name:
- Hollis_washington_0250E_30159.pdf
- Size:
- 56.12 MB
- Format:
- Adobe Portable Document Format
