Handling Missing Values in Mass Spectrometry Proteomics

relationships.isAuthorOf

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

Missing values are a persistent problem in mass spectrometry (MS) proteomics. Missing values refer to proteins that are present in the sample but are not identified and quantified due to various technical reasons. Missing values can make it difficult to compare across MS samples or runs, hindering reproducibility in MS proteomics research. Missing values can also reduce statistical power and wash out signal from low-abundance but biologically meaningful proteins. Here we present new strategies for handling missing values in two major MS acquisition strategies: data-dependent acquisition, or DDA, and data-independent acquisition, or DIA. These include a relatively large-scale deep learning-based method for imputing, or estimating, missing values in quants matrices derived from MS proteomic experiments. This deep learning-based method is theoretically applicable to any MS acquisition strategy. We also describe a novel strategy for handling missing values in DIA based on imputing peptide retention times rather than protein quantitations. We show that both of these imputation methods are capable of generating novel and potentially important biological insights. Finally, we present new conceptual ways of thinking about missing values in MS proteomics. We compare and contrast missing values in each of the major MS acquisition strategies and offer parallels and lessons from the related field of single-cell transcriptomics.

Description

Thesis (Ph.D.)--University of Washington, 2026

Citation

DOI

Collections