Toward Equitable Speech Recognition: Linguistic Structure, Representation, and Fairness in Automatic Speech Recognition

relationships.isAuthorOf

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

Automatic Speech Recognition (ASR) systems have achieved near-human transcription accuracy on high-resource languages, yet their benefits remain unevenly distributed across the world’s languages and speaker communities. This dissertation investigates three inter-locking sources of ASR inequity: data scarcity for endangered and low-resource languages, internal representational biases in multilingual ASR architectures, and the socially-situated harms experienced by speakers of underrepresented language varieties. At the data and adaptation level, the dissertation benchmarks fine-tuned multilingual ASR models, MMS-1B (Pratap et al., 2024) and XLS-R-300m (Babu et al., 2021), on five typologically diverse endangered languages drawn from linguistic fieldwork archives with systematic control of training data duration. Results show that parameter-efficient adapter fine-tuning could outperform full fine-tuning under one hour of data, that both models reach comparable performance at approximately one hour, and that both struggle persistently with complex phonological inventories—such as tone, nasality, and consonant length—in ways that aggregate error metrics obscure. At the model-internal level, a set of experiments probe the decoding mechanisms of Whisper, a large multilingual ASR model, across languages in Latin and diverse scripts. Sub-token probing reveals that higher-resource languages have higher decoder confidence, lower predictive entropy, better-ranked correct tokens, and more diverse alternative candidates. Sub-token usage patterns cluster typologically, not merely by resource tier, showing that linguistic structure shapes internal representations in addition to training scale. A scaled follow-up study across languages in diverse scripts finds that sub-token vocabulary activation is largely independent of pre-training hours, converging instead after approximately two hours of audio, and that rank-frequency distributions can be explained by shared orthography in addition to resource differences. At the social and user level, a mixed-methods survey of speakers across four U.S. dialect communities documents the lived experiences of ASR bias: invisible labor (code-switching, hyper-articulation, emotional management), internalized self-blame despite critical awareness of systemic exclusion, and varying rates of technology abandonment. These findings demonstrate that accuracy-centric fairness metrics miss critical dimensions of harm borne by affected users. Together, the studies in this dissertation argue that equitable speech recognition requires linguistically-grounded design and evaluation at every level of the pipeline—from input data and phonological structure, through model-internal representations, to the social consequences of system implementation.

Description

Thesis (Ph.D.)--University of Washington, 2026

Citation

DOI

Collections