Language Family Data Scheduling for Low-Resource Speech Translation
Date
relationships.isAuthorOf
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
Low-resource languages often underperform on natural language processing tasks when compared to high-resource languages due to their limited amount of data. In particular, supervised speech translation tasks require parallel language data in the form of source language audio paired with target translation language text, which is very difficult to obtain for languages which already lack abundant resources. Given a low-resource language, we study the effects of incorporating high-resource language data with the low-resource language data in the training of low-resource speech translation models. We choose the high-resource language by selecting one which is the closest in the phylogenetic language family tree to the low-resource language, in this case focusing on the high-resource language of Welsh and low-resource language of Irish. By exploring the effects of incorporating the Welsh data with the Irish data through various data scheduling configurations, we find that data schedules which utilize both Welsh and Irish data bolster performance on the naturalistic Irish dataset. However, when the Irish dataset is augmented with synthetic data, it is more beneficial to solely use the Irish data.
Description
Thesis (Master's)--University of Washington, 2026
