Language Family Data Scheduling for Low-Resource Speech Translation

relationships.isAuthorOf

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

Low-resource languages often underperform on natural language processing tasks when compared to high-resource languages due to their limited amount of data. In particular, supervised speech translation tasks require parallel language data in the form of source language audio paired with target translation language text, which is very difficult to obtain for languages which already lack abundant resources. Given a low-resource language, we study the effects of incorporating high-resource language data with the low-resource language data in the training of low-resource speech translation models. We choose the high-resource language by selecting one which is the closest in the phylogenetic language family tree to the low-resource language, in this case focusing on the high-resource language of Welsh and low-resource language of Irish. By exploring the effects of incorporating the Welsh data with the Irish data through various data scheduling configurations, we find that data schedules which utilize both Welsh and Irish data bolster performance on the naturalistic Irish dataset. However, when the Irish dataset is augmented with synthetic data, it is more beneficial to solely use the Irish data.

Description

Thesis (Master's)--University of Washington, 2026

Citation

DOI

Collections