Investigating Generalization of Unlike Coordination in Language Models via Filtered-Corpus Training

relationships.isAuthorOf

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

A long-standing debate in theoretical linguistics concerns the nature of coordination. A common view holds that only same-category constituents can be conjoined, which has been challenged by the many grammatical cases of unlike coordination found in natural language. This thesis investigates how language models (LMs) learn and generalize the linguistic phenomenon of coordination by Filtered-Corpus Training, which creates an environment where models have only been exposed to alike coordination during training. We evaluate and compare these models to counterparts trained on an unfiltered corpus. Our results suggest unlike coordination is not a general exception to LMs and can be learned only with indirect information, although direct exposure may still be needed for the more challenging cases. They further indicate that LMs process unlike coordination by treating the conjoined elements as belonging to similar structural categories or through a mechanism akin to deletion, both of which appear learnable from exposure to alike coordination alone. This work contributes to the growing understanding of how language models internally represent linguistic structure, while also adding to the broader debate on coordination by showing that how models generalize and process unlike coordination without direct exposure.

Description

Thesis (Master's)--University of Washington, 2026

Citation

DOI

Collections