<?xml version="1.0" encoding="UTF-8"?><?xml-stylesheet type="text/xsl" href="static/style.xsl"?><OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd"><responseDate>2026-09-19T04:49:08Z</responseDate><request verb="GetRecord" identifier="oai:digital.lib.washington.edu:1773/53512" metadataPrefix="dim">https://digital.lib.washington.edu/server/oai/request</request><GetRecord><record><header><identifier>oai:digital.lib.washington.edu:1773/53512</identifier><datestamp>2026-06-09T17:13:21Z</datestamp><setSpec>com_1773_4888</setSpec><setSpec>col_1773_4909</setSpec></header><metadata><dim:dim xmlns:dim="http://www.dspace.org/xmlns/dspace/dim" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:doc="http://www.lyncode.com/xoai" xsi:schemaLocation="http://www.dspace.org/xmlns/dspace/dim http://www.dspace.org/schema/dim.xsd">
   <dim:field mdschema="dc" element="contributor" qualifier="advisor">Hajishirzi, Hannaneh</dim:field>
   <dim:field mdschema="dc" element="contributor" qualifier="advisor">Smith, Noah A.</dim:field>
   <dim:field mdschema="dc" element="contributor" qualifier="author">Wang, Yizhong</dim:field>
   <dim:field mdschema="dc" element="date" qualifier="accessioned">2025-08-01T22:19:48Z</dim:field>
   <dim:field mdschema="dc" element="date" qualifier="available">2025-08-01T22:19:48Z</dim:field>
   <dim:field mdschema="dc" element="date" qualifier="issued">2025-08-01</dim:field>
   <dim:field mdschema="dc" element="date" qualifier="submitted">2026</dim:field>
   <dim:field mdschema="dc" element="identifier" qualifier="other">Wang_washington_0250E_28556.pdf</dim:field>
   <dim:field mdschema="dc" element="identifier" qualifier="uri">https://hdl.handle.net/1773/53512</dim:field>
   <dim:field mdschema="dc" element="description">Thesis (Ph.D.)--University of Washington, 2026</dim:field>
   <dim:field mdschema="dc" element="description" qualifier="abstract">Pretrained Language Models (LMs) have demonstrated remarkable general-purpose capabilities by encoding vast amounts of knowledge from the internet. However, effectively steering these models to serve diverse downstream applications, such as following instructions, chatting with users, using tools, or performing complex reasoning, poses another set of challenges that require diverse, high-quality, and increasingly costly training data. This dissertation explores scalable paradigms for structuring, creating, and optimizing data to facilitate the broader generalization of language models and enhance their critical capabilities.First, through the creation of the SuperNaturalInstructions benchmark—a large-scale dataset with over 1,600 NLP tasks—I demonstrate that unifying NLP tasks via natural language instructions enables model generalization at the task level. Second, I propose Self-Instruct, a novel framework where LMs generate their own instructional data to train themselves, thereby demonstrating model self-improvement. Third, I develop HyPER, a framework that routes preference annotation tasks between humans and AI to optimize data quality and collection efficiency for preference-based learning. Finally, I systematically study the impact of diverse open instruction-tuning datasets on LM capabilities, leading to the development of the Tülu series of openly available and highly capable models. Together, these efforts—unifying task structures, leveraging model-generated synthetic data, optimizing human-AI data partnerships, and fostering open data ecosystems—have demonstrated an effective path to building a strong, scalable, and community-driven data foundation for post-training language models. Finally, I envision future directions that can further enhance this data foundation for building more advanced and sustainable AI systems.</dim:field>
   <dim:field mdschema="dc" element="format" qualifier="mimetype">application/pdf</dim:field>
   <dim:field mdschema="dc" element="language" qualifier="iso">en_US</dim:field>
   <dim:field mdschema="dc" element="rights">CC BY</dim:field>
   <dim:field mdschema="dc" element="subject">Data Paradigms</dim:field>
   <dim:field mdschema="dc" element="subject">Data-centric AI</dim:field>
   <dim:field mdschema="dc" element="subject">Instruction Tuinng</dim:field>
   <dim:field mdschema="dc" element="subject">Language Models</dim:field>
   <dim:field mdschema="dc" element="subject">Post-Training</dim:field>
   <dim:field mdschema="dc" element="subject">Synthetic Data</dim:field>
   <dim:field mdschema="dc" element="subject">Artificial intelligence</dim:field>
   <dim:field mdschema="dc" element="subject" qualifier="other">Computer science and engineering</dim:field>
   <dim:field mdschema="dc" element="title">Scalable Data Paradigms for Steering General-Purpose Language Models</dim:field>
   <dim:field mdschema="dc" element="type">Thesis</dim:field>
   <dim:field mdschema="dc" element="embargo" qualifier="terms">Open Access</dim:field>
   <dim:field mdschema="others" element="access-status">open.access</dim:field>
</dim:dim>
</metadata></record></GetRecord></OAI-PMH>