Semantic Heterogeneity in the Age of Foundation Models: Benchmarks and Systems for Opaque Enterprise Schemas
Date
relationships.isAuthorOf
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
Semantic heterogeneity—the fact that the same real-world concept is represented differently across databases—is the principal obstacle to integrating, querying, and analyzing data across organizational boundaries. Its cost is concrete: analysts spend 40% of their working hours loading and cleaning data, and 60–70% of enterprise data is never used for analytics. The difficulty scales with the schema: enterprise systems routinely expose more than 10^5 attributes under opaque, proprietary vocabularies that no analyst—and no publicly trained model—can resolve by lookup. This dissertation asks how foundation models change the treatment of semantic heterogeneity: which data-management tasks they help solve, and how they compare against four decades of schema-matching and data-integration techniques. It answers through three systems, each closing a distinct gap. CHORUS brings foundation models to data discovery, unifying table-class detection, column-type annotation, and join-column prediction in a single prompt-based architecture with an anchoring mitigation against hallucination; it matches or exceeds task-specific deep models trained on 10^5-10^6 labelled examples, reaching 0.93, 0.89, and 0.90 F1 on the three tasks with no per-task training. Goby confronts the enterprise gap that public benchmarks conceal: it releases the first data-integration benchmark drawn from real, semi-private enterprise data, quantifies the performance drop that both LLM- and ML-based methods suffer on it, and closes that drop by learning a working ontology from data samples and serializing its full hierarchy into the model. Query Blueprints attacks the entanglement of structure and vocabulary in natural-language-to-query translation: by letting the user specify query structure exactly and grounding only the unknown vocabulary in data, it answers queries over opaque schemas where end-to-end LLM agents collapse—34% accuracy versus under 10% for four state-of-the-art agents on a scrambled-schema benchmark, at $0.40 versus $13.03 in cost. Together the three systems trace the contemporary heterogeneity stack: foundation models can replace task-specific training on public schemas, but enterprise data demands that structure be factored from vocabulary, and that vocabulary be grounded in data rather than recalled from training.
Description
Thesis (Ph.D.)--University of Washington, 2026
