Semantic Heterogeneity in the Age of Foundation Models: Benchmarks and Systems for Opaque Enterprise Schemas

dc.contributor.advisorSuciu, Dan
dc.contributor.authorKayali, Moe
dc.date.accessioned2026-09-16T18:24:53Z
dc.date.issued2026-09-16
dc.date.submitted2026
dc.descriptionThesis (Ph.D.)--University of Washington, 2026
dc.description.abstractSemantic heterogeneity—the fact that the same real-world concept is represented differently across databases—is the principal obstacle to integrating, querying, and analyzing data across organizational boundaries. Its cost is concrete: analysts spend 40% of their working hours loading and cleaning data, and 60–70% of enterprise data is never used for analytics. The difficulty scales with the schema: enterprise systems routinely expose more than 10^5 attributes under opaque, proprietary vocabularies that no analyst—and no publicly trained model—can resolve by lookup. This dissertation asks how foundation models change the treatment of semantic heterogeneity: which data-management tasks they help solve, and how they compare against four decades of schema-matching and data-integration techniques. It answers through three systems, each closing a distinct gap. CHORUS brings foundation models to data discovery, unifying table-class detection, column-type annotation, and join-column prediction in a single prompt-based architecture with an anchoring mitigation against hallucination; it matches or exceeds task-specific deep models trained on 10^5-10^6 labelled examples, reaching 0.93, 0.89, and 0.90 F1 on the three tasks with no per-task training. Goby confronts the enterprise gap that public benchmarks conceal: it releases the first data-integration benchmark drawn from real, semi-private enterprise data, quantifies the performance drop that both LLM- and ML-based methods suffer on it, and closes that drop by learning a working ontology from data samples and serializing its full hierarchy into the model. Query Blueprints attacks the entanglement of structure and vocabulary in natural-language-to-query translation: by letting the user specify query structure exactly and grounding only the unknown vocabulary in data, it answers queries over opaque schemas where end-to-end LLM agents collapse—34% accuracy versus under 10% for four state-of-the-art agents on a scrambled-schema benchmark, at $0.40 versus $13.03 in cost. Together the three systems trace the contemporary heterogeneity stack: foundation models can replace task-specific training on public schemas, but enterprise data demands that structure be factored from vocabulary, and that vocabulary be grounded in data rather than recalled from training.
dc.embargo.termsOpen Access
dc.format.mimetypeapplication/pdf
dc.identifier.otherKayali_washington_0250E_30182.pdf
dc.identifier.urihttps://hdl.handle.net/1773/57748
dc.language.isoen_US
dc.rightsnone
dc.subjectComputer science
dc.subject.otherComputer science and engineering
dc.titleSemantic Heterogeneity in the Age of Foundation Models: Benchmarks and Systems for Opaque Enterprise Schemas
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Kayali_washington_0250E_30182.pdf
Size:
2.93 MB
Format:
Adobe Portable Document Format