Towards Grounded Visual Understanding in Vision–Language Models

dc.contributor.advisorKrishna, Ranjay R
dc.contributor.advisorRatner, Alexander A
dc.contributor.authorZhang, Jieyu
dc.date.accessioned2026-08-11T19:26:35Z
dc.date.issued2026-08-11
dc.date.submitted2026
dc.descriptionThesis (Ph.D.)--University of Washington, 2026
dc.description.abstractVision-language models have made rapid progress, but strong performance on standard benchmarks does not necessarily indicate grounded visual understanding: the ability to connect language to the actual objects, attributes, relations, and dynamics present in visual input. This thesis studies how to build vision-language models with stronger grounded visual understanding, with contributions spanning evaluation, supervision, representation, and reasoning. First, this thesis shows that existing benchmarks can overestimate grounded understanding because models may exploit dataset artifacts rather than visual evidence. To address this limitation, it develops more reliable and fine-grained evaluation tools, including benchmarks designed to reduce annotation bias and a programmatic benchmark generation engine that enables query-driven analysis of model capabilities across a broad range of image and video tasks. Second, this thesis studies how richer supervision can improve grounding. It introduces a scalable pipeline for generating vision-language instruction data from structured visual representations, covering objects, attributes, relations, segmentation, and depth. The resulting large-scale instruction data substantially improves performance on diverse visual understanding benchmarks. Third, this thesis investigates visual representation learning for grounding, with a focus on tokenizations that better reflect underlying visual structure. In video, it develops object- and trajectory-centric tokenization methods that replace standard patch-based representations with more grounded visual units, improving efficiency and downstream performance while better preserving object identity over time. Finally, this thesis shows how grounding can serve as a foundation for multimodal reasoning. It presents open vision-language models that support pointing, tracking, and spatially precise reasoning over images and videos, enabled by new grounded datasets and training strategies. Together, these contributions advance a unified view of grounded visual understanding: models should not only describe visual content, but also localize, track, and reason over it in ways that are faithful to the depicted world. This thesis argues that such grounding is essential for making vision-language models more reliable, interpretable, and useful in real-world settings.
dc.embargo.termsOpen Access
dc.format.mimetypeapplication/pdf
dc.identifier.otherZhang_washington_0250E_29375.pdf
dc.identifier.urihttps://hdl.handle.net/1773/57235
dc.language.isoen_US
dc.rightsCC BY
dc.subjectComputer Vision
dc.subjectLanguage Models
dc.subjectMachine Learning
dc.subjectMultimodal Models
dc.subjectVision-language Models
dc.subjectArtificial intelligence
dc.subjectComputer science
dc.subject.otherComputer science and engineering
dc.titleTowards Grounded Visual Understanding in Vision–Language Models
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Zhang_washington_0250E_29375.pdf
Size:
41.28 MB
Format:
Adobe Portable Document Format