Towards Grounded Visual Understanding in Vision–Language Models

relationships.isAuthorOf

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

Vision-language models have made rapid progress, but strong performance on standard benchmarks does not necessarily indicate grounded visual understanding: the ability to connect language to the actual objects, attributes, relations, and dynamics present in visual input. This thesis studies how to build vision-language models with stronger grounded visual understanding, with contributions spanning evaluation, supervision, representation, and reasoning. First, this thesis shows that existing benchmarks can overestimate grounded understanding because models may exploit dataset artifacts rather than visual evidence. To address this limitation, it develops more reliable and fine-grained evaluation tools, including benchmarks designed to reduce annotation bias and a programmatic benchmark generation engine that enables query-driven analysis of model capabilities across a broad range of image and video tasks. Second, this thesis studies how richer supervision can improve grounding. It introduces a scalable pipeline for generating vision-language instruction data from structured visual representations, covering objects, attributes, relations, segmentation, and depth. The resulting large-scale instruction data substantially improves performance on diverse visual understanding benchmarks. Third, this thesis investigates visual representation learning for grounding, with a focus on tokenizations that better reflect underlying visual structure. In video, it develops object- and trajectory-centric tokenization methods that replace standard patch-based representations with more grounded visual units, improving efficiency and downstream performance while better preserving object identity over time. Finally, this thesis shows how grounding can serve as a foundation for multimodal reasoning. It presents open vision-language models that support pointing, tracking, and spatially precise reasoning over images and videos, enabled by new grounded datasets and training strategies. Together, these contributions advance a unified view of grounded visual understanding: models should not only describe visual content, but also localize, track, and reason over it in ways that are faithful to the depicted world. This thesis argues that such grounding is essential for making vision-language models more reliable, interpretable, and useful in real-world settings.

Description

Thesis (Ph.D.)--University of Washington, 2026

Citation

DOI