Building Multimodal AI Systems for Perceiving and Reasoning in 3D Space
Date
relationships.isAuthorOf
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
This dissertation examines how multimodal AI systems can perceive, represent, and reason about three-dimensional spatial environments. Three-dimensional scene understanding presents a structural challenge for contemporary vision-language models: a system must process geometrically large visual inputs without discarding the spatial evidence they carry, ground its reasoning in explicit intermediate steps rather than opaque end-to-end prediction, and be evaluated in ways that distinguish genuine spatial competence from benchmark-specific shortcut learning. These three requirements---evidence-preserving representation, reasoning-explicit architecture, and evidence-sensitive evaluation---form the organizing framework of this work. The first part develops efficient representations for large-scale 3D perception. Multi-view RGB-D scenes can span tens of thousands of visual tokens, most of which are spatially redundant. Dynamic Token Compression addresses this by lifting visual tokens into a shared 3D coordinate frame and merging them by spatial proximity and semantic similarity, enabling zero-shot 3D question answering with over 90% token reduction while maintaining competitive performance on OpenEQA and ScanQA. Token Merging with Spatial Awareness extends this principle inside the visual encoder itself, using depth-derived spatial tokens to guide token merging in Vision Transformers and producing more spatially coherent compressed representations for embodied and spatial question answering. Together, these methods demonstrate that token reduction must be guided by geometric structure---not merely visual similarity or temporal order---to faithfully preserve the spatial evidence needed for downstream reasoning. The second part moves from representation to explicit spatial reasoning. Even with compact and geometrically faithful inputs, many spatial queries require multi-step inference: identifying candidate objects, evaluating relational geometry, and grounding conclusions in scene evidence. Reason3DVG demonstrates that a language model fine-tuned on automatically generated structured reasoning traces---rather than answer-only labels---can perform 3D visual grounding more reliably; a model trained on 3.2K such examples outperforms systems trained on datasets sixty times larger. A complementary spatial agent equips a reasoning language model with executable distance estimation, inclusion classification, and 3D geometry reconstruction tools, enabling it to handle complex warehouse and outdoor spatial queries through structured function calls without large-scale end-to-end fine-tuning. The third part examines whether standard benchmark accuracy is an adequate proxy for spatial understanding quality. Analysis reveals that some spatially fine-tuned models improve direct question-answering scores while simultaneously degrading their ability to describe camera motion---a pattern indicating shortcut learning rather than improved spatial competence. CaMo introduces the Spatial Narrative Score, an evaluation protocol that decouples evidence generation from answer prediction: a model first produces scene and camera-motion descriptions without seeing the question, and a separate proxy language model then reasons over the narrative to produce an answer. Models that rely on benchmark priors cannot succeed under this protocol. CaMo-3B, trained on camera-motion and semantic narrative supervision, achieves consistent gains across spatial question answering and camera motion captioning benchmarks without the divergent accuracy-evidence behavior observed in competing fine-tuned models. Taken together, the three parts trace a unified investigation from the efficient extraction of spatial evidence, to its explicit use in multi-step reasoning, to the verification that spatial evidence---and not statistical answer priors---underlies a model's outputs. The central argument is that reliable 3D spatial understanding requires coordinated attention to all three stages: how evidence is preserved during compression, how it is manipulated during reasoning, and how its presence or absence is measured during evaluation.
Description
Thesis (Ph.D.)--University of Washington, 2026
