Towards Human-AI Symbiotic Spatial Perception

relationships.isAuthorOf

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

Artificial intelligence has become an indispensable part of digital life. People routinely use AI to write, program, search, summarize, and make decisions. Yet this emerging human-AI symbiosis remains concentrated in the digital realm, where people work primarily with language through digital devices. It is far less developed in the physical world, where people navigate and act under real stakes. People must still move through unfamiliar places, manipulate objects, assess risks, and make decisions through their own situated perception and at the cost of their own time and effort. This limitation does not arise from a lack of AI research on the physical world. Over several decades, computer vision, robotics, augmented reality, and embodied AI have produced systems that recognize objects, reconstruct environments, navigate, plan motion, and increasingly predict how real-world scenes may evolve. The challenge is to translate these capabilities into assistance that fits human activity. Spatial evidence is viewpoint-dependent, incomplete, multimodal, and subject to change. Its usefulness also depends on who will act, what that person is trying to accomplish, and how an inferred result is communicated. A model can therefore perform well on a general benchmark while remaining too slow, insufficiently grounded, mismatched to a person's capabilities, or too opaque for a consequential spatial decision. This dissertation advances a vision of human-AI symbiotic spatial perception: a perceptual partnership in which people and AI systems combine complementary knowledge and capabilities to acquire, interpret, and communicate information about physical space. I argue that useful spatial assistance depends on designing AI around both people and the physical world. Such systems must deliver timely, in-context output; acquire task-relevant spatial evidence beyond a person's direct presence; align spatial judgments with human capabilities and goals; and represent spatial data in forms that people can perceive, inspect, and manipulate. The dissertation investigates these requirements through four research frontiers and six projects. Frontier I, In-Context, Real-Time Personalized Spatial Assistance, examines RASSAR, a mobile augmented-reality system for accessibility and safety assessment. Frontier II, Spatial Perception Beyond Physical Presence, explores automated and human-guided spatial capture through RAIS and FlyMeThrough. Frontier III, Aligning AI Spatial Reasoning with Human Needs, introduces CapNav, a benchmark of capability-conditioned indoor navigation. Frontier IV, Human-Centered Spatial Representation and Interaction, investigates object- and event-centered techniques for spatial authoring through SonifyAR and DepthScape. Across the projects, the findings show that useful assistance can emerge even when AI capabilities remain imperfect. RASSAR demonstrates the feasibility of mobile, stakeholder-informed accessibility scanning while exposing the need for personalization and human verification. RAIS and FlyMeThrough show that robots and drones can extend the reach of spatial capture, but that predefined task criteria or sparse human annotations are needed to preserve relevance. CapNav reveals that strong vision-language models degrade when navigation depends on embodiment-specific constraints, metric estimates, and evidence integrated across views. SonifyAR and DepthScape show how intermediate representations, including contextual event descriptions and parametric spatial anchors, let people inspect and shape AI-generated outputs. Together, these contributions establish a human-centered design agenda for spatial AI that is timely, far-reaching, capability-aligned, expressive, and accountable to the people who use it.

Description

Thesis (Ph.D.)--University of Washington, 2026

Citation

DOI