Robotics Foundation Models with Reasoning in the Loop
Date
relationships.isAuthorOf
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
Recent advances in generative AI have demonstrated the power of scaling: large language and vision models trained on internet-scale data now exhibit remarkable capabilities in perception, generation, and reasoning, often generalizing to tasks and domains far beyond those seen during training. These successes have inspired growing interest in bringing foundation-model paradigms to robotics, with the goal of moving beyond narrow, task-specific autonomy in constrained environments toward general-purpose robots that perceive, plan, and act robustly across diverse open-world settings. However, robotics fundamentally differs from language and vision in ways that resist a direct transfer of the scaling recipe. Robot learning cannot rely on the passive accumulation of internet data, since embodied interaction data must be physically collected through hardware, human teleoperation, or simulation, each of which is expensive, slow, and difficult to scale to the diversity of real-world tasks. Moreover, embodied data is inherently long-tailed: rare events, edge cases, and failure modes are precisely the situations in which robustness matters most, yet the hardest to obtain. As a result, simply scaling data and model parameters is insufficient. Building general-purpose, robust robotics foundation models therefore demands a different question: how can robots learn more from less data, generalize beyond their training distribution, and continue to improve over time through their own experience? This thesis argues that reasoning in the loop offers a promising path forward. Rather than treating reasoning as a downstream capability applied after learning, I show how reasoning can be integrated directly into the learning process itself, shaping how data is represented, how models are trained, and how policies are refined through deployment. By reasoning over structured feedback, temporal context, and failure, robots can extract substantially more signal from each interaction, compensating for data scarcity and improving generalization to unseen objects, scenes, and tasks. I develop this agenda along three complementary axes. First, I introduce approaches for spatial reasoning that enable robots to ground language in 3D space and reason explicitly about geometric and relational structure among objects, supporting precise manipulation in cluttered and previously unseen environments. Second, I develop temporal reasoning through memory-centric models that retain, query, and reason over past observations and actions, allowing robots to maintain coherent behavior across extended time horizons and to perform high-precision, long-horizon tasks that exceed the scope of purely reactive policies. Third, I show how reasoning over failures allows robots to diagnose why actions fail, attribute errors to specific causes, and use that understanding to self-improve, increasing robustness without additional human supervision or demonstrations. Taken together, these contributions reframe robotics foundation models as systems that learn through reasoning, closing the loop between perception, action, and structured inference. This perspective points toward a new generation of self-improving autonomous systems capable of operating reliably in the open world, and lays a foundation for scalable, data-efficient embodied intelligence.
Description
Thesis (Ph.D.)--University of Washington, 2026
