Not All Policies Are Equal: Training Flow-Based Policies for Efficient Downstream Steering

relationships.isAuthorOf

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

Generative policies based on diffusion and flow matching can represent multimodal robotaction distributions and have become a practical foundation for visuomotor imitation learning. However, these policies are usually trained as action generators rather than as controllable latent- action maps. This thesis studies the gap between imitation performance and downstream steerability in flow-based robot policies. The central question is whether a pretrained flow policy’s latent variables provide a stable and useful coordinate system for downstream control, including supervised latent prediction and reinforcement-learning-based latent steering. A key finding is that low flow-matching loss does not guarantee a steerable latent space. Vanilla flow policies can achieve good action reconstruction while learning ill-conditioned maps: small action differences can produce large changes in recovered noise, observation perturbations can be amplified through inversion, and many different initial noise samples can decode to a narrow set of action trajectories. These effects make the latent variable unreliable for downstream policies that operate in noise space. To study this failure mode, diagnostics are developed for latent geometry, including inverse noise norms, noise round-trip consistency, finite-step contraction and expansion, forward and inverse sensitivity, observation-induced action drift, action coverage under fixed observations, and gradient conflict between flow-matching and regularization objectives. Geometry-aware regularizers are evaluated, including Bi-Lipschitz penalties on the flow state and observation-sensitivity penalties on visual and proprioceptive conditioning. In custom flow policies, these regularizers improve latent conditioning, produce broader but structured action coverage, and enable successful downstream steering in real-world manipulation tasks. In contrast, scaling the same ideas to π0-style vision- language-action flow policies on LIBERO exposes an important trade-off: although regularization often improves inversion stability, it can also reduce direct policy performance by interfering with task-relevant visual grounding, grasp timing, and placement behavior. The LIBERO results therefore serve not as a complete solution, but as evidence that geometry-aware regularization must be applied selectively. Overall, the current evidence suggests that steerable flow policies should not be regularized uniformly across all timesteps and inputs. Instead, regularization should be applied selectively: stronger constraints can help during reaching and grasping, where stable latent control is important, while weaker constraints are needed during placement, where the policy must preserve precise visual grounding and fine action control. Similarly, visual, proprioceptive, and language inputs should be regularized differently, because they carry different kinds of task-relevant information.

Description

Thesis (Master's)--University of Washington, 2026

Citation

DOI