Structured Indoor Reconstruction from Sparse Unposed Panoramas: Coarse Prediction to Generative Geometric Refinement

relationships.isAuthorOf

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

Accurate indoor floor plans and camera poses are important for applications such as digital twins, augmented reality, real estate visualization, and immersive walkthrough generation. However, reconstructing structured indoor layouts from sparse, unposed, wide-baseline panoramic images remains difficult due to pose uncertainty, limited cross-view overlap, occlusion, and the inherently global nature of floor-plan structure. This thesis studies learning-based methods for structured indoor reconstruction in this setting, with the goal of recovering precise camera pose, geometry, and multi-room floor-plan layouts from sparse panoramic observations. The thesis begins by establishing this problem setting through practical reconstruction pipelines, public datasets, and targeted studies of wide-baseline pose and layout estimation. Early human-in-the-loop systems show that accurate cross-view alignment is essential for constructing precise global floor plans, while also revealing the limitations of pipelines that primarily stitch together single-view layouts. To support systematic study, the thesis develops public benchmarks and cross-view annotations, and investigates targeted subproblems such as wide-baseline pose estimation and joint multi-view layout prediction. These early efforts establish two key insights: reliable indoor reconstruction depends critically on cross-view reasoning, and direct prediction of full multi-view, multi-room layouts is difficult to learn under sparse data, pose ambiguity, and noisy supervision. Motivated by these observations, the thesis reformulates the problem through Multi-view Ordered Wall Instance Segmentation (MOIS), which represents layouts using cross-view wall instances and within-room connectivity rather than direct global wall-sequence prediction. Under this formulation, the thesis introduces MvFFN, a feed-forward multi-view model that jointly predicts coarse camera poses, dense global geometry, and wall-instance-level structural outputs, which are assembled into globally consistent multi-room floor plans. This makes the reconstruction problem substantially more tractable and enables automated recovery of structured layouts from sparse panoramas. The thesis then studies refinement beyond feed-forward prediction. While MvFFN provides strong coarse reconstructions, challenging scenes with wide-baseline ambiguity, sparse observations, and noisy boundary evidence still require scene-level correction. This motivates BADGR, whose central insight is that learned generative priors and iterative geometric optimization should not be treated as separate stages, but as complementary processes that can be learned to improve one another. BADGR combines structural generation and bundle-adjustment-style refinement to jointly improve pose and layout consistency under global scene constraints. Experiments show improved wall, junction, and pose accuracy over prior baselines on challenging indoor benchmarks. Together, these results show that accurate indoor reconstruction from sparse panoramas requires the integration of cross-view reasoning, structured intermediate representations, learned structural priors, and global optimization.

Description

Thesis (Ph.D.)--University of Washington, 2026

Citation

DOI