High-Accuracy Intraoperative 3D Reconstruction with Neural Rendering Technology
Date
relationships.isAuthorOf
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
Intraoperative three-dimensional (3D) understanding of patient anatomy is a long-standing requirement in image-guided surgery, yet routine intraoperative imaging remains limited by the radiation, cost, and workflow disruption associated with intraoperative computed tomography (iCT), and by the small baseline and calibration drift of physical stereo endoscopes. This dissertation argues that recent neural rendering techniques—principally Neural Radiance Fields (NeRF) and 3D Gaussian Splatting—can be redesigned from photorealistic view-synthesis tools into quantitative metrology instruments, recovering near-CT geometric accuracy over the surgically exposed field directly from routine monocular endoscopic video. The dissertation contributes two methods and supports two applications. The first method, software-defined stereo, fits a NeRF to monocular video and queries it as a virtual stereo rig at a computationally chosen baseline; a foundation stereo matcher then recovers metric depth from the synthesized pair. On the SCARED laparoscopic benchmark, this reduces absolute relative depth error from 0.0326 (calibrated physical stereo) to 0.0207, the first reported case of a virtual stereo pipeline outperforming a calibrated physical stereo endoscope. In a five-patient in vivo endoscopic sinus surgery (ESS) study, mean distance-to-mesh error reaches 0.411 ± 0.056 mm, with 94.3% of points within 1 mm of preoperative CT, surpassing the contemporaneous self-supervised benchmark of Mangulabnan et al. (0.91 mm) and the deep monocular baseline of Liu et al. The per-frame nature of this first method, however, limits the spatial extent of the reconstruction: dense depth maps from individual frames cannot be aggregated into a single, complete, large-volume 3D model because pose errors and multiview depth inconsistencies accumulate during fusion. The second method, NeVStereo, resolves this aggregation limitation by closing a loop of render–stereo–feedback–fusion. It combines confidence-guided multiview depth voting, NeRF-coupled bundle adjustment built on DROID-SLAM, and iterative depth–radiance-field refinement so that pose, depth, and the radiance field improve jointly. Across indoor (ScanNet++ and Replica), tabletop (NVIDIA-HOPE and MobileBrick), and aerial (WildUAV) benchmarks, it surpasses both feed-forward 3D systems and 3D Gaussian Splatting-based NVS-stereo pipelines on all four output modalities—depth, pose, novel-view synthesis, and reconstructed mesh. Two clinical applications are supported by these two methods. (1) NeRF-based measurements during sinus endoscopy: in cadaveric ethmoid anatomy, this approach reaches submillimeter one-dimensional measurement accuracy using routine monocular video. Because the same diagnostic workflow is not constrained by an intraoperative time budget, it also motivates a longitudinal measurement workflow for tracking polyp size and geometry during medical-therapy follow-up of chronic rhinosinusitis, pending prospective validation. (2) Endoscopic guidance of ESS: under an intraoperative time budget, the same reconstruction is co-registered to the preoperative CT so that the displayed anatomy reflects the evolving operative cavity. On 3D-printed Fusetec sinonasal phantoms, this achieves a Dice score of 0.930 and an average Hausdorff distance of 0.273 mm across six sinonasal sides. The high-accuracy optical reconstruction also underpins the Virtual Intraoperative CT (viCT) workflow of Gunderson et al., which sequentially updates the preoperative CT volume in its native DICOM grid throughout cadaveric ESS without radiation, ancillary hardware, or 30–40 minutes of workflow disruption. Together, these contributions establish a new design point for intraoperative 3D reconstruction in which rendering is engineered to serve geometry, a continuous neural scene representation can act as an active measurement instrument, and the algorithmic ceiling on large-volume, visible-surface monocular 3D reconstruction is no longer the per-frame aggregation barrier. NeVStereo is demonstrated on general indoor, tabletop, and aerial multiview benchmarks; its migration to monocular sinonasal endoscopic video, which would carry the same large-volume guarantee into the in vivo ESS setting, is identified as the principal next step of this research program.
Description
Thesis (Ph.D.)--University of Washington, 2026
