Generative AI for Lifelike Digital Garment Visualization
Date
relationships.isAuthorOf
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
The fashion industry is undergoing a rapid digital transformation. With consumers increasinglyshopping for clothes online and brands seeking more engaging, cost-effective tools for showcasing
clothing online, the demand for generative AI tools for garment visualization is growing. Despite
rapid progress in this space, particularly in virtual try-on, existing methods still struggle to accurately
capture physical realism, such as how a garment drapes and flows in motion or how a garment fits
on a real body. In this thesis, I introduce a suite of generative AI-based methods that address five
core dimensions of this challenge – fabric dynamics, multi-view consistency, precise conditioning,
temporal smoothness, and garment fit – advancing the state-of-the-art in realistic digital garment
visualization. First, I present DreamPose, a diffusion-based method for synthesizing animated fashion videosfrom a still image. Given an input image and a driving pose sequence, DreamPose generates a video
depicting realistic human and fabric motion. The novelty of this work is leveraging a pretrained textto-
image model (Stable Diffusion) for image animation – proposing novel architectural adaptations
to ensure temporal consistency and enable joint image-and-pose conditioning. A key challenge with
this method is maintaining precise input image fidelity. We address this through a subject-specific
finetuning strategy that improves person and garment identity, without sacrificing generalization to
new pose sequences. DreamPose produces state-of-the-art results on the fashion video animation
task across a diverse range of clothing styles and poses. Next, I introduce Fashion-VDM, a video diffusion model for video virtual try-on. Given a garmentimage and a person video, the model synthesizes a temporally consistent try-on video that renders
the target garment onto the person, while preserving the person’s identity and motion. Although
image-based virtual try-on has achieved impressive results, existing video try-on methods rely on
multiple networks and suffer from poor quality due to warping artifacts. Fashion-VDM addresses
these limitations through a unified diffusion transformer architecture augmented with temporal
blocks, split classifier-free guidance, and a progressive temporal training strategy. Fashion-VDM is
the first diffusion-based video virtual try-on method, and qualitative and quantitative experiments
demonstrate that it outperforms existing methods by a wide margin. I then discuss HoloGarment, a novel view synthesis method for in-the-wild garments. Given oneto three images or a video of a person wearing a garment, HoloGarment generates 360◦ views
of that garment in a canonical pose. This is a challenging task due to the significant occlusions,
body pose variations, and cloth deformations present in real-world imagery. Prior methods fail to
handle these conditions, because existing 3D datasets mostly consist of static, unoccluded, and
rigid objects. Our key insight is to bridge this domain gap through an implicit training paradigm
that jointly leverages large-scale real video data and small-scale synthetic 3D data to optimize a
shared 2D-3D garment embedding space. At inference time, this shared space enables HoloGarment
to transform real-world garments into static, canonical 360◦ spin visualizations. Furthermore,
HoloGarment can leverage this shared space to handle dynamic video-to-360◦ novel view synthesis
via a garment-specific “atlas” finetuning stage that captures garment geometry and texture across all
viewpoints into a single garment embedding. Extensive experiments demonstrate that HoloGarment
achieves state-of-the-art performance on novel view synthesis of in-the-wild garments, robustly
handling wrinkling, pose variation, and occlusion while maintaining photorealism, view consistency,
and geometric accuracy. Finally, I describe how we take the first steps toward fit-aware virtual try-on – depicting notjust how a garment looks, but how it actually fits a specific person. Despite being crucial to
the shopping experience, fit is largely ignored by existing virtual try-on methods, which instead
default to generating well-fitted results regardless of body or garment size. This limitation is
largely due to the lack of VTO training data that includes both size information and ill-fit scenarios
(where the garment is too tight or too loose on the person). To bridge this gap, we introduce FIT
(Fit-Inclusive Try-on), a large-scale dataset of over one million try-on samples accompanied by
precise body and garment measurements. FIT is constructed with a scalable synthetic pipeline: 3D
garments are programmatically generated and draped using physics simulation, then transformed
into photorealistic images through a novel re-texturing framework that strictly preserves person and
garment geometry. Leveraging FIT, we train Fit-VTO, a baseline fit-aware virtual try-on model
Description
Thesis (Ph.D.)--University of Washington, 2026
