Generative AI for Lifelike Digital Garment Visualization

relationships.isAuthorOf

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

The fashion industry is undergoing a rapid digital transformation. With consumers increasinglyshopping for clothes online and brands seeking more engaging, cost-effective tools for showcasing clothing online, the demand for generative AI tools for garment visualization is growing. Despite rapid progress in this space, particularly in virtual try-on, existing methods still struggle to accurately capture physical realism, such as how a garment drapes and flows in motion or how a garment fits on a real body. In this thesis, I introduce a suite of generative AI-based methods that address five core dimensions of this challenge – fabric dynamics, multi-view consistency, precise conditioning, temporal smoothness, and garment fit – advancing the state-of-the-art in realistic digital garment visualization. First, I present DreamPose, a diffusion-based method for synthesizing animated fashion videosfrom a still image. Given an input image and a driving pose sequence, DreamPose generates a video depicting realistic human and fabric motion. The novelty of this work is leveraging a pretrained textto- image model (Stable Diffusion) for image animation – proposing novel architectural adaptations to ensure temporal consistency and enable joint image-and-pose conditioning. A key challenge with this method is maintaining precise input image fidelity. We address this through a subject-specific finetuning strategy that improves person and garment identity, without sacrificing generalization to new pose sequences. DreamPose produces state-of-the-art results on the fashion video animation task across a diverse range of clothing styles and poses. Next, I introduce Fashion-VDM, a video diffusion model for video virtual try-on. Given a garmentimage and a person video, the model synthesizes a temporally consistent try-on video that renders the target garment onto the person, while preserving the person’s identity and motion. Although image-based virtual try-on has achieved impressive results, existing video try-on methods rely on multiple networks and suffer from poor quality due to warping artifacts. Fashion-VDM addresses these limitations through a unified diffusion transformer architecture augmented with temporal blocks, split classifier-free guidance, and a progressive temporal training strategy. Fashion-VDM is the first diffusion-based video virtual try-on method, and qualitative and quantitative experiments demonstrate that it outperforms existing methods by a wide margin. I then discuss HoloGarment, a novel view synthesis method for in-the-wild garments. Given oneto three images or a video of a person wearing a garment, HoloGarment generates 360◦ views of that garment in a canonical pose. This is a challenging task due to the significant occlusions, body pose variations, and cloth deformations present in real-world imagery. Prior methods fail to handle these conditions, because existing 3D datasets mostly consist of static, unoccluded, and rigid objects. Our key insight is to bridge this domain gap through an implicit training paradigm that jointly leverages large-scale real video data and small-scale synthetic 3D data to optimize a shared 2D-3D garment embedding space. At inference time, this shared space enables HoloGarment to transform real-world garments into static, canonical 360◦ spin visualizations. Furthermore, HoloGarment can leverage this shared space to handle dynamic video-to-360◦ novel view synthesis via a garment-specific “atlas” finetuning stage that captures garment geometry and texture across all viewpoints into a single garment embedding. Extensive experiments demonstrate that HoloGarment achieves state-of-the-art performance on novel view synthesis of in-the-wild garments, robustly handling wrinkling, pose variation, and occlusion while maintaining photorealism, view consistency, and geometric accuracy. Finally, I describe how we take the first steps toward fit-aware virtual try-on – depicting notjust how a garment looks, but how it actually fits a specific person. Despite being crucial to the shopping experience, fit is largely ignored by existing virtual try-on methods, which instead default to generating well-fitted results regardless of body or garment size. This limitation is largely due to the lack of VTO training data that includes both size information and ill-fit scenarios (where the garment is too tight or too loose on the person). To bridge this gap, we introduce FIT (Fit-Inclusive Try-on), a large-scale dataset of over one million try-on samples accompanied by precise body and garment measurements. FIT is constructed with a scalable synthetic pipeline: 3D garments are programmatically generated and draped using physics simulation, then transformed into photorealistic images through a novel re-texturing framework that strictly preserves person and garment geometry. Leveraging FIT, we train Fit-VTO, a baseline fit-aware virtual try-on model

Description

Thesis (Ph.D.)--University of Washington, 2026

Citation

DOI