LiveVVT: High-Fidelity Video Virtual Try-On in Real Time
cs.CV, cs.AI
Submitted: 2026-08-27
Updated: 2026-08-28
Comments: 16 pages, 13 figures,
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Diffusion-based Video Virtual Try-On (VVT) achieves high visual fidelity through bidirectional spatio-temporal modeling, but complete-clip dependence incurs prohibitive latency and computational
Terminology
Abstract
Diffusion-based Video Virtual Try-On (VVT) achieves high visual fidelity through bidirectional spatio-temporal modeling, but complete-clip dependence incurs prohibitive latency and computational overhead in practical continuous deployment. Naively enforcing causality disrupts pretrained bidirectional priors and substantially degrades synthesis quality. We introduce LiveVVT, a rolling streaming diffusion framework that preserves bounded bidirectional modeling within causal recurrent generation. Within a fixed-size window, LiveVVT jointly denoises multiple video chunks under bounded look-ahead, preserving local bidirectional interactions while emitting one clean chunk per iteration. Beyond the window, two complementary memories sustain long-term consistency: a bounded temporal memory propagates recent dynamics and occlusion context, whereas a persistent global appearance memory, constructed once from the target garment and a frontal try-on keyframe, anchors garment details and dressed appearance throughout the stream. We further introduce a progressive distillation framework integrating bidirectional VVT learning, teacher-trajectory regression for causal few-step adaptation, and Collaborative Matching Distillation, which couples teacher-distribution matching with rolling flow matching on real videos to align optimization with recurrent inference. Experiments on paired and unpaired long-sequence benchmarks demonstrate superior generation quality over similarly sized models, with 26 times lower latency and 11 times higher throughput, enabling high-fidelity real-time streaming VVT.
Sources
- Qwen2.5-VL Technical Report
- UniVVT: A Unified End-to-End Framework for High-Fidelity Video Virtual Try-on
- CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation
- ViViD: Video Virtual Try-on using Diffusion Models
- FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space
- MagicTryOn: Harnessing Diffusion Transformer for Garment-Preserving Video Virtual Try-on
- FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization
- Wan: Open and Advanced Large-Scale Video Generative Models
- VideoGPT: Video Generation using VQ-VAE and Transformers
- Eevee: Towards Close-up High-resolution Video-based Virtual Try-on
- Causal Forcing++: Scalable Few-Step Autoregressive Diffusion Distillation for Real-Time Interactive Video Generation
- iTryOn: Mastering Interactive Video Virtual Try-On with Spatial-Semantic Guidance
- Video Virtual Try-on with Conditional Diffusion Transformer Inpainter
- DreamVVT: Mastering Realistic Video Virtual Try-On in the Wild via a Stage-Wise Diffusion Transformer Framework
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models