TexTailor: Texture-Preserving Video Virtual Try-On via Adaptive Garment Conditioning
cs.CV
Submitted: 2026-09-30
Updated: 2026-09-30
Terminology
Sources
- Dress&Dance: Dress up and Dance as You Like It - Technical Preview
- CatVTON: Concatenation Is All You Need for Virtual Try-On with Diffusion Models
- CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation
- MAGREF: Masked Guidance for Any-Reference Video Generation with Subject Disentanglement
- ViViD: Video Virtual Try-on using Diffusion Models
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning
- PhysEditWorld: A Large-Scale Dataset Toward Physics-Editable World Models
- Identity-Consistent Video Generation under Large Facial-Angle Variations
- LoRA: Low-Rank Adaptation of Large Language Models
- Embedding-perturbed Exploration Preference Optimization for Flow Models
- HunyuanVideo: A Systematic Framework For Large Video Generative Models
- MagicTryOn: Harnessing Diffusion Transformer for Garment-Preserving Video Virtual Try-on
- RealVVT: Towards Photorealistic Video Virtual Try-on via Spatio-Temporal Consistency
- Flow Matching for Generative Modeling
- Controllable Layer Decomposition for Reversible Multi-Layer Image Generation
- Dreamix: Video Diffusion Models are General Video Editors
- DINOv2: Learning Robust Visual Features without Supervision
- Once Is Enough: Lightweight DiT-Based Video Virtual Try-On via One-Time Garment Appearance Injection
- ModaFlow: Modality-Aware Flow Matching for High-Fidelity Virtual Try-On
- DINOv3
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models