Beyond Lip Sync: Reference-Grounded Oral Refinement for Audio-Driven Portrait Animation
cs.CV
Submitted: 2026-09-29
Updated: 2026-09-29
Code: https://github.com/MRzzm/HDTF
Terminology
Sources
- MultiDiffusion: Fusing Diffusion Paths for Controlled Image Generation
- VideoReTalking: Audio-based Lip Synchronization for Talking Head Video Editing In the Wild
- Generative Adversarial Networks
- Deep Residual Learning for Image Recognition
- From Inpainting to Editing: Unlocking Robust Mask-Free Visual Dubbing via Generative Bootstrapping
- ReF-LDM: A Latent Diffusion Model for Reference-based Face Image Restoration
- Image-to-Image Translation with Conditional Adversarial Networks
- Robust Reference-based Super-Resolution via C2-Matching
- Perceptual Losses for Real-Time Style Transfer and Super-Resolution
- Ensembling Off-the-shelf Models for GAN Training
- MaskGAN: Towards Diverse and Interactive Facial Image Manipulation
- LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision
- PhotoMaker: Customizing Realistic Human Photos via Stacked ID Embedding
- Identity-Preserving Video Dubbing Using Motion Warping
- FaceMe: Robust Blind Face Restoration with Personal Identification
- The Contextual Loss for Image Transformation with Non-Aligned Data
- Conditional Generative Adversarial Nets
- cGANs with Projection Discriminator
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models