ComplexSync: High-Fidelity and Real-Time Lip Sync in Complex Scenarios
cs.CV
Submitted: 2026-09-24
Updated: 2026-09-24
Code: https://github.com/Playmate111/ComplexSync
Terminology
Sources
- KeySync: A Robust Approach for Leakage-free Lip Synchronization in High Resolution
- SAM 3: Segment Anything with Concepts
- TalkVid: A Large-Scale Diversified Dataset for Audio-Driven Talking Head Synthesis
- Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion
- Auto-Encoding Variational Bayes
- Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation
- LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision
- Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference
- SayAnything: Audio-Driven Lip Synchronization with Conditional Video Diffusion
- Playmate: Flexible Control of Portrait Animation via 3D-Implicit Space Guided Diffusion
- Playmate2: Training-Free Multi-Character Audio-Driven Animation via Diffusion Transformer with Reward Feedback
- DINOv2: Learning Robust Visual Features without Supervision
- SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
- DreamFusion: Text-to-3D using 2D Diffusion
- Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model
- FORA: Fast-Forward Caching in Diffusion Transformer Acceleration
- Efficient Diffusion Models: A Survey
- DINOv3
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Denoising Diffusion Implicit Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models