Asynchronous Pipeline Parallelism for Real-Time Multilingual Lip Synchronization in Video Communication Systems
cs.MM, cs.AI, cs.CV, cs.DC, cs.NI
Submitted: 2025-12-20
Updated: 2025-12-20
Comments: Accepted to IEEE Big Data 2025, AIDE4IoT Workshop. Copyright \c{opyright} 2025 IEEE
DOI: 10.1109/BigData66926.2025.11400775
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In The Wild
- Listen, Attend and Spell
- Deep Speech 2: End-to-End Speech Recognition in English and Mandarin
- Conformer: Convolution-augmented Transformer for Speech Recognition
- Sequence to Sequence Learning with Neural Networks
- Effective Approaches to Attention-based Neural Machine Translation
- Tacotron: Towards End-to-End Speech Synthesis
- FastSpeech: Fast, Robust and Controllable Text to Speech
- FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
- YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot Voice Conversion for everyone
- AudioPaLM: A Large Language Model That Can Speak and Listen
- Diff2Lip: Audio Conditioned Diffusion Models for Lip-Synchronization
- LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision
- AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head Synthesis
- SyncTalk: The Devil is in the Synchronization for Talking Head Synthesis
- A Bridge from Audio to Video: Phoneme-Viseme Alignment Allows Every Face to Speak Multiple Languages
- S$^3$FD: Single Shot Scale-invariant Face Detector
Related papers
- ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling
- Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models
- Mitigating GenAI-powered Evidence Pollution for Out-of-Context Multimodal Misinformation Detection
- A Rate-Distortion-Classification Approach for Lossy Image Compression