Not Quite My Tempo: Voice Activity-aware Speech Synthesis for Lip-Synchronous Dubbing
eess.AS, cs.AI
Submitted: 2026-09-22
Updated: 2026-09-22
Comments: accepted at Interspeech 2026
Code: https://github.com/snakers4/silero-vad
Project page: https://anondemos.github.io/NotQuiteMyTempo
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Automatic lip-synchronous dubbing requires a speech synthesis model to generate alternating voice and silence patterns in the target language that match the timing of the source clip precisely to
Terminology
Abstract
Automatic lip-synchronous dubbing requires a speech synthesis model to generate alternating voice and silence patterns in the target language that match the timing of the source clip precisely to ensure an optimal viewing experience. Prior works address this problem by conditioning the speech synthesis process on lip movements extracted from the video signal. In this work, we condition the speech generation on a binary voice-activity signal, which has a lightweight representation and can be produced in multiple ways. We show that the model follows the voice-activity signal with high accuracy while maintaining natural prosody and semantically appropriate pause placement within sentences, as demonstrated through extensive objective and subjective evaluations. By randomly masking this condition during training, we make the feature entirely optional during inference, allowing editors to enforce or relax lip-sync constraints when desired.
Sources
- ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching
- SyncVoice: Simple and Effective Automatic Video Dubbing with Vision-Augmented TTS
Related papers
- X-VC: Zero-shot Streaming Voice Conversion in Codec Space
- Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
- Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
- Towards Audio Token Compression in Large Audio Language Models
- WaveScat: Wavelet Scattering Front-Ends with Self-Supervised Features for Speech Deepfake Detection
- ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions