Learnable Classifier-Free Guidance Null Embeddings for Enhanced Controllable Speech Synthesis
eess.AS, cs.LG
Submitted: 2026-09-21
Updated: 2026-09-21
Comments: Accepted paper at Interspeech 2026
Project page: https://airtimemedia.github.io/IS2026-LearnableCFG
License: http://creativecommons.org/licenses/by/4.0/
The gist: Classifier-free Guidance (CFG) is widely adopted in text-to-speech (TTS) systems to enhance generation quality and conditioning fidelity by interpolating between conditioned and unconditioned
Terminology
Abstract
Classifier-free Guidance (CFG) is widely adopted in text-to-speech (TTS) systems to enhance generation quality and conditioning fidelity by interpolating between conditioned and unconditioned predictions. A common unconditional technique is to use an empty representation, in the form of a fixed null vector. In this work, we propose replacing this representation with a learnable unconditional embedding, optimized to represent a meaningful unconditional state. Objective and subjective evaluations demonstrate that learnable null embeddings consistently outperform fixed null embeddings across speaker similarity, speech stability, and expressiveness, while exhibiting greater robustness to larger guidance scales. We further show that learning a distinct unconditional embedding for each of the TTS conditioning modalities allows fine-grained control over speaker and text guidance, showcasing the trade-off between similarity and quality, and stability and expressiveness in the generated speech.
Sources
- Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
- Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis
- Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations
- MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
- VibeVoice Technical Report
- VoxCPM2 Technical Report
- CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
- IndexTTS 2.5 Technical Report
- Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
- MOSS-TTS Technical Report
- dots.tts Technical Report
- FCPE: A Fast Context-based Pitch Estimation Model
Related papers
- X-VC: Zero-shot Streaming Voice Conversion in Codec Space
- Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
- Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
- Towards Audio Token Compression in Large Audio Language Models
- WaveScat: Wavelet Scattering Front-Ends with Self-Supervised Features for Speech Deepfake Detection
- ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions