A Comprehensive Study of Content Representations for Speech Synthesis
eess.AS, cs.LG, cs.SD
Submitted: 2026-09-25
Updated: 2026-09-25
Code: https://github.com/interactiveaudiolab/penn7https:
Terminology
Sources
- CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
- AudioPaLM: A Large Language Model That Can Speak and Listen
- Rethinking Discrete Speech Representation Tokens for Accent Generation
- Canary-1B-v2 & Parakeet-TDT-0.6B-v3: Efficient and High-Performance Models for Multilingual ASR and AST
- Cross-domain Neural Pitch and Periodicity Estimation
Related papers
- X-VC: Zero-shot Streaming Voice Conversion in Codec Space
- Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
- Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
- Towards Audio Token Compression in Large Audio Language Models
- WaveScat: Wavelet Scattering Front-Ends with Self-Supervised Features for Speech Deepfake Detection
- ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions