Same Bit Width, Different Outcomes: Post-Training Quantization of Text-to-Speech Across Architectures
eess.AS, cs.LG, cs.SD
Submitted: 2026-09-24
Updated: 2026-09-24
Code: https://github.com/uxfacdev/tts-ptq-map
Terminology
Sources
- OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models
- Streaming Sequence-to-Sequence Learning with Delayed Streams Modeling
- VoxCPM: Tokenizer-Free TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
- No Language Left Behind: Scaling Human-Centered Machine Translation
Related papers
- X-VC: Zero-shot Streaming Voice Conversion in Codec Space
- Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
- Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
- Towards Audio Token Compression in Large Audio Language Models
- WaveScat: Wavelet Scattering Front-Ends with Self-Supervised Features for Speech Deepfake Detection
- ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions