CTC-TTS: LLM-based dual-streaming text-to-speech with CTC alignment
eess.AS, cs.AI, cs.SD
Submitted: 2026-02-23
Updated: 2026-09-07
Comments: Fix two typos in Figure 2 in INTERSPEECH 2026 version
Project page: https://ctctts.github.io
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- SyncSpeech: Efficient and Low-Latency Text-to-Speech based on Temporal Masked Transformer
- Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
- CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
- CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
- Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming
- Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis
- SpeakStream: Streaming Text-to-Speech with Interleaved Data
Related papers
- X-VC: Zero-shot Streaming Voice Conversion in Codec Space
- Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
- Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
- Towards Audio Token Compression in Large Audio Language Models
- WaveScat: Wavelet Scattering Front-Ends with Self-Supervised Features for Speech Deepfake Detection
- ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions