Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning
eess.AS, cs.CL, cs.SD
Submitted: 2026-09-29
Updated: 2026-09-29
Code: https://github.com/textstat/textstat
Terminology
Sources
- Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
- Seamless: Multilingual Expressive and Streaming Speech Translation
- SoundStorm: Efficient Parallel Audio Generation
- IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
- CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
- CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
- FunASR: A Fundamental End-to-End Speech Recognition Toolkit
- Distilling the Knowledge in a Neural Network
- BigVGAN: A Universal Neural Vocoder with Large-Scale Training
- Flow Matching for Generative Modeling
- Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow
- FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
- UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022
- Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
- CAM++: A Fast and Efficient Network for Speaker Verification Using Context-Aware Masking
- Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Related papers
- X-VC: Zero-shot Streaming Voice Conversion in Codec Space
- Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
- Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
- Towards Audio Token Compression in Large Audio Language Models
- WaveScat: Wavelet Scattering Front-Ends with Self-Supervised Features for Speech Deepfake Detection
- ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions