HoliTok: A Continuous Holistic Tokenization with Robust Dual Capabilities of Speech Generation and Understanding
cs.SD, cs.AI, eess.AS
Submitted: 2026-05-28
Updated: 2026-09-05
Comments: 14 pages, 2 figures, 8 tables; Accepted by EMNLP 2026 Main Conference
Code: https://github.com/bovod-sjtu/HoliTok
License: http://creativecommons.org/licenses/by/4.0/
The gist: Unified speech foundation models require a holistic tokenization space that is both learnable by language models and decodable into high-quality waveforms.
Terminology
Abstract
Unified speech foundation models require a holistic tokenization space that is both learnable by language models and decodable into high-quality waveforms. Existing speech tokenizers, however, often fail to satisfy these requirements simultaneously, leading to increased architectural complexity and more involved training designs. We propose HoliTok, a continuous Holistic speech Tokenization model designed for unified generation-understanding modeling. HoliTok encodes 48 kHz speech into a compact 25 Hz sequence of 128-dimensional latents. It is trained with a progressive strategy that jointly preserves signal-level fidelity, incorporates semantic information, and maintains strong latent learnability. Based on this tokenization, we build a unified AR+DiT model for speech synthesis and recognition, where the same latent sequence supports both generation-specific and unified generation-understanding tasks. Experiments show that HoliTok achieves competitive reconstruction fidelity, improves generative learnability for high-quality and controllable synthesis, and, among the evaluated representations, is the only one that operates robustly in our unified generation-understanding architecture without additional optimization tricks. These results suggest that HoliTok serves as an effective speech tokenizer and a foundational representation interface for unified spoken language modeling. The code is available at: https://github.com/bovod-sjtu/HoliTok.
Sources
- MusicLM: Generating Music From Text
- Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations
- SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
- On the Distillation Loss Functions of Speech VAE for Unified Reconstruction, Understanding, and Generation
- DashengTokenizer: One layer is enough for unified audio understanding and generation
- AISHELL-2: Transforming Mandarin ASR Research Into Industrial Scale
- HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units
- CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
- High Fidelity Neural Audio Compression
- Unified Autoregressive Visual Generation and Understanding with Continuous Tokens
- Kimi-Audio Technical Report
- Qwen2.5 Technical Report
- Autoregressive Diffusion Transformer for Text-to-Speech Synthesis
- Semantic-VAE: Semantic-Alignment Latent Representation for Better Speech Synthesis
- GPT-4o System Card
- Qwen2.5-Omni Technical Report
- Qwen3-Omni Technical Report
- Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation
- WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling
Related papers
- Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
- SoundWeaver: Compositional Warm-Starting for Text-to-Audio Diffusion Serving
- WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
- Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment