MELD: Mel-Spectrogram-Based Speech Language Modeling with Discrete Latent Variables
eess.AS, cs.CL
Submitted: 2026-05-28
Updated: 2026-09-01
Code: https://github.com/descriptinc/descript-audio-codec
Terminology
Sources
- On The Landscape of Spoken Language Models: A Comprehensive Survey
- dMel: Speech Tokenization made Simple
- VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
- Moshi: a speech-text foundation model for real-time dialogue
- Multimodal Latent Language Modeling with Next-Token Diffusion
- Speech Synthesis From Continuous Features Using Per-Token Latent Diffusion
- Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
- VoxCPM: Tokenizer-Free TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
Related papers
- X-VC: Zero-shot Streaming Voice Conversion in Codec Space
- Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
- Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
- Towards Audio Token Compression in Large Audio Language Models
- WaveScat: Wavelet Scattering Front-Ends with Self-Supervised Features for Speech Deepfake Detection
- ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions