Decoder-Side Semantic Conditioning for Low-Bitrate Neural Speech Compression
cs.SD, cs.CL, cs.LG
Submitted: 2025-12-25
Updated: 2026-09-06
Comments: Accepted to APSIPA ASC 2026
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: Speech codecs are usually optimized for waveform fidelity, allocating bits to acoustic detail that can be inferred from linguistic structure.
Terminology
Abstract
Speech codecs are usually optimized for waveform fidelity, allocating bits to acoustic detail that can be inferred from linguistic structure. This leads to inefficient compression and degraded recognition performance. We propose SemDAC, a semantic-aware neural speech codec that adds hierarchical semantic conditioning to residual vector quantization (RVQ). The first RVQ quantizer is distilled from HuBERT features to produce semantic tokens capturing phonetic content, while later quantizers encode residual acoustics. The decoder is conditioned on semantic tokens via feature-wise linear modulation (FiLM), steering reconstruction toward information not explained by semantic abstraction. At 0.95 kbps, SemDAC matches or surpasses a 2.5 kbps DAC baseline on PESQ, STOI, SI-SNR, and Whisper WER, with comparable ViSQOL, and achieves higher subjective MOS than higher-bitrate DAC baselines. Results show that explicit semantic conditioning, rather than token disentanglement or increased model size alone, improves compression efficiency and recognition robustness.
Sources
- High Fidelity Neural Audio Compression
- AudioPaLM: A Large Language Model That Can Speak and Listen
- HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec
- SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models
- Decoupled Weight Decay Regularization
Related papers
- Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
- SoundWeaver: Compositional Warm-Starting for Text-to-Audio Diffusion Serving
- WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
- Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment