LoSATok: Low-Dimensional Semantic-Acoustic Tokenizer for Cross-Domain Audio Understanding and Generation
eess.AS, cs.AI, cs.SD
Submitted: 2026-05-27
Updated: 2026-09-12
Comments: Accepted to EMNLP 2026 Findings
Code: https://github.com/wxzyd123/LoSATok
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- On the Distillation Loss Functions of Speech VAE for Unified Reconstruction, Understanding, and Generation
- MiDashengLM: Efficient Audio Understanding with General Audio Captions
- DashengTokenizer: One layer is enough for unified audio understanding and generation
- LP-MusicCaps: LLM-Based Pseudo Music Captioning
- MusicLM: Generating Music From Text
- Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
- XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs
- SemanticVocoder: Bridging Audio Generation and Audio Understanding via Semantic Latents
- UniFlow-Audio: Unified Flow Matching for Audio Generation from Omni-Modalities
- Semantic-VAE: Semantic-Alignment Latent Representation for Better Speech Synthesis
- Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation
- WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling
- Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis
- LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech
- UniSRCodec: Unified and Low-Bitrate Single Codebook Codec with Sub-Band Reconstruction
Related papers
- X-VC: Zero-shot Streaming Voice Conversion in Codec Space
- Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
- Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
- Towards Audio Token Compression in Large Audio Language Models
- WaveScat: Wavelet Scattering Front-Ends with Self-Supervised Features for Speech Deepfake Detection
- ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions