Code-Switching Spoken Language Identification as Multi-Label Set Prediction
eess.AS, cs.CL
Submitted: 2026-10-01
Updated: 2026-10-01
Code: https://github.com/espnet/espnet
Terminology
Sources
- CS-YODAS: A Mined Dataset of In-the-Wild Code-Switched Speech
- Adversarial synthesis based data-augmentation for code-switched spoken language identification
- XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Related papers
- X-VC: Zero-shot Streaming Voice Conversion in Codec Space
- Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
- Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
- Towards Audio Token Compression in Large Audio Language Models
- WaveScat: Wavelet Scattering Front-Ends with Self-Supervised Features for Speech Deepfake Detection
- ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions