Long-Tail Rebalancing for Non-Verbal Vocalization-Aware ASR: A Track 1 System for the NVVSpeech Challenge
eess.AS, cs.CL, cs.SD
Submitted: 2026-09-20
Updated: 2026-09-23
Comments: Accepted by ISCSLP 2026, NVVSpeech Challenge Track 1
Project page: https://nvvspeech-challenge.github.io
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Beyond Words: Towards Effective Modeling of Non-Verbal Vocalizations in ASR
- Decoupling Representation and Classifier for Long-Tailed Recognition
- Preference Optimization for Non-Verbal Vocalization Synthesis
- Balanced Meta-Softmax for Long-Tailed Visual Recognition
- NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations
- A Scalable Pipeline for Enabling Non-Verbal Speech Generation and Understanding
- WESR: Scaling and Evaluating Word-level Event-Speech Recognition
- Qwen3-ASR Technical Report
Related papers
- X-VC: Zero-shot Streaming Voice Conversion in Codec Space
- Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
- Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
- Towards Audio Token Compression in Large Audio Language Models
- WaveScat: Wavelet Scattering Front-Ends with Self-Supervised Features for Speech Deepfake Detection
- ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions