Sometin Beta Pass Notin: Improving Multilingual ASR for Nigerian Languages via Knowledge Distillation
cs.CL, eess.AS
Submitted: 2026-05-18
Updated: 2026-09-26
Comments: Accepted at Proc. SLT 2026, 7 pages
Code: https://github.com/snakers4/silero-vad
License: http://creativecommons.org/licenses/by/4.0/
The gist: Although modern multilingual Automatic Speech Recognition (ASR) systems support several Nigerian languages, their performance consistently lags behind resource-rich languages such as English and
Terminology
Abstract
Although modern multilingual Automatic Speech Recognition (ASR) systems support several Nigerian languages, their performance consistently lags behind resource-rich languages such as English and French. Nigerian languages present unique modelling hurdles, including acute data scarcity, inconsistent orthography, tonal diacritics, diverse accents, frequent code-switching, and localised named entities. To address these challenges, we developed a multilingual ASR framework using a two-stage distillation process. First, we employed student-teacher knowledge distillation from existing monolingual models, conditioned on robust language-specific N-gram language models. Second, we performed iterative self improvement using pseudo-labelled data to further refine accuracy. Our method significantly bridges the performance gap, achieving on average a reduction in the relative Word Error Rate (WER) of 29% over the monolingual baselines. Our models also outperform state-of-the-art multilingual models across major benchmarks, including Common Voice and FLEURS. We introduce Sometin Beta Pass Notin (SBPN), a multilingual foundational ASR model that covers Yorùbá, Hausa, Igbo, Nigerian Pidgin, and Nigerian English.
Sources
- Seamless: Multilingual Expressive and Streaming Speech Translation
- GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
- Distil-Whisper: Robust Knowledge Distillation via Large-Scale Pseudo Labelling
- Sequence Transduction with Recurrent Neural Networks
- NeMo: a toolkit for building AI applications using Neural Modules
- Towards End-to-End Training of Automatic Speech Recognition for Nigerian Pidgin
- Canary-1B-v2 & Parakeet-TDT-0.6B-v3: Efficient and High-Performance Models for Multilingual ASR and AST
- MUSAN: A Music, Speech, and Noise Corpus
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering