Sometin Beta Pass Notin: Improving Multilingual ASR for Nigerian Languages via Knowledge Distillation

arXiv:2605.17710 · cs.CL, eess.AS · Submitted 2026-05-18 · Read on arXiv

cs.CL, eess.AS

Submitted: 2026-05-18

Updated: 2026-09-26

Comments: Accepted at Proc. SLT 2026, 7 pages

Code: https://github.com/snakers4/silero-vad

License: http://creativecommons.org/licenses/by/4.0/

The gist: Although modern multilingual Automatic Speech Recognition (ASR) systems support several Nigerian languages, their performance consistently lags behind resource-rich languages such as English and

Terminology

Abstract

Although modern multilingual Automatic Speech Recognition (ASR) systems support several Nigerian languages, their performance consistently lags behind resource-rich languages such as English and French. Nigerian languages present unique modelling hurdles, including acute data scarcity, inconsistent orthography, tonal diacritics, diverse accents, frequent code-switching, and localised named entities. To address these challenges, we developed a multilingual ASR framework using a two-stage distillation process. First, we employed student-teacher knowledge distillation from existing monolingual models, conditioned on robust language-specific N-gram language models. Second, we performed iterative self improvement using pseudo-labelled data to further refine accuracy. Our method significantly bridges the performance gap, achieving on average a reduction in the relative Word Error Rate (WER) of 29% over the monolingual baselines. Our models also outperform state-of-the-art multilingual models across major benchmarks, including Common Voice and FLEURS. We introduce Sometin Beta Pass Notin (SBPN), a multilingual foundational ASR model that covers Yorùbá, Hausa, Igbo, Nigerian Pidgin, and Nigerian English.

Sources

Related papers