Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech Recognition
summary
In short
The episode analyzes 'Comparative Analysis of Multilingual Pre-trained Models for Nepali ASR,' written by Suman Paudel and Sarbin Sayami. The hosts compare six models (like Whisper and IndicWav2Vec) to determine the best approach for Nepali Automatic Speech Recognition, concluding that linguistic proximity can substitute for raw model scale.
Key concepts
- Automatic Speech Recognition (ASR)
- The technology that converts spoken language into written text. The paper focuses on improving ASR specifically for the Nepali language, which is a low-resource language with limited available transcribed audio.
- Pre-trained Models
- Large AI models (like Whisper) that have already been trained on massive amounts of general audio data. These models are then 'fine-tuned' using limited Nepali data to adapt them for the specific language task.
- Real-Time Factor
- A measure of speed used in ASR deployment, indicating how long it takes to process one second of audio. A lower factor means the model is faster and better suited for live applications like voice assistants.
- Proximity vs. Scale
- The core finding of the paper: 'Proximity' refers to training on languages similar to Nepali, while 'Scale' refers to sheer data size. The hosts discuss how linguistic proximity can achieve performance comparable to massive scale.
Terminology used across episodes
This episode discusses
- Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech Recognition · Paper Radio
- Vakyansh: ASR Toolkit for Low Resource Indic languages
The paper
Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech Recognition · Read on arXiv
Suman Paudel, Sarbin Sayami
Tribhuvan University · Tribhuvan University
Multilingual pretrained models nominally support Nepali, yet no controlled benchmark has compared them under a single fine-tuning protocol. We fine-tune six pretrained models (XLSR-53, IndicWav2Vec, MMS-1B, Whisper-Medium, Whisper-Large-v3-Turbo, and Conformer-Hi) spanning CTC self-supervised, autoregressive encoder-decoder, and hybrid Conformer-CTC architectures, on the OpenSLR SLR54 Nepali corpus (165 hours) using identical preprocessing, splits, optimizer, and family-matched learning-rate schedules. We evaluate Word Error Rate (WER), Character Error Rate (CER), and Real-Time Factor (RTF) on three independent test sets (OpenSLR, FLEURS, Common Voice). Whisper-Large-v3-Turbo (14.76% WER) and IndicWav2Vec (14.89% WER) tie at the top despite a 9x parameter gap and 40x pretraining-data gap, providing direct empirical evidence that language-family proximity in pretraining can substitute for raw scale for in-domain Nepali. CTC decoders run up to 29x faster than autoregressive Whisper at the same accuracy, flipping the practical deployment preference toward CTC under any latency budget. Massively multilingual pretraining (MMS-1B) yields the smallest out-of-domain degradation on FLEURS (+12.55 pp), indicating that scale buys robustness rather than peak in-domain accuracy. The resulting benchmark provides the first standardized, multi-model, efficiency-aware reference numbers for Nepali ASR.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech Recognition".
Jane: The paper was written by Suman Paudel and Sarbin Sayami from Tribhuvan University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Welcome back, everyone. Today we're looking at a paper that's been making waves in the speech recognition world — it's called "Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech Recognition." Jane, when I first saw this title, I thought, okay, another benchmark paper. But this one's got some real teeth to it.
Jane: Oh, absolutely, Tom. And I think the title tells you exactly what you're getting. They took six different pre-trained models — these are models that have already learned to understand speech in general — and they fine-tuned them specifically for Nepali. That's the key word: comparative. Nobody had done this properly for Nepali before.
Tom: Right, and that's the part that got me excited. You've got models like Whisper from OpenAI, which is this massive system trained on hundreds of thousands of hours of audio. Then you've got IndicWav2Vec, which is much smaller but trained specifically on Indian languages. And the paper's asking a really simple question: does bigger always mean better?
Jane: And the answer, spoiler alert, is no. But let me back up for our listeners who might not be deep in the weeds here. Nepali is spoken by about thirty-two million people, but it has very little transcribed audio available for training — around one hundred sixty-five hours. Compare that to English, where you have hundreds of thousands of hours. So you can't just build a model from scratch. You have to start with something pre-trained and adapt it.
Tom: Exactly. And the authors, Suman Paudel and Sarbin Sayami from Tribhuvan University in Nepal, they did something really smart. They kept everything else the same — same data, same preprocessing, same training setup — so the only thing that varied was the model itself. That's how you get a fair comparison.
Jane: I love that discipline. Too often people compare models but change five other things at the same time, and then you don't know what actually made the difference. Here, they controlled for all of that. And what they found is that the model trained on Indic languages — IndicWav2Vec — matches the performance of Whisper's largest model, even though it's nine times smaller.
Tom: Nine times smaller. That's like comparing a compact car to a semi-truck and finding they both get you to work at the same time. That's going to matter a lot for people who actually want to deploy this technology in Nepal, where you might not have a massive server farm.
Jane: And that's the practical hook. But there's also a scientific hook here, which is about how we think about pre-training. The paper suggests that linguistic proximity — training on languages that are similar to your target — can substitute for raw scale. That's a big deal for the whole field of low-resource speech recognition.
Tom: Yeah, and we're going to dig into exactly how they set up this comparison and what the numbers actually show. Stick around, because the results might surprise you.
Summary: Tom: We're back with "Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech Recognition," and Jane, I want to get into the summary because this paper is dense with results.
Jane: It really is. Let me lay out the headline numbers. They fine-tuned six models on the OpenSLR SLR54 Nepali corpus, which is that one hundred sixty-five-hour dataset I mentioned. Then they tested on three different datasets: the OpenSLR test set, which is clean and in-domain; FLEURS, which is a curated multilingual benchmark; and Common Voice, which is crowd-sourced and messy.
Tom: And the results on the in-domain test set — that's OpenSLR — are fascinating. Whisper-Turbo gets a Word Error Rate of fourteen point seven six percent. IndicWav2Vec gets fourteen point eight nine percent. Those are essentially tied. But Whisper-Turbo has eight hundred nine million parameters, and IndicWav2Vec has ninety-four point four million. That's a nine-fold difference.
Jane: And here's the kicker — Whisper-Turbo was pre-trained on roughly six hundred eighty thousand hours of audio. IndicWav2Vec was pre-trained on about seventeen thousand hours. So you have a forty-fold difference in pre-training data, and they still tie. That's the proximity-over-scale finding, and it's the heart of this paper.
Tom: But it's not just about accuracy. They also measured speed. And this is where things get really interesting. They measured something called Real-Time Factor, which is basically how long it takes to process one second of audio. IndicWav2Vec runs at about zero point zero zero two six — that's roughly four hundred times faster than real-time. Whisper-Turbo runs at about zero point zero seven six — still fast, but about twenty-nine times slower.
Jane: So for the same accuracy, you get a twenty-nine-times speed difference. That changes the deployment calculus completely. If you're building a voice assistant for a Nepali-language app, you want that fast model. The paper even says this flips the practical preference toward CTC-style decoders over the autoregressive approach Whisper uses.
Tom: And then there's the generalization story. MMS-1B, which was trained on over one thousand one hundred languages, doesn't win on the in-domain test. It gets about twenty-seven percent WER, which is mid-tier. But when you move to out-of-domain data, it degrades the least. The gap from in-domain to FLEURS is only about twelve point five percentage points, compared to over thirty points for some other models.
Jane: That's the scale-buys-robustness finding. Big multilingual pre-training doesn't give you peak accuracy on clean speech, but it makes the model more resilient when the audio gets messy. That's a different kind of value.
Tom: So we've got three distinct takeaways: proximity beats scale for in-domain accuracy, CTC beats autoregressive for speed at equal accuracy, and massive multilingual pre-training buys robustness. That's a lot of practical guidance in one paper.
Jane: And we haven't even talked about the zero-shot results, which are honestly kind of brutal. Before fine-tuning, most of these models were useless on Nepali. Whisper was hallucinating Hindi and English. IndicWav2Vec had WERs over two hundred percent. Only MMS-1B produced anything usable, and only on FLEURS at about thirty-two percent WER. That's a strong argument that fine-tuning is non-negotiable.
Tom: Right, and that sets up the question of how they actually did the fine-tuning. Let's get into the methodology next.
Improvements: Tom: We're continuing with "Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech Recognition," and Jane, I want to talk about what this paper actually improves. Because it's not just a bunch of numbers — it's a blueprint.
Jane: Exactly. And I think the biggest improvement is standardization. Before this paper, if you wanted to build a Nepali ASR system, you'd have to dig through scattered papers, each using different datasets, different preprocessing, different evaluation metrics. You couldn't compare anything. This paper fixes that by putting six models through the exact same pipeline.
Tom: Same preprocessing, same optimizer, same batch size, same learning-rate schedules matched to each model family. That's the kind of rigor that lets you actually trust the comparisons.
Jane: And they also released everything — all six fine-tuned checkpoints and a per-utterance benchmark dataset on the Hugging Face Hub. That's huge for reproducibility. Anyone can download these models and verify the numbers or build on top of them.
Tom: But I think the deeper improvement is the conceptual one. The paper isolates two variables that usually get conflated: pretraining proximity and pretraining scale. Proximity means the languages the model saw during pre-training are close to your target language. Scale means how much total data and how many parameters. This paper shows they trade off — they don't compound.
Jane: That's a genuinely new contribution. And it has practical implications for how we think about building ASR for other low-resource languages. If you have a language that's related to a well-resourced language family, you might be better off with a small, family-specific model than a giant, generic one.
Tom: And then there's the efficiency angle. They're the first to publish Real-Time Factor numbers for Nepali ASR. That's not glamorous, but for anyone deploying a real product, it's essential. You need to know if your model can keep up with live speech.
Jane: Right. And the paper gives concrete deployment recommendations based on the measurements. If you need real-time performance on edge devices, use IndicWav2Vec. If you need cross-domain robustness on clean recordings, use Whisper-Turbo. If you're dealing with messy, out-of-domain audio, use MMS-1B.
Tom: That's the kind of actionable guidance that practitioners actually need. It's not just "here are some results" — it's "here's what to use and why."
Jane: And I think there's one more improvement worth mentioning. They compared against a prior published result from Ghimire et al., who reported a CER of six point seven seven percent for MMS-1B using active learning. This paper gets six point zero six percent CER without active learning, just with a clean fine-tuning protocol. That suggests a lot of the gains people attribute to clever data selection might actually come from just doing the basics right.
Tom: That's a subtle but important point. Sometimes the boring stuff — consistent preprocessing, proper hyperparameters — matters more than fancy tricks.
Jane: So the improvements here are methodological, practical, and conceptual. Now let's get into the first page of the paper and see how they set all of this up.
First Page: Tom: We're deep into "Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech Recognition" now, and Jane, I want to look at the first page because it sets up the whole problem so well.
Jane: It does. The abstract is really well-written. It frames the core question immediately: multilingual pre-trained models nominally support Nepali, but nobody had actually compared them under a single fine-tuning protocol. That's the gap they're filling.
Tom: And I love the opening framing — that self-supervised and weakly supervised pre-training have driven ASR on resource-rich languages to near-human accuracy, but the gains haven't propagated uniformly to the world's roughly seven thousand languages. That's the big-picture motivation.
Jane: Right. And they give you the linguistic context for Nepali specifically. It's an Indo-Aryan language with about thirty-two million native speakers. It uses Devanagari script. It has contrastive aspirated and unaspirated stops, conjunct-heavy orthography, agglutinative morphology, and free word order. That last one — free word order — is actually a real challenge for language models.
Tom: And then the data reality check. Only about one hundred sixty-five hours of openly licensed transcribed Nepali speech exists. That's an order of magnitude less than English. So you can't just train from scratch.
Jane: The authors also point out that previous Nepali ASR results came from disjoint single-model studies — each paper evaluated one model on one dataset with different preprocessing. So cross-model comparison was impossible. This paper is the first to fix that.
Tom: And the contributions list is worth reading. First, the first standardized multi-model, multi-dataset benchmark for Nepali ASR. Second, empirical isolation of pretraining proximity from pretraining scale. Third, first-of-kind Real-Time Factor measurements for Nepali. Fourth, per-scenario deployment recommendations. Fifth, public release of all checkpoints and the benchmark dataset.
Jane: That's a complete package. And I think the phrase that stuck with me from the abstract is "language-family proximity in pretraining can substitute for raw scale." That's the thesis statement of the whole paper.
Tom: It's a bold claim, and they back it up with data. But I also want to note the authors — Suman Paudel from the School of Mathematical Sciences and Sarbin Sayami from the Central Department of Computer Science and Information Technology, both at Tribhuvan University in Kathmandu. This is research done in Nepal, about Nepali, by people who understand the language and the context.
Jane: That matters. A lot of low-resource language research is done by outsiders who don't speak the language. Here, you have local researchers who can actually evaluate the qualitative outputs and understand the error patterns. That's a huge advantage.
Tom: And it's a model for how other under-resourced languages could be studied. You don't need a massive lab to do impactful work — you need a clear protocol and the right questions.
Jane: Absolutely. Now let's wrap this up and think about what this means for the broader world.
Conclusion: Tom: We're closing out our discussion of "Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech Recognition," and Jane, I think we should pull it all together.
Jane: Let's do it. The core finding is that for in-domain Nepali speech, a small model trained on Indic languages matches a massive model trained on hundreds of thousands of hours of audio. IndicWav2Vec and Whisper-Turbo tie at around fourteen point eight percent WER, but IndicWav2Vec is nine times smaller and twenty-nine times faster.
Tom: And that's not just a Nepali story. It's a story about how we should think about pre-training for any low-resource language. If you have a language family with decent resources, you might not need the biggest model on the shelf.
Jane: But the paper also shows that scale isn't useless — it buys robustness. MMS-1B degrades the least on out-of-domain data. So the choice depends on your use case. Clean studio recordings? Go small and fast. Noisy crowd-sourced audio? Go big and robust.
Tom: And the practical impact is real. For Nepal, this means voice assistants, transcription tools, and accessibility applications are now feasible with off-the-shelf hardware. You don't need a data center to run Nepali ASR.
Jane: The authors also released everything — models and benchmark data — so the next team can build on this without starting from scratch. That's how progress compounds.
Tom: And I think the broader implication is about who gets to do AI research. This paper came from a university in Kathmandu, not a tech giant. That's a signal that the field is opening up.
Jane: It really is. And we should also acknowledge the limitations they were honest about — single GPU training, no external language model during decoding, only read speech evaluated. There's room to grow.
Tom: But that's what makes it exciting. This is a foundation, not a finish line. The next step could be adding a Nepali language model for decoding, or continued pre-training on unlabeled Nepali audio.
Jane: So here's our send-off for this paper. It gave us a fair comparison, a clear recommendation, and a public resource. That's a rare combination.
Tom: Well said. We're saying goodbye to "Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech Recognition," and I'm genuinely excited to see what comes next in this space.
Jane: Same here. Thanks for listening, everyone. We'll be back with the next paper soon.
More episodes
- 2610.10857-Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization