Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech Recognition

arXiv:2608.12327 · cs.CL, cs.AI, cs.LG · Submitted 2026-05-31 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech Recognition".

Jane: The paper was written by Suman Paudel and Sarbin Sayami from Tribhuvan University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Welcome back, everyone. Today we're looking at a paper that's been making waves in the speech recognition world — it's called "Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech Recognition." Jane, when I first saw this title, I thought, okay, another benchmark paper. But this one's got some real teeth to it.

Jane: Oh, absolutely, Tom. And I think the title tells you exactly what you're getting. They took six different pre-trained models — these are models that have already learned to understand speech in general — and they fine-tuned them specifically for Nepali. That's the key word: comparative. Nobody had done this properly for Nepali before.

Tom: Right, and that's the part that got me excited. You've got models like Whisper from OpenAI, which is this massive system trained on hundreds of thousands of hours of audio. Then you've got IndicWav2Vec, which is much smaller but trained specifically on Indian languages. And the paper's asking a really simple question: does bigger always mean better?

Jane: And the answer, spoiler alert, is no. But let me back up for our listeners who might not be deep in the weeds here. Nepali is spoken by about thirty-two million people, but it has very little transcribed audio available for training — around one hundred sixty-five hours. Compare that to English, where you have hundreds of thousands of hours. So you can't just build a model from scratch. You have to start with something pre-trained and adapt it.

Tom: Exactly. And the authors, Suman Paudel and Sarbin Sayami from Tribhuvan University in Nepal, they did something really smart. They kept everything else the same — same data, same preprocessing, same training setup — so the only thing that varied was the model itself. That's how you get a fair comparison.

Jane: I love that discipline. Too often people compare models but change five other things at the same time, and then you don't know what actually made the difference. Here, they controlled for all of that. And what they found is that the model trained on Indic languages — IndicWav2Vec — matches the performance of Whisper's largest model, even though it's nine times smaller.

Tom: Nine times smaller. That's like comparing a compact car to a semi-truck and finding they both get you to work at the same time. That's going to matter a lot for people who actually want to deploy this technology in Nepal, where you might not have a massive server farm.

Jane: And that's the practical hook. But there's also a scientific hook here, which is about how we think about pre-training. The paper suggests that linguistic proximity — training on languages that are similar to your target — can substitute for raw scale. That's a big deal for the whole field of low-resource speech recognition.

Tom: Yeah, and we're going to dig into exactly how they set up this comparison and what the numbers actually show. Stick around, because the results might surprise you.

Summary: Tom: We're back with "Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech Recognition," and Jane, I want to get into the summary because this paper is dense with results.

Jane: It really is. Let me lay out the headline numbers. They fine-tuned six models on the OpenSLR SLR54 Nepali corpus, which is that one hundred sixty-five-hour dataset I mentioned. Then they tested on three different datasets: the OpenSLR test set, which is clean and in-domain; FLEURS, which is a curated multilingual benchmark; and Common Voice, which is crowd-sourced and messy.

Tom: And the results on the in-domain test set — that's OpenSLR — are fascinating. Whisper-Turbo gets a Word Error Rate of fourteen point seven six percent. IndicWav2Vec gets fourteen point eight nine percent. Those are essentially tied. But Whisper-Turbo has eight hundred nine million parameters, and IndicWav2Vec has ninety-four point four million. That's a nine-fold difference.

Jane: And here's the kicker — Whisper-Turbo was pre-trained on roughly six hundred eighty thousand hours of audio. IndicWav2Vec was pre-trained on about seventeen thousand hours. So you have a forty-fold difference in pre-training data, and they still tie. That's the proximity-over-scale finding, and it's the heart of this paper.

Tom: But it's not just about accuracy. They also measured speed. And this is where things get really interesting. They measured something called Real-Time Factor, which is basically how long it takes to process one second of audio. IndicWav2Vec runs at about zero point zero zero two six — that's roughly four hundred times faster than real-time. Whisper-Turbo runs at about zero point zero seven six — still fast, but about twenty-nine times slower.

Jane: So for the same accuracy, you get a twenty-nine-times speed difference. That changes the deployment calculus completely. If you're building a voice assistant for a Nepali-language app, you want that fast model. The paper even says this flips the practical preference toward CTC-style decoders over the autoregressive approach Whisper uses.

Tom: And then there's the generalization story. MMS-1B, which was trained on over one thousand one hundred languages, doesn't win on the in-domain test. It gets about twenty-seven percent WER, which is mid-tier. But when you move to out-of-domain data, it degrades the least. The gap from in-domain to FLEURS is only about twelve point five percentage points, compared to over thirty points for some other models.

Jane: That's the scale-buys-robustness finding. Big multilingual pre-training doesn't give you peak accuracy on clean speech, but it makes the model more resilient when the audio gets messy. That's a different kind of value.

Tom: So we've got three distinct takeaways: proximity beats scale for in-domain accuracy, CTC beats autoregressive for speed at equal accuracy, and massive multilingual pre-training buys robustness. That's a lot of practical guidance in one paper.

Jane: And we haven't even talked about the zero-shot results, which are honestly kind of brutal. Before fine-tuning, most of these models were useless on Nepali. Whisper was hallucinating Hindi and English. IndicWav2Vec had WERs over two hundred percent. Only MMS-1B produced anything usable, and only on FLEURS at about thirty-two percent WER. That's a strong argument that fine-tuning is non-negotiable.

Tom: Right, and that sets up the question of how they actually did the fine-tuning. Let's get into the methodology next.

Improvements: Tom: We're continuing with "Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech Recognition," and Jane, I want to talk about what this paper actually improves. Because it's not just a bunch of numbers — it's a blueprint.

Jane: Exactly. And I think the biggest improvement is standardization. Before this paper, if you wanted to build a Nepali ASR system, you'd have to dig through scattered papers, each using different datasets, different preprocessing, different evaluation metrics. You couldn't compare anything. This paper fixes that by putting six models through the exact same pipeline.

Tom: Same preprocessing, same optimizer, same batch size, same learning-rate schedules matched to each model family. That's the kind of rigor that lets you actually trust the comparisons.

Jane: And they also released everything — all six fine-tuned checkpoints and a per-utterance benchmark dataset on the Hugging Face Hub. That's huge for reproducibility. Anyone can download these models and verify the numbers or build on top of them.

Tom: But I think the deeper improvement is the conceptual one. The paper isolates two variables that usually get conflated: pretraining proximity and pretraining scale. Proximity means the languages the model saw during pre-training are close to your target language. Scale means how much total data and how many parameters. This paper shows they trade off — they don't compound.

Jane: That's a genuinely new contribution. And it has practical implications for how we think about building ASR for other low-resource languages. If you have a language that's related to a well-resourced language family, you might be better off with a small, family-specific model than a giant, generic one.

Tom: And then there's the efficiency angle. They're the first to publish Real-Time Factor numbers for Nepali ASR. That's not glamorous, but for anyone deploying a real product, it's essential. You need to know if your model can keep up with live speech.

Jane: Right. And the paper gives concrete deployment recommendations based on the measurements. If you need real-time performance on edge devices, use IndicWav2Vec. If you need cross-domain robustness on clean recordings, use Whisper-Turbo. If you're dealing with messy, out-of-domain audio, use MMS-1B.

Tom: That's the kind of actionable guidance that practitioners actually need. It's not just "here are some results" — it's "here's what to use and why."

Jane: And I think there's one more improvement worth mentioning. They compared against a prior published result from Ghimire et al., who reported a CER of six point seven seven percent for MMS-1B using active learning. This paper gets six point zero six percent CER without active learning, just with a clean fine-tuning protocol. That suggests a lot of the gains people attribute to clever data selection might actually come from just doing the basics right.

Tom: That's a subtle but important point. Sometimes the boring stuff — consistent preprocessing, proper hyperparameters — matters more than fancy tricks.

Jane: So the improvements here are methodological, practical, and conceptual. Now let's get into the first page of the paper and see how they set all of this up.

First Page: Tom: We're deep into "Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech Recognition" now, and Jane, I want to look at the first page because it sets up the whole problem so well.

Jane: It does. The abstract is really well-written. It frames the core question immediately: multilingual pre-trained models nominally support Nepali, but nobody had actually compared them under a single fine-tuning protocol. That's the gap they're filling.

Tom: And I love the opening framing — that self-supervised and weakly supervised pre-training have driven ASR on resource-rich languages to near-human accuracy, but the gains haven't propagated uniformly to the world's roughly seven thousand languages. That's the big-picture motivation.

Jane: Right. And they give you the linguistic context for Nepali specifically. It's an Indo-Aryan language with about thirty-two million native speakers. It uses Devanagari script. It has contrastive aspirated and unaspirated stops, conjunct-heavy orthography, agglutinative morphology, and free word order. That last one — free word order — is actually a real challenge for language models.

Tom: And then the data reality check. Only about one hundred sixty-five hours of openly licensed transcribed Nepali speech exists. That's an order of magnitude less than English. So you can't just train from scratch.

Jane: The authors also point out that previous Nepali ASR results came from disjoint single-model studies — each paper evaluated one model on one dataset with different preprocessing. So cross-model comparison was impossible. This paper is the first to fix that.

Tom: And the contributions list is worth reading. First, the first standardized multi-model, multi-dataset benchmark for Nepali ASR. Second, empirical isolation of pretraining proximity from pretraining scale. Third, first-of-kind Real-Time Factor measurements for Nepali. Fourth, per-scenario deployment recommendations. Fifth, public release of all checkpoints and the benchmark dataset.

Jane: That's a complete package. And I think the phrase that stuck with me from the abstract is "language-family proximity in pretraining can substitute for raw scale." That's the thesis statement of the whole paper.

Tom: It's a bold claim, and they back it up with data. But I also want to note the authors — Suman Paudel from the School of Mathematical Sciences and Sarbin Sayami from the Central Department of Computer Science and Information Technology, both at Tribhuvan University in Kathmandu. This is research done in Nepal, about Nepali, by people who understand the language and the context.

Jane: That matters. A lot of low-resource language research is done by outsiders who don't speak the language. Here, you have local researchers who can actually evaluate the qualitative outputs and understand the error patterns. That's a huge advantage.

Tom: And it's a model for how other under-resourced languages could be studied. You don't need a massive lab to do impactful work — you need a clear protocol and the right questions.

Jane: Absolutely. Now let's wrap this up and think about what this means for the broader world.

Conclusion: Tom: We're closing out our discussion of "Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech Recognition," and Jane, I think we should pull it all together.

Jane: Let's do it. The core finding is that for in-domain Nepali speech, a small model trained on Indic languages matches a massive model trained on hundreds of thousands of hours of audio. IndicWav2Vec and Whisper-Turbo tie at around fourteen point eight percent WER, but IndicWav2Vec is nine times smaller and twenty-nine times faster.

Tom: And that's not just a Nepali story. It's a story about how we should think about pre-training for any low-resource language. If you have a language family with decent resources, you might not need the biggest model on the shelf.

Jane: But the paper also shows that scale isn't useless — it buys robustness. MMS-1B degrades the least on out-of-domain data. So the choice depends on your use case. Clean studio recordings? Go small and fast. Noisy crowd-sourced audio? Go big and robust.

Tom: And the practical impact is real. For Nepal, this means voice assistants, transcription tools, and accessibility applications are now feasible with off-the-shelf hardware. You don't need a data center to run Nepali ASR.

Jane: The authors also released everything — models and benchmark data — so the next team can build on this without starting from scratch. That's how progress compounds.

Tom: And I think the broader implication is about who gets to do AI research. This paper came from a university in Kathmandu, not a tech giant. That's a signal that the field is opening up.

Jane: It really is. And we should also acknowledge the limitations they were honest about — single GPU training, no external language model during decoding, only read speech evaluated. There's room to grow.

Tom: But that's what makes it exciting. This is a foundation, not a finish line. The next step could be adding a Nepali language model for decoding, or continued pre-training on unlabeled Nepali audio.

Jane: So here's our send-off for this paper. It gave us a fair comparison, a clear recommendation, and a public resource. That's a rare combination.

Tom: Well said. We're saying goodbye to "Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech Recognition," and I'm genuinely excited to see what comes next in this space.

Jane: Same here. Thanks for listening, everyone. We'll be back with the next paper soon.

Suman Paudel, Sarbin Sayami

Tribhuvan University · Tribhuvan University

cs.CL, cs.AI, cs.LG

Submitted: 2026-05-31

Updated: 2026-08-14

Comments: 9 pages, 6 figures, 7 tables. Based on M.Sc. thesis (Institute of Science and Technology, Tribhuvan University). Code and models: https://github.com/p-sumann/nepali-asr-benchmark

Code: https://github.com/p-sumann/nepali-asr-benchmark

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

Importance score: 69/100

Key concepts

Automatic Speech Recognition (ASR)
The technology that converts spoken language into written text. The paper focuses on improving ASR specifically for the Nepali language, which is a low-resource language with limited available transcribed audio.
Pre-trained Models
Large AI models (like Whisper) that have already been trained on massive amounts of general audio data. These models are then 'fine-tuned' using limited Nepali data to adapt them for the specific language task.
Real-Time Factor
A measure of speed used in ASR deployment, indicating how long it takes to process one second of audio. A lower factor means the model is faster and better suited for live applications like voice assistants.
Proximity vs. Scale
The core finding of the paper: 'Proximity' refers to training on languages similar to Nepali, while 'Scale' refers to sheer data size. The hosts discuss how linguistic proximity can achieve performance comparable to massive scale.

Terminology

Summary

Summary

This paper presents the first standardized, multi-model, multi-dataset benchmark for Nepali Automatic Speech Recognition (ASR). The authors fine-tune six multilingual pretrained models from three architectural families—CTC self-supervised (XLSR-53, IndicWav2Vec, MMS-1B), autoregressive encoder–decoder (Whisper-Medium, Whisper-Large-v3-Turbo), and hybrid Conformer-CTC (Conformer-Hi)—on the OpenSLR SLR54 Nepali corpus (165 hours) using identical preprocessing, splits, optimizer, and family-matched learning-rate schedules. Models are evaluated on three independent test sets (OpenSLR, FLEURS, Common Voice) along three orthogonal axes: accuracy (WER, CER), inference efficiency (RTF), and Nepali-specific error patterns.

Key results. Whisper-Large-v3-Turbo (14.76% WER) and IndicWav2Vec (14.89% WER) tie at the top on OpenSLR despite a 9× parameter gap and 40× pretraining-data gap, providing direct empirical evidence that language-family proximity in pretraining can substitute for raw scale for in-domain Nepali. CTC decoders run up to 29× faster than autoregressive Whisper at the same accuracy, flipping the practical deployment preference toward CTC under any latency budget. Massively multilingual pretraining (MMS-1B) yields the smallest out-of-domain degradation on FLEURS (+12.55 pp), indicating that scale buys robustness rather than peak in-domain accuracy.

Zero-shot performance. Only MMS-1B produced independently usable output, and only on FLEURS (31.75% WER). Whisper models predominantly hallucinated Hindi or English. IndicWav2Vec WER > 200% reflects an uninitialised Nepali CTC head. These results justify fine-tuning as essential for every model.

Fine-tuned in-domain results. The top three models (Whisper-Turbo, IndicWav2Vec, Whisper-Medium) fall within roughly one percentage point of each other on validation WER despite spanning a 9× range in parameter count and using fundamentally different decoders. IndicWav2Vec reaches its best WER with 94.4 M parameters in under six hours of training, whereas Whisper-Turbo requires 809 M parameters and approximately 34 hours. Conformer-Hi (30.5 M, Hindi-only pretraining) outperforms both XLSR-53 (317 M) and MMS-1B (965 M), reinforcing that linguistic proximity in pretraining matters more than raw model capacity.

Multi-test-set benchmark. On the fine-tuned benchmarks, Whisper-Turbo achieves the lowest in-domain WER (14.76%), with IndicWav2Vec (14.89%) and Whisper-Medium (15.57%) within a single percentage point. Conformer-Hi, XLSR-53, and MMS-1B form a middle tier around 26–27% WER. On FLEURS, Whisper-Medium leads narrowly (39.06% WER), but MMS-1B is the only model with <11% CER on FLEURS. On Common Voice, every model degrades sharply; even the best model crosses 48% WER, establishing crowd-sourced Nepali audio as the largest remaining open problem.

Inference efficiency. All six models operate well below the real-time threshold. IndicWav2Vec averages RTF ≈0.0026, approximately 400× real-time; Conformer-Hi is marginally faster owing to its smaller parameter count. CTC decoding is structurally cheaper than autoregressive decoding because it requires a single encoder forward pass rather than per-token generation. Whisper-Turbo, despite matching IndicWav2Vec on accuracy, is approximately 29× slower (0.076 average RTF vs. 0.0026), which is decisive for any deployment with strict latency budgets.

Generalization. The WER increase from in-domain (OpenSLR) to out-of-domain test sets exposes the generalization gap directly. MMS-1B shows the smallest FLEURS gap (+12.55 pp), roughly half of the next-best model. The plausible explanation: MMS-1B’s >1,100-language pretraining exposes it to acoustic conditions broader than any single in-domain corpus, so the in-domain to out-of-domain shift is closer to in-distribution from its perspective. Whisper-Turbo is the most consistent performer across all three test sets among the high-accuracy models.

Training dynamics. Most models converge within the first one to three epochs; early stopping triggered on Whisper-Turbo (epoch 3), Whisper-Medium (epoch 6), and MMS-1B (epoch 4). The smooth monotone descent on every validation-WER panel indicates no training instability and no overfitting within the controlled budget.

Comparison with prior published Nepali results. Ghimire et al. (2023) reported MMS-1B Nepali CER of 6.77% with active-learning-based data selection on an in-house Nepali set. The MMS-1B CER of 6.06% reported here on the OpenSLR test partition is lower despite no active-learning intervention, suggesting that the controlled fine-tuning protocol used in this study is at least competitive with active-learning data selection. Pratap et al. (2024) and Javed et al. (2022) report only aggregated Indic numbers; the per-language Nepali measurements here fill that gap.

Practical recommendations. Three deployment-time recommendations follow directly from the measurements. IndicWav2Vec is the preferred choice for real-time and edge use where its 94.4 M parameter count and fast CTC decoder are decisive. Whisper-Turbo is preferred when cross-domain robustness on cleanly recorded speech matters more than latency. MMS-1B is preferred when out-of-domain generalization, rather than peak in-domain accuracy, is the priority. All six fine-tuned checkpoints are released on the Hugging Face Hub for direct download or programmatic loading.

Limitations. Experiments were carried out on a single NVIDIA L4 GPU (24 GB VRAM), with select models replicated on an A100 80 GB instance for time-bound runs. Hardware constraints limited per-device batch sizes and, for the largest models, the number of training epochs. Conformer-Hi required a separate training pipeline (NVIDIA NeMo) which, while configured to match the shared preprocessing protocol, is not bit-for-bit identical to the Hugging Face fine-tuning pipeline used for the other five models. All evaluation data is read speech from curated corpora; performance on spontaneous conversational speech, code-switched (Nepali–English/Hindi) speech, and dialectal variation remains outside this study’s scope. No external language model was used during decoding, so reported numbers reflect purely acoustic-model performance and may understate what shallow-fusion or rescoring approaches could achieve. Publicly released pretrained checkpoints are used as-is; continued self-supervised pretraining on unlabelled Nepali audio was not attempted and could plausibly close the residual gap to the strongest models.

Conclusion. The two top models (Whisper-Turbo, IndicWav2Vec) tie within 0.13 pp despite a 9× parameter gap, providing direct evidence that language-family proximity in pretraining can substitute for raw scale on in-domain Nepali. CTC decoding is up to 29× faster than autoregressive Whisper at equivalent accuracy, flipping the practical deployment preference toward CTC under any latency budget. Massively multilingual pretraining (MMS-1B) buys out-of-domain robustness rather than peak accuracy. The resulting benchmark and per-scenario recommendations supply the empirically grounded reference numbers that have been missing from the Nepali ASR literature.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement and what the improved AI system can do:

Improvement: Implement a model-selection algorithm that prioritizes pretraining language-family proximity over raw parameter count or total pretraining data volume. The system will automatically rank candidate ASR models by (a) whether the target language's family is represented in pretraining, (b) the proportion of pretraining data from that family, and (c) architectural compatibility.

Capability: The system can now select the optimal pretrained model for any low-resource language (e.g., Nepali) without expensive fine-tuning trials. It will correctly prefer a 94.4M-parameter Indic-proximate model over a 809M-parameter broadly multilingual model when in-domain accuracy is the goal, saving up to 9× compute and 5.7× training time.

Improvement: Build a routing layer that automatically selects between CTC and autoregressive decoders based on the deployment constraint (latency vs. robustness). The router uses the measured 29× RTF gap to make decisions: if the latency budget is <0.01× real-time, route to CTC; if cross-domain robustness on clean audio is prioritized, route to Whisper-family.

Improvement: Implement a two-phase fine-tuning protocol that first fine-tunes on in-domain data (e.g., OpenSLR) and then applies a small amount of fine-tuning on diverse out-of-domain samples (e.g., FLEURS, Common Voice) to compress the generalization gap. The system will monitor the WER gap (e.g., +12.55 pp for MMS-1B vs. +30.93 pp for XLSR-53) and stop when the gap stabilizes.

Improvement: Add a pre-fine-tuning assessment module that predicts whether a given pretrained model will produce usable zero-shot output for a target language. The predictor uses the paper's finding that only MMS-1B produced independently usable Nepali output (31.75% WER on FLEURS) while Whisper hallucinated Hindi/English and IndicWav2Vec produced >200% WER due to uninitialized heads.

Improvement: Implement a predictive model that estimates total fine-tuning time and early-stopping epoch based on model architecture, parameter count, and pretraining proximity. Using the paper's measured durations (IndicWav2Vec: 5.9h, Whisper-Turbo: 33.9h, MMS-1B: 19.2h), the system will schedule GPU resources optimally.

Improvement: Add a post-processing layer that applies Nepali-specific orthographic corrections based on the paper's error patterns (e.g., काने vs. खाने, दे रै vs. धे रै, टाउँ vs. ठाउँ). The system will learn confusion pairs from the released per-utterance benchmark data and apply character-level corrections.

Improvement: Build an automated evaluation pipeline that always tests on three independent sets (in-domain, curated out-of-domain, crowd-sourced out-of-domain) and reports WER, CER, and RTF together. The system will flag models that show >25 pp generalization gaps as not deployment-ready.

  1. Deploy Nepali ASR in production with a 94.4M-parameter model running at 400× real-time with 14.89% WER, suitable for edge devices and real-time applications.

  2. Select the right model for any low-resource language in under 10 minutes using the proximity-vs-scale predictor, avoiding weeks of trial-and-error fine-tuning.

  3. Serve mixed workloads where some requests need sub-10ms latency (CTC) and others need maximum accuracy (autoregressive), automatically routing each request.

  4. Maintain accuracy on noisy, crowd-sourced audio by combining MMS-1B's robustness (+12.55 pp gap) with error-pattern post-processing, reducing Common Voice WER from 58.65% toward 50%.

  5. Benchmark any new ASR model against six baselines across three test sets with standardized metrics, producing deployment-ready recommendations in a single run.

  6. Predict training costs before starting, enabling optimal GPU scheduling and budget allocation across multiple candidate models.

Abstract

Multilingual pretrained models nominally support Nepali, yet no controlled benchmark has compared them under a single fine-tuning protocol. We fine-tune six pretrained models (XLSR-53, IndicWav2Vec, MMS-1B, Whisper-Medium, Whisper-Large-v3-Turbo, and Conformer-Hi) spanning CTC self-supervised, autoregressive encoder-decoder, and hybrid Conformer-CTC architectures, on the OpenSLR SLR54 Nepali corpus (165 hours) using identical preprocessing, splits, optimizer, and family-matched learning-rate schedules. We evaluate Word Error Rate (WER), Character Error Rate (CER), and Real-Time Factor (RTF) on three independent test sets (OpenSLR, FLEURS, Common Voice). Whisper-Large-v3-Turbo (14.76% WER) and IndicWav2Vec (14.89% WER) tie at the top despite a 9x parameter gap and 40x pretraining-data gap, providing direct empirical evidence that language-family proximity in pretraining can substitute for raw scale for in-domain Nepali. CTC decoders run up to 29x faster than autoregressive Whisper at the same accuracy, flipping the practical deployment preference toward CTC under any latency budget. Massively multilingual pretraining (MMS-1B) yields the smallest out-of-domain degradation on FLEURS (+12.55 pp), indicating that scale buys robustness rather than peak in-domain accuracy. The resulting benchmark provides the first standardized, multi-model, efficiency-aware reference numbers for Nepali ASR.

Sources

Related papers