Pashto Common Voice: Building the First Open Speech Corpus for a 60-Million-Speaker Low-Resource Language

arXiv:2603.27021 · cs.CL · Submitted 2026-03-27 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Pashto Common Voice".

Jane: The Pashto Common Voice corpus represents the first large-scale, openly licensed speech resource for Pashto, addressing its absence in open speech technology despite having over 60 million native speakers.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: We’ve been looking at the details of the Pashto Common Voice: Building the First Open Speech Corpus for a sixty-Million-Speaker Low-Resource Language paper, and it really starts with who put this together <ref:2603.27021#pg0,Pashto Common Voice: Building the First Open Speech Corpus for a 60>. Jane The authors are Hanif Rahman, Shafeeq ur Rehman Pashto DAO, and they’re presenting this as their contribution to solving the problem of Pashto being missing from open speech technology despite having over sixty million native speakers <ref:2603.27021#pg0,Hanif Rahman, Shafeeq ur Rehman Pashto DAO>.

Lu: It’s interesting to see the focus on a language with such a massive speaker base when it's so far outside the mainstream open speech domain. Meng I wonder how they managed to get such diverse participation across ten different Mozilla Common Voice releases over those three years.

Lalam: The title itself emphasizes that this is the first large-scale, openly licensed speech resource for Pashto, which sets a pretty high bar for what’s expected of future language corpora. Tom That open licensing aspect is huge because it means anyone can use this data to build their own tools without needing special permissions.

Jane: It’s about making that data accessible to the community, and they are showing how a dedicated effort over several years can yield a significant result in addressing those kinds of gaps. Lu They’re essentially creating the first big open resource for this language, which is what makes this paper so important for anyone interested in low-resource language AI.

Tom: So, to summarize, they are highlighting the creation of a resource that directly addresses the lack of open speech tools for Pashto speakers. Jane That’s right; it sets up a new baseline for training ASR systems on this specific language.

Lu: The authors really laid out their plan for how they would go from an initial small collection to this large-scale resource using a systematic, multi-phase methodology. Meng I'm curious if that phased approach was strictly necessary, or if they could have jumped straight to larger collections.

Lalam: Their methodology involved interface localization, Wikipedia sentence extraction with automated filtering, and phonemically targeted contributions for the four most frequently dropped Pashto characters. Tom That sounds like a very comprehensive plan to ensure both volume and quality from the start.

Jane: It’s about being intentional about every step, ensuring that the data collected is not just a random collection of audio clips but something structured for actual use in ASR systems. Lu And that systematic approach is what allows them to track where their growth came from, which is really valuable for understanding how to build these things sustainably.

Meng: From an engineering viewpoint, tracking the progression across CV14 through CV23 shows a clear trajectory of effort over time, and that timeline gives us some good context on the human effort involved. Tom It helps us understand the human side of building a language resource for AI.

The paper's summary: Tom: Now we’re talking about what this Pashto Common Voice: Building the First Open Speech Corpus for a sixty-Million-Speaker Low-Resource Language paper actually says about their process <ref:2603.27021#pg0,Pashto Common Voice: Building the First Open Speech Corpus for a 60>. Jane Basically, they detail how they took a small starting point and systematically expanded it over three years into a resource with one hundred forty-seven total hours and one thousand four hundred eighty-three unique speakers <ref:2603.27021#pg0,147 total hours and 1,483 unique speakers>.

Lu: They describe the initial state as just one point five hours and five contributors before it grew to that final scale through various community initiatives <ref:2603.27021#pg0>. Meng It’s the growth trajectory itself that they emphasize, showing how participation can increase significantly with targeted campaigns, like that VOA Pashto broadcast campaign between CV17 and CV18.

Lalam: The core of their summary is the sustained community effort spanning from two thousand twenty-two to two thousand twenty-five which resulted in a corpus with one hundred forty-seven point zero seven hours total hours and sixty thousand three hundred thirty-seven validated clips by MCV23 <ref:2603.27021#pg0>. Tom That final number gives us a solid measure of what they managed to achieve by the end of their project.

Jane: They also broke down exactly how they curated the data, including using Wikipedia for sentence extraction with automated filtering to get about twenty-seven thousand sentences <ref:2603.27021#pg1>. Lu That shows they weren't relying on just one source; they were building a pipeline that pulled in different types of content.

Meng: And the focus on phonemically targeted contributions for those four most frequently dropped characters shows a deep understanding of the language’s unique orthographic challenges. Tom That kind of targeted effort is what really makes this corpus useful for specialized ASR work, especially when dealing with informal text.

Lalam: In short, the summary is about transforming scattered community contributions into a structured, validated resource ready for serious research applications in Pashto speech recognition.

Jane: It really boils down to showing how dedicated effort can create a valuable asset where none existed before for this specific language in the open speech technology landscape. Lu It’s a testament to what sustained, focused community work can accomplish when applied strategically.

Meng: And from an engineering standpoint, having the final statistics on validated hours and clips gives us tangible metrics to measure the success of their data collection pipeline. Tom Those metrics are essential for showing that the project was successful in its stated goal.

The paper's improvements: Tom: Let’s shift gears now to what the authors suggest as improvements or future directions for this Pashto Common Voice corpus, because they don't just stop at the final count. Jane They point out a few key areas where they think future efforts should focus, mostly concerning domain coverage and speaker demographics.

Lu: One major suggestion is that while Technology and News are currently dominant domains in MCV23, researchers should prioritize small-scale, targeted sentence collection for under-represented content domains like Religion and Biography. Tom That makes perfect sense; it helps create a more balanced model that isn't just skewed toward news and tech topics.

Lalam: They also noted a demographic gap: the speaker demographics showed a skew toward younger contributors, with thirty-seven percent being in their twenties, but there’s a critical metadata gap where gender is unreported for ninety-eight percent of speakers <ref:2603.27021#pg1>. Jane That lack of gender information could be an issue for future research if we want to build models that aren't biased based on speaker demographics.

Meng: From a practical standpoint, addressing that metadata gap means future collection efforts should incorporate systems to collect demographic info upfront, which is something we can definitely build into the pipeline. Tom It moves us from just collecting audio to collecting data ready for more nuanced ASR applications later on.

Lu: The suggestion about the knowledge broker—that intermediary role connecting technical work with the community—is something that could significantly accelerate future growth in order of magnitude. Jane If we can establish those partnerships early, it means we could see much faster corpus expansion than relying solely on purely technical pipelines.

Tom: It seems like they’re advocating for a more holistic strategy that combines technical refinement with strategic community engagement to keep the momentum going for this project. Lalam They are essentially saying that finding that person early in the process might be more impactful than just optimizing the pipeline itself.

Jane: That's a good summary; it’s about acknowledging the limitations of a static dataset and planning how to proactively address them moving forward with targeted efforts.

Conclusion: Tom: Alright, we’ve covered the scope of this paper on Pashto Common Voice: Building the First Open Speech Corpus for a sixty-Million-Speaker Low-Resource Language <ref:2603.27021#pg0,Pashto Common Voice: Building the First Open Speech Corpus for a 60>. We’ve seen how they built something substantial and what they suggest for future work. Jane They're wrapping up by summarizing that this corpus provides a critical, openly available resource for building consumer hardware speech technology on Pashto.

Lu: The main implication is establishing a new baseline performance metric for ASR systems on this language using the fine-tuned Whisper Base model results we discussed earlier in the show. Meng That thirteen point four percent WER compared to the zero-shot rate of ninety-nine point zero percent really shows a tangible improvement in what’s possible with current AI models when applied correctly <ref:2603.27021#pg0>.

Lalam: It underscores the importance of community engagement and targeted data curation as essential ingredients for successfully addressing low-resource languages in the open speech technology space. Tom It’s about showing that a dedicated effort can produce usable assets for real-world applications on consumer hardware.

Jane: Exactly; this paper gives us a concrete example of how to tackle the challenges inherent in building resources for languages with limited digital presence.

Lu: I think the structural opportunity they identified regarding knowledge brokers is perhaps the most interesting long-term concept we can take away from this research, as it suggests a way to scale up data acquisition efforts.

Meng: And from an engineering standpoint, integrating that broker framework into the pipeline could be a really smart move for future projects aiming for rapid growth.

Lalam: This entire effort highlights how crucial targeted data curation is when we are working on languages that don't have much existing digital infrastructure.

Pashto DAO

cs.CL

Submitted: 2026-03-27

Updated: 2026-03-27

Comments: Submitted to Interspeech 2026

Journal ref: Proc. Interspeech 2026, pp. 2879-2883

DOI: 10.21437/Interspeech.2026-1432

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 87/100

The gist: The Pashto Common Voice corpus represents the first large-scale, openly licensed speech resource for Pashto, addressing its absence in open speech technology despite having over 60 million native

Key concepts

Corpus Growth Trajectory
This tracks how the Pashto speech collection expanded over time. It started small but saw massive growth, especially after a VOA broadcast campaign. This shows that community engagement and external media can significantly boost volunteer participation in language data projects.
Dual-Review Gate
This is the quality control system used to validate audio clips. A clip needs two independent upvotes to be accepted and two downvotes to be rejected. This rigorous process ensures that the final dataset used for training is high-quality, though it results in a large volume of unreviewed material.
Downstream Impact on ASR
The corpus was used to improve speech recognition technology, specifically by fine-tuning Whisper Base. Achieving a 13.4% Word Error Rate (WER) on the test set shows that this large dataset enables advanced AI models to perform much better than zero-shot methods previously possible for Pashto.

Terminology

Summary

The Pashto Common Voice corpus represents the first large-scale, openly licensed speech resource for Pashto, addressing its absence in open speech technology despite having over 60 million native speakers. The corpus was built through a sustained community effort spanning from 2022 to 2025 and provides a new baseline for training ASR systems on this low-resource language.

Corpus Growth and Scale

The project grew from an initial state of 1.5 hours and 5 contributors to 147 total hours and 1,483 unique speakers across ten Mozilla Common Voice releases (CV14–CV23). This growth trajectory is detailed in Figure 1, which shows the progression of total recorded hours and unique speakers across the ten releases. A significant milestone was noted between CV17 and CV18, where Speaker participation increased approximately 108-fold, coinciding with a VOA Pashto broadcast campaign. By MCV23, the corpus reached 147.07 h total hours and included 60,337 validated clips across 13 content domains.

Methodology for Data Curation

The construction of the corpus involved seven distinct phases over three years to ensure quality and coverage. These phases included:

  1. Interface Localisation (2022–Feb. 2023), where the full web interface (over 1,200 strings) was translated into Pashto to enable speaker participation.

  2. Initial Sentence Corpus (Feb. 2023), curated via Mozilla’s Sentence Col1 and involving an initial set of 1,053 sentences.

  3. Wikipedia Sentence Extraction (2023), where a pipeline extracted approximately 27,000 sentences from Pashto Wikipedia, filtered by removing special characters and retaining only sentences of 15 words or fewer.

  4. Phonemically Targeted Contributions (2023–2024), specifically addressing the four fricative and affricate characters (shin, zhe, dze, tse) that are most frequently dropped in informal digital text.

  5. Sentence Expansion (2023–2024), which brought the validated sentence pool to 13 content domains by MCV23.

Quality Control and Characterisation

The quality assurance process utilized Mozilla Common Voice’s dual-review gate, requiring two independent upvotes to validate, two downvotes to reject. As of MCV23, this resulted in 60,337 clips (56.0%) being validated. The high fraction of unreviewed material (42,265 clips (39.2%)) was attributed to a volume bottleneck, rather than quality failure. Characterisation revealed that while Technology (8,609 clips) and News (5,264) dominate the domain distribution, areas like Religion and Biography are under-represented targets for future collection. Furthermore, the speaker demographics showed a skew toward younger contributors (37% in their twenties) with a critical metadata gap: Gender is unreported for 98% of speakers.

Downstream Impact on ASR

The corpus has established a new baseline for speech recognition performance. The authors fine-tuned Whisper Base on the MCV20 train split, achieving 13.4% WER on the MCV20 test split, which is a significant improvement over the published Whisper Base zero-shot WER of 99.0% on Pashto. This demonstrates that full fine-tuning rather than few-shot adaptation is now possible at this scale, enabling ASR work that was previously unfeasible. The corpus also provides sufficient audio for self-supervised pre-training frameworks such as wav2vec 2.0 for future research.

Key Findings and Transferable Insights

The study identifies several key findings regarding the process of building low-resource language resources:

The largest single growth event came from broadcast media.

That broadcast media via an engaged community intermediary can produce order-of-magnitude volunteer growth is the paper’s primary transferable finding.

A structural opportunity identified is the role of a knowledge broker, which acts as an intermediary connecting technical work to a community, suggesting that finding that person early may matter more than optimising the pipeline. Limitations noted include the corpus being primarily composed of prompted read speech and uncontrolled dialectal coverage.

Conclusion

The Pashto Common Voice corpus is openly available at commonvoice.mozilla.org/ps under a CC-0 licence, providing a critical resource for building open-source Pashto voice technology on consumer hardware. The creation of this corpus underscores the importance of community engagement and targeted data curation in addressing the needs of low-resource languages.

Improvements for AI systems

Here are the specific improvements to AI systems that can be made based on this research:

  1. The Whisper Base model, when fine-tuned on a Pashto speech corpus (MCV20), achieves a Word Error Rate (WER) of 13.4% on the test split, significantly outperforming its zero-shot performance of 99.0%.

  2. This fine-tuned system enables practical deployment for Pashto Automatic Speech Recognition (ASR) on consumer hardware like Apple MacBook Pros, making high-quality ASR accessible without massive computational resources typically required for large proprietary models.

  3. The development of a phonemically targeted sentence curation methodology allows researchers to systematically address the language's unique orthographic challenges (eight keyboard-absent characters).

  4. This technique provides a reliable training signal by ensuring that the ASR model is specifically exposed to data covering the Pashto fricatives and affricates—sounds most frequently dropped or misrepresented in informal digital text.

  5. The corpus characterization allows for targeted future data collection efforts, specifically addressing identified gaps:

Small-scale, targeted sentence collection can be prioritized for under-represented content domains (e.g., Religion, Biography) and demographic groups (e.g., gender-aware contributions), leading to more robust and unbiased ASR models in those specific areas.

  1. A knowledge broker framework can be integrated into the data acquisition pipeline to strategically partner with community intermediaries (journalists, media contacts). This allows for the identification of high-leverage, order-of-magnitude growth events that drive corpus expansion faster than purely technical efforts alone.

Abstract

We present the Pashto Common Voice corpus -- the first large-scale, openly licensed speech resource for Pashto, a language with over 60 million native speakers largely absent from open speech technology. Through a community effort spanning 2022-2025, the corpus grew from 1.5 hours and 5 contributors to 147 total hours and 1,483 unique speakers across ten Mozilla Common Voice releases (CV14-CV23). Speaker participation increased approximately 108-fold between CV17 and CV18, coinciding with a VOA Pashto broadcast campaign. We describe the full methodology: interface localisation, Wikipedia-based sentence extraction with automated filtering, phonemically targeted contributions for the four most frequently dropped Pashto characters, and multi-channel community outreach. MCV23 contains 107,781 clips (60,337 validated; 82.33 validated hours) across 13 content domains. Fine-tuning Whisper Base on the MCV20 yields 13.4% WER on the MCV20 test split, against the published Whisper Base zero-shot WER of 99.0% on Pashto.

Sources

Related papers