DonorRank: Donor Language Selection for Low-Resource Cross-Lingual Speech Recognition

arXiv:2608.11441 · cs.CL · Submitted 2026-08-11 · Read on arXiv

Akriti Dhasmana, Aarohi Srivastava, David Chiang

University of Notre Dame

cs.CL

Submitted: 2026-08-11

Updated: 2026-08-13

Comments: 11 pages, 4 figures, 12 tables

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: DonorRank: Donor Language Selection for Low-Resource Cross-Lingual Speech Recognition Abstract Low-resource automatic speech recognition (ASR) commonly relies on cross-lingual transfer, where models

Terminology

Summary

DonorRank: Donor Language Selection for Low-Resource Cross-Lingual Speech Recognition

Abstract

Low-resource automatic speech recognition (ASR) commonly relies on cross-lingual transfer, where models are adapted from higher-resource donor languages. However, selecting donors remains challenging for spontaneous speech from under-resourced language communities, due to linguistic variation, evolving orthographic conventions, and uneven resource availability. We present DonorRank, a learning-to-rank framework for predicting effective donor languages for zero-shot ASR. We evaluate DonorRank on two multilingual speech corpora of Indic and African language families. It accurately predicts donor language rankings and improves donor selection over common heuristics based on genetic similarity or high-resource languages. Beyond improving transfer, we show how DonorRank is a general framework for analyzing donor language selection itself. Our analyses show that the composition of the donor set determines which linguistic cues are useful in predicting successful transfer. We also identify transfer patterns that provide practical guidance for multilingual ASR in low-resource settings.

Introduction

Building automatic speech recognition (ASR) systems for the world’s thousands of languages is challenging since most languages lack sufficient transcribed speech for supervised training. Cross-lingual transfer has thus become a common strategy for low-resource ASR, where a model is fine-tuned on one or more higher-resource donor languages before being applied to an unseen target language. However, the transfer performance in this approach depends critically on the choice of donor languages, and can vary substantially even among closely related languages. Therefore, what makes a donor language effective is an important question for developing robust and inclusive multilingual speech technologies.

The paper makes the following contributions:

  1. It presents DonorRank, a learning-based framework for donor language selection in zero-shot ASR and evaluates it in two complementary multilingual settings.

  2. It provides an empirical analysis of the features that drive donor language selection, demonstrating that linguistic and dataset characteristics improve donor rankings over those based solely on genealogical relationships.

  3. It identifies recurring transfer patterns, including effective donor hubs and nonobvious donor languages, providing practical guidance for donor language selection in low-resource ASR scenarios where empirical evidence is limited.

Data and Languages

The paper evaluates donor language selection on two spontaneous speech corpora representing contrasting multilingual transfer settings: VAANI-D, a Devanagari-script subset of the VAANI corpus, and WAXAL. VAANI-D consists of closely related Indo-Aryan language varieties spoken across northern and central India, providing a linguistically controlled setting for studying transfer within a dense network of related languages. In contrast, WAXAL spans languages from multiple African language families spoken across Sub-Saharan Africa and written in multiple scripts, providing a complementary setting for studying donor selection across substantially greater typological diversity.

VAANI-D is a subset of the VAANI corpus, which contains over 150,000 hours of speech, approximately 10% of which is transcribed, from more than 156,000 speakers across 773 districts of India. The subset comprises 20 Indic languages and language varieties, including both higher- and lower-resource varieties. Among these, Awadhi (awa), Bhili (bhb), Garhwali (gbm), Halbi (hlb), Konkani (kok), and Marwari (mwr) serve as the primary low-resource target languages for evaluating zero-shot transfer. Restricting analysis to a common writing system provides a linguistically controlled setting in which differences in transfer performance are less likely to arise from script mismatch and instead reflect variation among closely related language varieties. The corpus has spontaneous code-mixed speech, and several language varieties exhibit limited orthographic standardization despite sharing the Devanagari script.

WAXAL contains approximately 1,250 hours of transcribed spontaneous speech across 19 African languages spoken by more than 100 million speakers. These languages are typologically diverse, spanning multiple language families (e.g., Bantu, Semitic) and written scripts (e.g., Latin, Ge’ez). Many languages in WAXAL are represented using practical Romanization rather than long-established written conventions, reflecting another common characteristic of low-resource ASR.

For both datasets, the paper fine-tunes the ASR model using the training split of a single donor language and evaluates zero-shot transfer on unseen target languages using the corresponding test splits. Development data is used for model selection. To facilitate fair comparison across donor languages, the amount of fine-tuning data is capped at seven hours per language and one hour of test data per language is used.

DonorRank Framework

DonorRank extends the LangRank framework. It constructs ground-truth donor rankings through pairwise transfer experiments. For every donor-target language pair, a multilingual ASR model (w2vBERT) is fine-tuned on speech from the donor language and evaluated in a zero-shot setting on the target language. The resulting transcription error rates provide an empirical measure of transfer quality, allowing donor languages to be ranked independently for each target language.

Using these empirical rankings as supervision, a LightGBM learning-to-rank model is trained with the LambdaRank objective. Given a target language and a set of candidate donors, the model predicts an ordering that maximizes agreement with the observed transfer rankings on the basis of linguistic and dataset feature vectors for the source and target language.

Each donor-target pair is represented using a combination of linguistic and dataset-level features. The linguistic features are obtained from URIEL/lang2vec and include genetic, geographic, syntactic, and phonological representations derived from WALS, SSWL, and PHOIBLE. These features encode different notions of language similarity, ranging from genealogical relationships to structural and phonological properties. In addition, dataset-specific features are incorporated: donor training hours, target training and test hours, target type-to-token ratio, and lexical overlap between donor and target transcriptions.

At inference time, DonorRank predicts a ranked list of donor languages for a new target language, given concatenated linguistic and dataset features for the language. The highest-ranked donors are then selected for ASR fine-tuning and subsequent zero-shot evaluation. DonorRank is evaluated using leave-one-target-language-out cross-validation. For each fold, donor-target pairs involving the held-out target language are excluded from training, and the ranking model predicts donor rankings for the unseen target language.

Experimental Setup

All experiments use the same pre-trained w2vBERT checkpoint and identical fine-tuning hyperparameters. Models are fine-tuned for 10 epochs with a batch size of 16 and a learning rate of 5 × 10−5. Ranking fidelity is evaluated using Normalized Discounted Cumulative Gain (NDCG), which measures agreement between predicted donor rankings and the empirical rankings obtained from pairwise transfer experiments. Zero-shot transfer performance is measured using character error rate (CER) and word error rate (WER) using Levenshtein distance between predicted and reference transcriptions.

DonorRank is compared against two baselines: selecting the closest phylogenetic neighbor for each target and a prominent high-resource language in each collection (Hindi for VAANI-D and Oromo for WAXAL). Multi-donor transfer experiments are performed on six primary low-resource target languages in VAANI-D: Awadhi, Bhili, Garhwali, Halbi, Konkani, and Marwari. For each target language, models are fine-tuned using the top-k ranked donor languages (k = 1,..., 5), distributing the available fine-tuning data approximately evenly across selected donors while maintaining a comparable training budget.

Results

DonorRank accurately predicts optimal donor languages. Across both VAANI-D and WAXAL, DonorRank achieves high agreement between predicted and empirical donor rankings. For VAANI-D, the CER-based ranker has a higher mean NDCG (0.984) than the WER-based ranker, and this difference is statistically significant (t(17) = 4.02, p < 0.001). This shows that CER is a more reliable signal for ranking in VAANI-D than WER, since VAANI-D includes languages where spellings and word boundaries are not standardized, so character-level differences may better capture cross-variety similarity than word-level measures. In contrast, the WER-based ranker performs better on WAXAL (0.948 versus 0.924), but the difference is not significant. WAXAL includes languages from different families and scripts with limited or no linguistic or phylogenetic relationship, and character-level information may be too fine-grained given the composition of the languages.

Accurate donor ranking improves zero-shot transfer. DonorRank consistently picks the optimal donor across all 19 varieties in VAANI-D compared to the genetic neighbor and Hindi baselines. In WAXAL, the genetically similar donor language performs at par or better with DonorRank’s top-1 language more often than it does for VAANI, but DonorRank often identifies donor languages that outperform both the genetic baseline and Oromo. These results suggest that the role of donor ranking differs based on the composition of languages being considered. In VAANI-D, the challenge lies in distinguishing among many closely related candidate donors, where learned ranking consistently identifies the strongest transfer language. In WAXAL, donor selection involves navigating a much broader multilingual landscape, where language relatedness provides a useful starting point but learned ranking can uncover more effective transfer partners.

Saturation effect in multiple donors. Across all six target languages, incorporating multiple donors generally improves performance over single-donor transfer. The highest-ranked languages tend to capture most of the transferable information (k = 2...4), while additional donors provide diminishing returns. This suggests a saturation effect: combining a small number of well-chosen donors is beneficial, but adding more languages yields limited gains and can introduce noise (as in the case of Bhili).

Understanding Donor Language Selection

What linguistic signals govern successful transfer? In VAANI-D, the ranking model relies on geographic proximity, lexical overlap, and the amount of available training data. These cues help distinguish among languages that are already closely related and some of which are mutually intelligible. In contrast, WAXAL places more emphasis on syntactic features while geographic information remains informative. Because WAXAL spans multiple language families and writing systems, structural linguistic properties are stronger indicators of transferability than lexical similarity alone. Notably, genetic similarity does not emerge as an important feature for either dataset, with normalized feature importance of 1.9% and 1.4% for VAANI and WAXAL, respectively.

Phylogenetic relatedness is not sufficient. Phylogenetic relatedness is a useful starting point for donor selection but isn’t sufficient to explain transfer performance across either dataset. In VAANI-D, several closely related language pairs perform well under both DonorRank and genealogy-based selection. However, other cases show that the nearest linguistic relative is not always the most effective donor. For example, DonorRank selects Rajasthani rather than Hindi for Haryanvi, and Marwari as the best donor for Hindi itself, yielding the best transcription error rates. This shows that higher-resource or more widely spoken languages are not necessarily optimal transfer sources. WAXAL exhibits a similar pattern. Closely related pairs such as Amharic–Tigrinya are successfully identified, while for other targets, DonorRank selects donors that differ from the nearest genealogical neighbor and yield improved transfer performance. For instance, Shona, a Bantu language, is the highest-ranked donor for Malagasy, an Austronesian language; these are completely unrelated languages phylogenetically, yet DonorRank’s selection of Shona results in lower transcription error rates (e.g., CER −10 points).

Successful transfer is characterized by complementary signals. Ablation experiments show that no single feature fully explains transfer, although several individual features produce strong rankings. In VAANI-D, phonological features provide the best performance (0.952), followed by geographic (0.950) and genetic features (0.945). In WAXAL, geographic features are the strongest individual predictor (0.953), while syntactic and genetic features also perform well (0.937 and 0.935). Combining linguistic features improves ranking performance beyond most individual features. In VAANI-D, the best linguistic combination reaches 0.963, compared with 0.952 for the strongest single feature, while the full model achieves the highest overall NDCG of 0.986. In WAXAL, several linguistic combinations closely match the strongest individual geographic model. The full model remains competitive at 0.948, but doesn’t outperform the best reduced configurations. These results show that transferability presents differently for the two corpora. VAANI-D benefits from linguistic and resource information, whereas WAXAL can be modeled using a smaller set of linguistic cues.

DonorRank discovers transfer hubs and widely effective donor languages. Some languages consistently rank highest as donors across many targets, acting as transfer hubs in their respective collections. In VAANI-D, languages like Chhattisgarhi (hne) and Magadhi (mag) emerge as effective donors despite not being the highest-resource or most mainstream languages. Similarly, languages such as Shona (sna), Maasai (mas), and Luganda (lug) rank as the strongest donors for a wide range of targets, though they aren’t as widely spoken as other languages in the WAXAL dataset. This suggests that some languages occupy central positions within the transfer landscape of a multilingual collection.

Conclusion

The paper presents DonorRank, a learning-to-rank framework for donor language selection in low-resource cross-lingual ASR. Using two complementary datasets, it shows that effective donor language selection can be learned in both closely related and typologically diverse language compositions. Beyond donor selection, DonorRank is a framework for understanding multilingual transfer itself. The analyses show that the linguistic cues associated with successful transfer depend on the composition of the multilingual dataset, while highlighting that transferability can’t be explained by any single notion of similarity. Instead, effective donor selection emerges from multiple complementary linguistic and corpus-level signals.

Limitations

The experiments were focused on Indic and African languages, and may not generalize to all language families. Additionally, the evaluation focuses on a fixed fine-tuning setup using the w2vBERT model, and other training strategies may affect transfer performance. Lastly, missing linguistic features for some languages in the lang2vec database reflecting the real-world gaps in documentation may influence the findings.

Improvements for AI systems

Improvements to AI Systems Based on This Paper:

  1. Adaptive Donor Selection Engine for Multilingual ASR: Implement a learning-to-rank module (LightGBM with LambdaRank) that automatically predicts optimal donor languages for any new low-resource target language, using linguistic features (genetic, geographic, syntactic, phonological) and dataset statistics (training hours, lexical overlap, type-to-token ratio). This replaces static heuristics like closest relative or most spoken language, improving zero-shot ASR accuracy by up to 10 CER points in cases like Shona→Malagasy.

  2. Feature-Importance-Aware Transfer Analyzer: Build a diagnostic tool that identifies which linguistic cues (e.g., geographic proximity vs. syntactic similarity) drive successful transfer for a given language cluster. The system can dynamically weight features based on dataset composition—e.g., prioritizing phonological and lexical features for closely related Indic languages, while emphasizing syntactic features for typologically diverse African languages—enabling more informed donor selection without exhaustive pairwise experiments.

  3. Transfer Hub Discovery and Recommendation System: Develop an AI component that mines empirical transfer rankings to identify hub languages (e.g., Chhattisgarhi, Shona, Luganda) that consistently serve as effective donors across many targets. The system can recommend a small set of universal donors for new languages in a region, reducing the need for per-target search and enabling rapid deployment in crisis or under-resourced settings.

  4. Multi-Donor Saturation Optimizer: Implement an algorithm that determines the optimal number of donors (typically 2–4) before performance plateaus, based on predicted ranking confidence and feature diversity. This avoids wasteful fine-tuning on redundant donors and mitigates noise from poorly chosen additions, as observed with Bhili. The system can allocate a fixed training budget (e.g., 7 hours) across the most complementary donors automatically.

  5. Orthography-Aware Evaluation and Ranking: For languages with non-standardized spelling (e.g., Devanagari varieties), the system uses CER-based ranking instead of WER to avoid penalizing legitimate orthographic variation. This improves donor selection fidelity and can be extended to other script-mismatched or code-mixed corpora, making the AI robust to real-world data noise.

  6. Cross-Lingual Transfer Pattern Predictor: Leverage the identified patterns (e.g., unrelated languages can be strong donors, genetic similarity is insufficient) to build a predictive model that scores donor-target pairs even when no empirical transfer data exists. This enables zero-shot donor selection for entirely unseen languages, using only linguistic features and learned transfer dynamics from prior multilingual collections.

Capabilities of the Improved AI System:

  • Automatically selects top-k donors for any new low-resource language with minimal human input.

  • Provides explainable rankings, showing which features (e.g., geography vs. syntax) drive each recommendation.

  • Adapts its selection strategy based on language family density and typological diversity.

  • Avoids overfitting to genealogical trees, uncovering non-obvious but effective transfer partners.

  • Handles orthographic variation gracefully, improving accuracy for under-standardized languages.

  • Scales to new corpora without retraining from scratch, using transferable feature importance patterns.

Abstract

Low-resource automatic speech recognition (ASR) commonly relies on cross-lingual transfer, where models are adapted from higher-resource donor languages. However, selecting donors remains challenging for spontaneous speech from under-resourced language communities, due to linguistic variation, evolving orthographic conventions, and uneven resource availability. We present DonorRank, a learning-to-rank framework for predicting effective donor languages for zero-shot ASR. We evaluate DonorRank on two multilingual speech corpora of Indic and African language families. It accurately predicts donor language rankings and improves donor selection over common heuristics based on genetic similarity or high-resource languages. Beyond improving transfer, we show how DonorRank is a general framework for analyzing donor language selection itself. Our analyses show that the composition of the donor set determines which linguistic cues are useful in predicting successful transfer. We also identify transfer patterns that provide practical guidance for multilingual ASR in low-resource settings.

Sources

Related papers