Expanding the Lexicon of Ge'ez Based African Languages: A Comparative Study of Amharic and Tigrinya
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Expanding the Lexicon of Ge'ez Based African Languages".
Jane: Multilingual pre-trained language models struggle with low-resource Ge’ezscript languages like Amharic and Tigrinya due to high out-of-vocabulary rates and excessive subword fragmentation from Latin-script tokenizers,
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So, Jane, we've got this fascinating paper on arXiv called "Expanding the Lexicon of Ge'ez Based African Languages: A Comparative Study of Amharic and Tigrinya." Basically, the authors are tackling a big problem where multilingual models like XLM-R struggle when it comes to languages like Amharic and Tigrinya because they get overwhelmed with words the model hasn't seen before.
Jane: That’s right, Tom; it tackles the high out-of-vocabulary rates and the excessive subword fragmentation that happens when these models use tokenizers built for Latin scripts. The core thesis of this paper is that existing large models aren't optimized for Ge’ez-script languages, and they propose a new approach called VEXMLM to fix it by targeting Amharic and Tigrinya specifically.
Lu: It’s really interesting how they systematically combine several different techniques into one framework; combining language-specific SentencePiece tokenization with vocabulary augmentation, mean initialization for embeddings, and a two-stage training procedure is quite a comprehensive strategy.
Meng: From an engineering standpoint, that level of systematic combination sounds complex to implement efficiently, but if it solves the OOV issue without retraining the whole model from scratch every time for a new language, that’s something we can actually get behind practically.
Lalam: I see this as a huge cultural tool; if we can get these models to truly understand and represent Amharic and Tigrinya better, it opens up incredible avenues for preserving and processing the rich linguistic heritage of Africa in ways current technology simply doesn't allow.
Tom: Exactly, Lalam; that focus on representation is what makes this important for languages that are often underrepresented in global AI efforts. The paper claims VEXMLM substantially outperforms both XLM-R and Glot500 across all the tasks they tested on Amharic and Tigrinya.
Jane: And those results are quite striking, Tom; specifically, they report particularly strong gains for out-of-vocabulary entity recognition in Amharic and Tigrinya, which suggests the vocabulary expansion really hit the mark where it needed to.
Lu: The authors also point out that this work differs from other vocabulary expansion efforts because they focus specifically on Ge’ez-script languages, which haven't received systematic attention in that literature despite their significance.
Meng: I noticed they mention they train a vocabulary of fifty thousand subword units for Tigrinya and thirty-two thousand for Amharic, which adds about three hundred eighty-one million parameters to the total vocabulary size compared to the baseline. That’s a pretty substantial increase in model complexity.
Paper summary: Tom: It is substantial, Meng; but they explain how they handle that extra size by only retaining tokens from those new tokenizers that were absent from XLM-R’s original vocabulary, which amounts to adding thirty thousand net new Ge’ez-derived subword tokens. That's smart parameter management.
Jane: And the embedding initialization strategy is also key; they use the mean of all source embeddings to initialize these new token embeddings, which positions them at what they call the geometric centroid of the source embedding distribution, which helps keep things aligned.
Lu: That centering technique prevents those new tokens from pulling the model's understanding in a random direction or creating some kind of bias toward one specific source language during that initial phase.
Lalam: It’s like giving every new word a balanced starting point in the model's knowledge space, ensuring that when it encounters something new, it doesn't immediately drift away from what it already knows about the existing languages.
Tom: And then they run a two-stage training process: first, continued masked language modeling on the curated corpora to update everything including those new embeddings, and then supervised fine-tuning for specific tasks like question answering and named entity recognition.
Meng: The two-stage approach sounds robust; the initial pretraining lets the model absorb the language patterns thoroughly before it gets specialized for those downstream applications. What about that second stage?
Jane: In the second stage, they update only the task-specific classification head and the embedding layer while keeping the transformer body partially frozen to prevent overfitting during that fine-tuning process. That careful approach is what lets them get high performance on those specific tasks without corrupting the general language understanding they built in Stage one.
Lu: The paper does touch upon other related work, mentioning things like EEVE-Korean vocabulary expansion and QTok for evaluating tokenization quality across languages, showing they are aware of the broader landscape but highlighting their unique focus on Ge’ez-script languages.
Tom: So, to wrap up this overview of "Expanding the Lexicon of Ge'ez Based African Languages: A Comparative Study of Amharic and Tigrinya," we see a comprehensive framework that addresses tokenization inefficiencies through targeted vocabulary expansion, careful embedding initialization, and a structured two-stage training pipeline.
Jane: The authors are making a strong case that this systematic combination provides an effective path to improve multilingual models for underrepresented languages without requiring them to retrain from scratch.
Lu: The implications here are huge because it shows that language-aware vocabulary adaptation can substantially improve the performance of these large multilingual models for African languages, which have historically been neglected in this research area.
Paper summary: Meng: From a practical standpoint, if this framework proves efficient enough, we might see a significant reduction in the data and compute needed to successfully deploy models for other low-resource script families down the line.
Lalam: Imagine the impact on accessibility; having better AI tools that genuinely understand Amharic and Tigrinya means these communities can access information, create content, and utilize technology with a level of fidelity they’ve never had before.
Tom: It really boils down to this: VEXMLM substantially outperforms XLM-R and Glot500 on Amharic and Tigrinya across various evaluations, with specific gains noted in out-of-vocabulary handling.
Jane: That performance parity is what makes the paper compelling; it shows that optimizing for a specific script family isn't just a marginal improvement but leads to measurable gains in entity recognition and overall task accuracy.
Lu: The comparative analysis also showed that while VEXMLM beats Glot500 on some metrics, Glot500 still achieves higher accuracy on named entity recognition, which gives us a nuanced view of where the trade-offs lie between scale and specificity.
Meng: The ablation study provided some good clarity too; it pointed out that while vocabulary expansion alone with random initialization only gave a modest gain, the continued pretraining provided the largest single improvement, which really validates that iterative training process.
Tom: So, to conclude this discussion on "Expanding the Lexicon of Ge'ez Based African Languages: A Comparative Study of Amharic and Tigrinya," we see that VEXMLM’s success is driven by two main factors: the enhanced vocabulary coverage through targeted Ge’ez-script subword augmentation, and continued MLM pretraining that updates representations to reflect the specific morphological and orthographic patterns of Amharic and Tigrinya.
Jane: The authors conclude that this work demonstrates that language-aware vocabulary adaptation can substantially improve multilingual language models for underrepresented African languages, which really underscores the importance of focusing research efforts on these areas.
Lu: It highlights how integrating language-specific adaptations into a unified framework can be a very powerful way to address the representational gaps in current large-scale pretraining methods.
Meng: I think the main impact is demonstrating that we can achieve meaningful performance improvements for languages like Amharic and Tigrinya using targeted, systematic techniques instead of just relying on massive amounts of general data to cover everything.
Lalam: This work means that the future of AI tools won't just be about covering a few major languages; it will be about building models that are truly inclusive and capable of understanding the diverse linguistic realities across the continent.
Conclusion: Tom: So, we've been deep into the technical nitty-gritty of VEXMLM and how it tackles those tokenization headaches for Amharic and Tigrinya, and now we need to wrap up what this whole paper is really saying about its title and the authors.
Jane: Absolutely, Tom; the paper is titled "Expanding the Lexicon of Ge'ez Based African Languages: A Comparative Study of Amharic and Tigrinya," which just tells us exactly what they’re focused on—the vocabulary expansion for these specific languages.
Lu: That title perfectly frames the research because it highlights that this isn't just about making a model bigger; it’s about deliberately enriching the model's understanding of a specific linguistic domain, which is such a creative way to think about language representation.
Meng: From an engineering standpoint, that emphasis on "Expanding the Lexicon" suggests they are prioritizing adding meaningful tokens rather than just throwing random vocabulary at the problem, which makes sense for practical deployment.
Lalam: I see it as a powerful act of cultural preservation; by expanding the lexicon specifically for these languages, the authors are ensuring that their unique structures and nuances get properly encoded into the digital world.
Tom: Exactly, Lalam; this research demonstrates that when you tailor the model's vocabulary to match a language's specific character, you get much better results than just using a generic multilingual tool.
Jane: And those results show that VEXMLM significantly outperforms existing models on tasks like out-of-vocabulary recognition for these languages, which is what the authors are really proving here.
Lu: The comparative study aspect is also vital because they didn't just test it in a vacuum; they compared it against XLM-R and Glot500, which gives us concrete evidence of where this specific adaptation makes a difference in performance across different architectures.
Meng: I’m interested in the implications for building models for other low-resource scripts; if this systematic method works, it suggests there might be a template we can use to improve representation for other African languages down the road.
Lalam: It means that the future of AI tools won't just be about covering a few major languages; it will be about building models that are truly inclusive and capable of understanding the diverse linguistic realities across the continent.
Tom: Right, Lalam; this paper is showing us a blueprint for how to make global AI more linguistically aware, moving beyond just scaling up general models.
Jane: It really boils down to this: language-aware vocabulary adaptation, done systematically, can substantially improve multilingual models for underrepresented African languages.
Lu: The authors conclude that by combining targeted subword augmentation with continued pretraining on the actual data, they’ve found a much more efficient way to handle these complex linguistic challenges than previous methods.
Meng: I just want to know if they flagged any limitations; does this approach still run into issues when we try to apply it to completely unrelated languages outside the Ge'ez family?
Lalam: That is a fair question, Meng; and while the focus is on Amharic and Tigrinya now, the underlying principles of targeted augmentation are very general, suggesting potential for broader application.
Tom: Well, that’s our next big topic—we need to look closely at those limitations and what the authors suggest for future work before we move on to how this actually impacts everyday life.
Hailay Kidu Teklehaymanot†, Debela Desalegn Yadeta‡, Wolfgang Nejdl†
L3S Research Center, Leibniz University Hannover, Germany · Addis Ababa University, Ethiopia
cs.CL
Submitted: 2026-07-16
Updated: 2026-09-28
Comments: 12 pages , 5 tables , 1 figurs
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 87/100
The gist: Multilingual pre-trained language models struggle with low-resource Ge’ezscript languages like Amharic and Tigrinya due to high out-of-vocabulary rates and excessive subword fragmentation from
Key concepts
- VEXMLM
- A vocabulary-extended variant of XLM-R designed to handle low-resource Ge'ez languages. It integrates language adaptation, new subword tokens, mean embedding initialization, and a two-stage training method to overcome tokenization and OOV issues.
- Ge'ezderived Subwords
- Custom vocabulary units created specifically for Amharic and Tigrinya based on their unique linguistic structure. This expansion adds thousands of new tokens that help the model better represent the specific words and morphemes found in these languages.
- Mean Initialization
- A technique where new token embeddings are set to the average of all existing source embeddings. This places new tokens at a central point in the embedding space, preventing them from being placed randomly or biased towards any single source language.
- Two-Stage Training
- A training procedure consisting of two steps: first, continued masked language modeling pretraining on the specific low-resource data to update all parameters; second, task-specific fine-tuning on downstream tasks like NER and QA using a frozen transformer body.
Terminology
Summary
Multilingual pre-trained language models struggle with low-resource Ge’ezscript languages like Amharic and Tigrinya due to high out-of-vocabulary rates and excessive subword fragmentation from Latin-script tokenizers, a problem this paper addresses by introducing VEXMLM, a vocabulary-extended variant of XLM-R. This work is significant because it demonstrates that combining language-specific tokenizer adaptation, vocabulary augmentation with Ge’ezderived subwords, mean initialization for embeddings, and a two-stage training procedure provides an effective and computationally efficient path to improve multilingual models for underrepresented languages without retraining from scratch.
The gist
VEXMLM substantially outperforms XLM-R and Glot500 across all evaluated tasks on Amharic and Tigrinya, with particularly strong gains on out-of-vocabulary entity recognition, while improvements on Amharic/Tigrinya transfer to 17 languages in Africa.
How it works
The VEXMLM framework is a systematic approach combining four key components to mitigate OOV and over-segmentation issues:
-
Language-specific SentencePiece tokenization and vocabulary augmentation with 30,000 Ge’ezderived subword tokens. The paper states that for Tigrinya, they train a vocabulary of 50,000 subword units, and for Amharic, 32,000 subword units. This results in a combined vocabulary of
381M parameters
(up from 279M in the baseline). The process involves retaining only tokens from the new tokenizers that are absent from XLM-R’s original vocabulary, yielding30,000 net new Ge’ez-derived subword tokens added.
-
Mean initialization of new token embeddings to preserve alignment with the existing embedding space. New embeddings are initialized via
the mean of all source embeddings,
which positions them at thegeometric centroid of the source embedding distribution
(Equation 2), reducing the risk of outlier placement and avoiding directional bias toward any specific source language. -
A two-stage training procedure comprising continued masked language modeling pretraining followed by task-specific fine-tuning.
(Stage 1)
Continued pretraining involves continued masked language modeling on the curated corpora
for Amharic and Tigrinya, where all model parameters including original and newly added embeddings are updated.
Low-resource data is upsampled to address corpus imbalance.
(Stage 2)
The model undergoes task-specific fine-tuning on three downstream tasks: named entity recognition (NER), sentiment analysis (SA), and question answering (QA). To mitigate overfitting, this stage updates only the task-specific classification head and the embedding layer while keeping the transformer body partially frozen.
Evaluation Metrics and Results
The model's effectiveness is assessed through intrinsic tokenization quality metrics and extrinsic downstream task performance across 19 low-resource languages in Africa. Intrinsic evaluations show VEXMLM achieves higher parity than XLM-R for Ge’ez-script languages, such as Tigrinya (0.27 vs 1.36). Fertility analysis confirms that the vocabulary expansion addresses tokenization inefficiency, with VEXMLM achieving substantially lower fertility for Tigrinya (2.1 → 1.3, 38% reduction).
Furthermore, OOV handling is improved; VEXMLM improves average OOV-word accuracy by 5.9 points (from 88.4% to 94.3%)
across African languages on NER tasks, with primary targets showing gains of +1.1 and +11.4 points respectively for Amharic and Tigrinya.
Comparative Analysis
When compared against XLM-R, VEXMLM shows superior performance on SA, QA (EM), and QA (F1). Against Glot500, VEXMLM outperforms it on SA (+0.33) and QA (EM: +0.03), though Glot500 achieves higher NER accuracy (0.92 vs 0.78). The ablation study isolates the contribution of each component: Vocabulary expansion alone with random initialization provides a modest gain (+1.2 points over baseline),
while Continued pretraining providing the largest single gain (+7.1 points)
is identified as the most critical component of the pipeline for achieving top OOV accuracy.
Conclusion and Implications
The findings demonstrate that VEXMLM’s improvements are driven by two main factors: enhanced vocabulary coverage through targeted Ge’ez-script subword augmentation, and continued MLM pretraining that updates representations to reflect the morphological and orthographic patterns of Amharic and Tigrinya.
The work concludes that language-aware vocabulary adaptation can substantially improve multilingual language models for underrepresented African languages,
highlighting the importance of "
Improvements for AI systems
Based on the provided scientific paper, here are specific improvements that can be made to existing AI systems, particularly Large Language Models (LLMs), by implementing the VEXMLM framework:
-
Enhance Performance for Low-Resource and Script-Specific Languages (Amharic and Tigrinya):
-
Improve Out-of-Vocabulary (OOV) Recognition Accuracy:
-
Optimize Tokenization Efficiency for Morphologically Rich Ge'ez Scripts:
-
Enable Robust Cross-Lingual Transfer to African Languages:
-
Enhance Performance for Low-Resource and Script-Specific Languages (Amharic and Tigrinya):
The improved system will demonstrate significantly better performance on Amharic and Tigrinya tasks (NER, SA, QA) compared to baseline models like XLM-R or Glot500. This is achieved by adapting the model's internal representations to capture the unique morphological and orthographic patterns of these languages through two mechanisms:
-
The system will utilize a vocabulary expanded with 30,000 Ge’ez-derived subword tokens, ensuring that complex words and specific character combinations are not broken down into meaningless fragments.
-
Continued Masked Language Modeling (MLM) pretraining on curated monolingual corpora for Amharic and Tigrinya will allow the model to learn richer, contextually relevant representations for these specific language structures.
- Improve Out-of-Vocabulary (OOV) Recognition Accuracy:
The system will significantly reduce the rate at which words or entities are mapped to the generic [UNK] symbol, leading to a substantial improvement in OOV word accuracy across all evaluated African languages (e.g., Swahili, Kinyarwanda, Nigerian Pidgin). Specifically for Amharic and Tigrinya NER, OOV accuracy is projected to improve by 1.1 and 11.4 points respectively compared to the baseline XLM-R model. This means the AI can accurately identify named entities even when they appear in novel or less frequent orthographic forms.
- Optimize Tokenization Efficiency for Morphologically Rich Ge'ez Scripts:
The system will achieve superior tokenization efficiency by employing language-specific SentencePiece tokenizers tailored to the specific character inventory of Amharic (32k units) and Tigrinya (50k units). This targeted approach drastically reduces subword fragmentation and sequence length compared to Latin-script-centric tokenizers. The resulting model will exhibit:
-
Significantly lower
Fertility
scores for Tigrinya (projected 38% reduction), meaning fewer tokens are needed to represent the same amount of language, leading to faster inference and reduced computational cost. -
Higher
Compression
metrics, indicating more efficient encoding of the text.
- Enable Robust Cross-Lingual Transfer to African Languages:
The system will facilitate stronger cross-lingual generalization across the 19 evaluated African languages by utilizing a mean-based embedding initialization strategy for the new Ge'ez tokens. This ensures that the newly added vocabulary is placed at the geometric centroid of the source embedding space, preventing destabilization and maintaining compatibility with XLM-R’s existing representation structure. Furthermore, continued pretraining on upsampled low-resource data will strengthen the model's ability to transfer knowledge from Amharic/Tigrinya to other African languages that were not explicitly targeted by vocabulary expansion. This results in improved performance on downstream tasks (SA, QA) across the entire benchmark of 19 languages.
Abstract
Multilingual pre-trained language models such as XLM-R perform well for major languages but struggle with low-resource Ge'ez-script languages, largely because Latin-script-centric tokenizers split their words into many subwords. We introduce VEXMLM, a vocabulary-extended variant of XLM-R targeting Amharic and Tigrinya. We train language-specific SentencePiece tokenizers on monolingual corpora, extend XLM-R's vocabulary with 30k Ge'ez-script subwords, and initialize each new embedding to the mean of the pretrained embeddings. VEXMLM undergoes two-stage training: (1) continued masked language modeling on the monolingual corpora and (2) supervised fine-tuning on question answering and named entity recognition (Amharic and Tigrinya) and sentiment analysis (Amharic). VEXMLM lowers tokenizer fertility below that of XLM-R and Glot500 on both languages, by 28.0% (Amharic) and 45.9% (Tigrinya) relative to XLM-R. Downstream, it modestly improves named entity recognition over XLM-R, scores below XLM-R on extractive question answering, and is comparable on sentiment analysis. An ablation on Tigrinya NER shows that vocabulary expansion alone lowers accuracy on out-of-vocabulary words (words that XLM-R's tokenizer cannot represent or splits into more pieces than the expanded tokenizer), and that continued pretraining is required for the expanded model to exceed the baseline. Vocabulary expansion thus makes Ge'ez-script tokenization substantially more efficient, while its downstream benefit depends on the task and on adapting the new embeddings through continued pretraining. Resources: GitHub repository Hugging Face model.
Sources
- Qtok: A Comprehensive Framework for Evaluating Multilingual Tokenizer Quality in Large Language Models
- Unpacking Tokenization: Evaluating Text Compression and its Correlation with Model Performance
- Accelerating Multilingual Language Model for Excessively Tokenized Languages
- Efficient and Effective Vocabulary Expansion Towards Multilingual Large Language Models
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- Byte Latent Transformer: Patches Scale Better Than Tokens
- Trans-Tokenization and Cross-lingual Vocabulary Transfers: Language Adaptation of LLMs for Low-Resource NLP
- How Multilingual Are Large Language Models Fine-Tuned for Translation?
- Improving Pre-Trained Multilingual Models with Vocabulary Expansion
- Investigating Multilingual Instruction-Tuning: Do Polyglot Models Demand for Multilingual Instructions?
- How Can We Effectively Expand the Vocabulary of LLMs with 0.01GB of Target Language Text?
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering