A Systematic Review of NLP for Ghanaian Languages: Datasets, Models, and a Research Roadmap
cs.CL
Submitted: 2024-05-10
Updated: 2026-09-16
Comments: 8 main pages. Includes an appendix with supplementary methodology and abstract translations in 15 languages
License: http://creativecommons.org/licenses/by/4.0/
The gist: Natural Language Processing (NLP) for Ghana's 73 living indigenous languages remains deeply fragmented, under-resourced, and heavily skewed toward a single language.
Terminology
Abstract
Natural Language Processing (NLP) for Ghana's 73 living indigenous languages remains deeply fragmented, under-resourced, and heavily skewed toward a single language. We present the first systematic review of the Ghanaian NLP landscape, screening 17,000+ publications across four academic databases to critically synthesize 36 core studies spanning datasets, model architectures, and evaluation paradigms. Our analysis exposes a severe resource imbalance: Twi-centric NLP has grown modestly, driven largely by religious-text alignment and crowdsourcing, while the remaining 70+ languages remain almost entirely unaddressed, and dataset releases, model checkpoints, and evaluation practices remain inconsistent and rarely shared across the field. We translate these findings into a prioritized roadmap targeting Ghana's acute regional constraints, dialectal variation, non-standardized orthographies, and the absence of shared infrastructure, offering a replicable template for systematic review and research prioritization in other low-resource language settings.
Sources
- NLP for Ghanaian Languages
- English-Twi Parallel Corpus for Machine Translation
- Contextual Text Embeddings for Twi
- PaLM: Scaling Language Modeling with Pathways
- On Measures of Biases and Harms in NLP
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- A Primer on Pretrained Multilingual Language Models
- LSTM: A Search Space Odyssey
- English2Gbe: A multilingual machine translation model for {Fon/Ewe}Gbe
- Survey of Low-Resource Machine Translation
- Low-resource Languages: A Review of Past Work and Future Challenges
- PidginUNMT: Unsupervised Neural Machine Translation from West African Pidgin to English
- GPT-4 Technical Report
- Masakhane -- Machine Translation For Africa
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
- When does Bias Transfer in Transfer Learning?
- AI4D -- African Language Program
- Are Pre-trained Convolutions Better than Pre-trained Transformers?
- Attention Is All You Need
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering