AfriSUD: A Dependency Treebank Collection for Evaluating Models on African Languages
cs.CL, cs.AI
Submitted: 2026-06-10
Updated: 2026-09-01
License: http://creativecommons.org/licenses/by/4.0/
The gist: Despite their linguistic diversity and global significance, African languages remain underrepresented in research and resources to support NLP.
Terminology
Abstract
Despite their linguistic diversity and global significance, African languages remain underrepresented in research and resources to support NLP. We aim to bridge this gap by introducing AfriSUD, the first large-scale collection of syntactically annotated treebanks for nine diverse African languages spanning major language families and regions across Sub-Saharan Africa. Using the Surface-Syntactic Universal Dependencies (SUD) framework, our community-led effort provides high-quality, native-speaker verified data that capture typological key features such as agglutination and tone. We evaluate a range of models on AfriSUD for part-of-speech tagging and dependency parsing including non-transformer baselines, multilingual pretrained encoders, and LLMs. Our results reveal a significant syntax gap, where models still show clear limitations across the nine languages, suggesting that existing architectures may not fully capture the structural diversity of African-language syntax.
Sources
- Lugha-Llama: Adapting Large Language Models for African Languages
- Deep Biaffine Attention for Neural Dependency Parsing
- GPT-4o System Card
- Neural Machine Translation with Dynamic Graph Convolutional Decoder
- Gemma 3 Technical Report
- AfriqueLLM: How Data Mixing and Model Architecture Impact Continued Pre-training for African Languages
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering