UrduMMLU: A Massive Multitask Benchmark for Urdu Language Understanding
cs.CL, cs.AI
Submitted: 2026-06-05
Updated: 2026-09-11
Comments: 30 pages, 18 figures, 19 tables, Published In Proceedings of The 2026 Conference on Empirical Methods in Natural Language Processing
Code: https://github.com/meta-llama/llama-models
License: http://creativecommons.org/licenses/by/4.0/
The gist: Meaningful multilingual evaluation must test models in the target language and educational context.
Terminology
Abstract
Meaningful multilingual evaluation must test models in the target language and educational context. Urdu, spoken by more than 230 million people, lacks a broad MMLU-style benchmark built from native educational sources. We introduce UrduMMLU, a benchmark of 26,389 Urdu MCQs across 26 subjects and five domains, collected from native Urdu MCQ banks and public examination PDFs. Unlike translation-based benchmarks, UrduMMLU combines academic subjects with content specific to Urdu and regional education. We label the exam-derived portion through dual human annotation with strict consensus filtering. We evaluate 30 LLMs under English and Urdu prompts, yielding 60 zero-shot evaluations, and further evaluate four open-source LLMs under multiple few-shot settings across both prompt languages. Gemini-3.5-Flash performs best, reaching 90.23% and 90.45% accuracy, while no other model exceeds 85%. The strongest open-source model trails by 7.78 and 9.12 points, and many models lose 25 to 40 points on Urdu-centered Humanities subjects compared with STEM. Few-shot prompting yields only modest gains. Results on UrduMMLU show that current LLMs have uneven Urdu knowledge, particularly for content grounded in the regional context.
Sources
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- Gemma 3 Technical Report
- The Llama 3 Herd of Models
- Qalb: Largest State-of-the-Art Urdu Large Language Model for 230M Speakers with Systematic Continued Pre-training
- UQuAD1.0: Development of an Urdu Question Answering Training Data for Machine Reading Comprehension
- IndicMMLU-Pro: Benchmarking Indic Large Language Models on Multi-Task Language Understanding
- Ministral 3
- Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
- OpenAI GPT-5 System Card
- Gemma 2: Improving Open Language Models at a Practical Size
- Qwen3 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering