UrduMMLU: A Massive Multitask Benchmark for Urdu Language Understanding

arXiv:2606.07167 · cs.CL, cs.AI · Submitted 2026-06-05 · Read on arXiv

cs.CL, cs.AI

Submitted: 2026-06-05

Updated: 2026-09-11

Comments: 30 pages, 18 figures, 19 tables, Published In Proceedings of The 2026 Conference on Empirical Methods in Natural Language Processing

Code: https://github.com/meta-llama/llama-models

License: http://creativecommons.org/licenses/by/4.0/

The gist: Meaningful multilingual evaluation must test models in the target language and educational context.

Terminology

Abstract

Meaningful multilingual evaluation must test models in the target language and educational context. Urdu, spoken by more than 230 million people, lacks a broad MMLU-style benchmark built from native educational sources. We introduce UrduMMLU, a benchmark of 26,389 Urdu MCQs across 26 subjects and five domains, collected from native Urdu MCQ banks and public examination PDFs. Unlike translation-based benchmarks, UrduMMLU combines academic subjects with content specific to Urdu and regional education. We label the exam-derived portion through dual human annotation with strict consensus filtering. We evaluate 30 LLMs under English and Urdu prompts, yielding 60 zero-shot evaluations, and further evaluate four open-source LLMs under multiple few-shot settings across both prompt languages. Gemini-3.5-Flash performs best, reaching 90.23% and 90.45% accuracy, while no other model exceeds 85%. The strongest open-source model trails by 7.78 and 9.12 points, and many models lose 25 to 40 points on Urdu-centered Humanities subjects compared with STEM. Few-shot prompting yields only modest gains. Results on UrduMMLU show that current LLMs have uneven Urdu knowledge, particularly for content grounded in the regional context.

Sources

Related papers