Test Set Quality in Multilingual LLM Evaluation
cs.CL
Submitted: 2025-08-04
Updated: 2025-11-13
Comments: to appear in the proceedings of Eval4NLP workshop at AACL 2025. Camera ready version
Journal ref: https://aclanthology.org/2025.eval4nlp-1.14/
DOI: 10.18653/v1/2025.eval4nlp-1.14
Code: https://github.com/nishkalavallabhi/testsetquality
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Several multilingual benchmark datasets have been developed in a semi-automatic manner in the recent past to measure progress and understand the state-of-the-art in the multilingual capabilities of
Terminology
Abstract
Several multilingual benchmark datasets have been developed in a semi-automatic manner in the recent past to measure progress and understand the state-of-the-art in the multilingual capabilities of Large Language Models. However, there is not a lot of attention paid to the quality of the datasets themselves, despite the existence of previous work in identifying errors in even fully human-annotated test sets. In this paper, we manually analyze recent multilingual evaluation sets in two languages - French and Telugu, identifying several errors in the process. We compare the performance difference across several LLMs with the original and revised versions of the datasets and identify large differences (almost 10% in some cases) in both languages). Based on these results, we argue that test sets should not be considered immutable and should be revisited, checked for correctness, and potentially versioned. We end with some recommendations for both the dataset creators as well as consumers on addressing the dataset quality issues.
Sources
- Annotation Errors and NER: A Study with OntoNotes 5.0
- Evaluating the Quality of Benchmark Datasets for Low-Resource Languages: A Case Study on Turkish
- MultiLoKo: a multilingual local knowledge benchmark for LLMs spanning 31 languages
- Spanish and LLM Benchmarks: is MMLU Lost in Translation?
- INCLUDE: Evaluating Multilingual Language Understanding with Regional Knowledge
- Measuring the Effect of Disfluency in Multilingual Knowledge Probing Benchmarks
- From Guidelines to Practice: A New Paradigm for Arabic Language Model Evaluation
- Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation
- IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation
- MILU: A Multi-task Indic Language Understanding Benchmark
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering