Challenges and Recommendations for LLM-as-a-Judge in Multilingual Settings and for Low-Resource Languages
cs.CL, cs.AI
Submitted: 2026-07-02
Updated: 2026-09-07
Comments: To appear at EMNLP Findings 2026 (camera-ready version)
License: http://creativecommons.org/licenses/by/4.0/
The gist: LLM-as-a-Judge has become the dominant evaluation paradigm for many natural language generation tasks (albeit mostly in English) due to shortcomings of conventional metrics and high correlations with
Terminology
Abstract
LLM-as-a-Judge has become the dominant evaluation paradigm for many natural language generation tasks (albeit mostly in English) due to shortcomings of conventional metrics and high correlations with human judgment. There are now attempts to extend LLM-as-a-Judge to multilingual settings including low-resource languages. However, LLMs have limited proficiency in low-resource languages, and there is often no adequate human validation in these settings. To highlight the scope of the problem and current practices, we explore the use of LLM-as-a-Judge evaluators in ACL Anthology papers focusing on multilingual settings and low-resource languages across a diverse set of tasks. Out of 650 papers mentioning LLM-as-a-judge, only 33 of them focus on low-resource or multilingual settings. Our in-depth analysis of these papers indicates inconsistent evaluation outcomes, a tendency to overtrust LLM judgments in multilingual settings, and the widespread reliance on a single judge model per study. To help the NLP community further, we conclude with a checklist of recommendations about how to use LLM-as-a-Judge in multilingual and low-resource settings
Sources
- TukaBench: A Culturally Grounded Jailbreak Benchmark for African Languages
- DeepSeek-V3 Technical Report
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- A Survey on LLM-as-a-Judge
- LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods
- LALM-as-a-Judge: Benchmarking Large Audio-Language Models for Safety Evaluation in Multi-Turn Spoken Dialogues
- Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation
- GPT-4 Technical Report
- Socially Responsible Data for Large Multilingual Language Models
- MM-Eval: A Multilingual Meta-Evaluation Benchmark for LLM-as-a-Judge and Reward Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering