Apples to Apples? Towards Comparable Crosslingual Language Model Evaluation
cs.CL
Submitted: 2026-08-25
Updated: 2026-08-30
Comments: EMNLP 2026 CR
Code: https://github.com/xiulinyang/multilingual-eval
License: http://creativecommons.org/licenses/by/4.0/
The gist: Crosslingual evaluation of language models that enables fair comparisons remains a fundamental challenge in multilingual NLP.
Terminology
Abstract
Crosslingual evaluation of language models that enables fair comparisons remains a fundamental challenge in multilingual NLP. Existing studies adopt a variety of downstream tasks and intrinsic metrics with different theoretical justifications, yet there has been little empirical investigation into whether these approaches yield meaningful crosslingual conclusions. We systematically examine crosslingual evaluation approaches using controlled monolingual language models trained on parallel data with varying tokenizer vocabulary sizes and model sizes, and further validate our findings on multilingual LLMs. We further discuss challenges in achieving comparable downstream evaluation across languages. Our results show that several widely used normalized metrics introduce crosslinguistic biases rooted in tokenization, encoding, and orthographic differences. In contrast, sentence-level negative log-likelihood computed over semantically equivalent sequences provides more meaningful and consistent crosslingual comparisons.
Sources
- Teuken-7B-Base & Teuken-7B-Instruct: Towards European LLMs
- SinLlama -- A Large Language Model for Sinhala
- When Is Multilinguality a Curse? Language Modeling for 250 High- and Low-Resource Languages
- No Language Left Behind: Scaling Human-Centered Machine Translation
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling
- The Llama 3 Herd of Models
- MaLA-500: Massive Language Adaptation of Large Language Models
- Scaling Language Models: Methods, Analysis & Insights from Training Gopher
- How Much is Enough? The Diminishing Returns of Tokenization Training Data
- The Roots of Performance Disparity in Multilingual Language Models: Intrinsic Modeling Difficulty or Design Choices?
- The Art of Breaking Words: Rethinking Multilingual Tokenizer Design
- The Bitter Lesson Learned from 2,000+ Multilingual Benchmarks
- Qwen3 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering