Apples to Apples? Towards Comparable Crosslingual Language Model Evaluation

arXiv:2608.25089 · cs.CL · Submitted 2026-08-25 · Read on arXiv

cs.CL

Submitted: 2026-08-25

Updated: 2026-08-30

Comments: EMNLP 2026 CR

Code: https://github.com/xiulinyang/multilingual-eval

License: http://creativecommons.org/licenses/by/4.0/

The gist: Crosslingual evaluation of language models that enables fair comparisons remains a fundamental challenge in multilingual NLP.

Terminology

Abstract

Crosslingual evaluation of language models that enables fair comparisons remains a fundamental challenge in multilingual NLP. Existing studies adopt a variety of downstream tasks and intrinsic metrics with different theoretical justifications, yet there has been little empirical investigation into whether these approaches yield meaningful crosslingual conclusions. We systematically examine crosslingual evaluation approaches using controlled monolingual language models trained on parallel data with varying tokenizer vocabulary sizes and model sizes, and further validate our findings on multilingual LLMs. We further discuss challenges in achieving comparable downstream evaluation across languages. Our results show that several widely used normalized metrics introduce crosslinguistic biases rooted in tokenization, encoding, and orthographic differences. In contrast, sentence-level negative log-likelihood computed over semantically equivalent sequences provides more meaningful and consistent crosslingual comparisons.

Sources

Related papers