English is Not All You Need: Systematically Exploring the Role of Multilinguality in LLM Post-Training
cs.CL, cs.AI
Submitted: 2026-04-14
Updated: 2026-09-20
License: http://creativecommons.org/licenses/by/4.0/
The gist: Despite the widespread multilingual deployment of large language models, post-training pipelines remain predominantly English-centric, contributing to performance disparities across languages.
Terminology
Abstract
Despite the widespread multilingual deployment of large language models, post-training pipelines remain predominantly English-centric, contributing to performance disparities across languages. We present a systematic, controlled study of the interplay between training language coverage, model scale, and task domain, based on 220 supervised fine-tuning runs on parallel translated multilingual data mixtures spanning mathematical reasoning and API calling tasks, with models up to 8B parameters. We find that English-only post-training is typically suboptimal: incorporating even a single non-English language improves both English performance and cross-lingual generalization. Increasing language diversity during post-training generally yields further gains, particularly for low-resource languages, while performance on high-resource languages tends to plateau rather than degrade. Moreover, greater language diversity enables strong zero-shot transfer to unseen languages, reducing the need for direct inclusion, though gains remain limited for typologically distant, low-resource languages.
Sources
- BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models
- Gemma 3 Technical Report
- mCoT: Multilingual Instruction Tuning for Reasoning Consistency in Language Models
- API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
- ATLAS: Adaptive Transfer Scaling Laws for Multilingual Pretraining, Finetuning, and Decoding the Curse of Multilinguality
- Do Multilingual LLMs Think In English?
- Multilingual Instruction Tuning With Just a Pinch of Multilinguality
- Language Models are Multilingual Chain-of-Thought Reasoners
- A Post-trainer's Guide to Multilingual Training Data: Uncovering Cross-lingual Transfer Dynamics
- Qwen3 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering