MultiSynt/MT: Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages
cs.CL
Submitted: 2026-07-01
Updated: 2026-09-08
Comments: EMNLP 2026 Camera-ready Version
Code: https://github.com/ellamind/inference-hive
License: http://creativecommons.org/licenses/by/4.0/
The gist: Open web-scale pre-training corpora remain concentrated in English, limiting multilingual LLM development.
Terminology
Abstract
Open web-scale pre-training corpora remain concentrated in English, limiting multilingual LLM development. We introduce MultiSynt/MT, an open synthetic parallel corpus with approximately 4.8 trillion target-language tokens across 36 languages, produced by translating 100 billion high-quality Nemotron-CC tokens with Tower+ and OPUS-MT/HPLT-MT systems. For many medium- and lower-resource European languages, this is the largest openly available pre-training resource. Across five high- and medium-resource languages, reference LLMs trained on MultiSynt/MT reach the final score of HPLT 2.0, a native-data baseline, using roughly 72% fewer pre-training tokens, and outperform it by approximately 15% relative at a matched 100B-token training budget. Our analyses also identify evaluation blind spots: standard multiple-choice benchmarks miss translation-quality differences that a fluency-sensitive LLM-as-judge protocol recovers on the trained LLMs without detecting a deficit relative to its native-data baseline, while Norwegian idiomatic and culturally grounded tasks remain better served by native data. We release the corpus, including row-aligned translations from multiple systems, to support controlled research on multilingual pre-training data and evaluation.
Sources
- Who Benchmarks the Benchmarks? Towards Comprehensive Evaluation of Commonsense Reasoning Benchmarks
- BLOOM: A 176B-Parameter Open-Access Multilingual Language Model
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Aya Expanse: Combining Research Breakthroughs for a New Multilingual Frontier
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling
- Gemma 3 Technical Report
- Massively Multilingual Adaptation of Large Language Models Using Bilingual Translation Data
- EMMA-500: Enhancing Massively Multilingual Adaptation of Large Language Models
- MultiBLiMP 1.0: A Massively Multilingual Benchmark of Linguistic Minimal Pairs
- FIN-bench-v2: A Unified and Robust Benchmark Suite for Evaluating Finnish Large Language Models
- Preliminary Ranking of WMT25 General Machine Translation Systems
- Rethinking Multilingual Continual Pretraining: Data Mixing for Adapting LLMs Across Languages and Resources
- Few-shot Learning with Multilingual Language Models
- Evaluating GPT-3.5 and GPT-4 Models on Brazilian University Admission Exams
- HPLT 3.0: Very Large-Scale Multilingual Resources for LLMs and MT. Mono- and Bi-lingual Data, Multilingual Evaluation, and Pre-Trained Models
- EuroLLM-9B: Technical Report
- FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
- Open-sci-ref-0.01: open and reproducible reference baselines for language model and dataset comparison
- Sometimes We Want Translationese
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering