One Form to Transfer Them All: Pretraining Multilingual Language Models Beyond Native Orthography
cs.CL
Submitted: 2026-08-26
Updated: 2026-09-08
Comments: EMNLP 2026 (Main Conference). 9 pages, 6 figures (plus appendix)
Code: https://github.com/karpathy/nanoGPT
Project page: http://skylion007.github.io/OpenWebTextCorpus
License: http://creativecommons.org/licenses/by/4.0/
The gist: Multilingual language models transfer knowledge across languages through shared subword vocabulary, a mechanism that breaks down when related languages use different writing systems.
Terminology
Abstract
Multilingual language models transfer knowledge across languages through shared subword vocabulary, a mechanism that breaks down when related languages use different writing systems. Prior work addresses this via script equalization (romanization or IPA transcription), but direct comparisons are rare; the focus has been on encoder-only models, with most work adapting existing pretrained models. We systematically compare different input representations in autoregressive multilingual pretraining, comparing orthographic text, IPA, and romanization in a controlled setup across three scales (467M, 709M, and 1.03B) on eight languages in four typologically motivated pairs. Across a wide range of downstream tasks on seen and unseen languages, romanized pretraining yields the strongest cross-lingual transfer, and the advantage over text widens with scale. IPA improves over text in most settings but trails romanization. Surprisingly, finetuning a text-pretrained model on romanized data hurts performance on languages already covered by the base model, only marginally helping when the model lacks script coverage. Our results indicate that for multilingual models spanning typologically diverse scripts, to obtain maximum benefits, romanization should be treated as a core design choice applied at pretraining rather than a post hoc fix.
Sources
- Aya 23: Open Weight Releases to Further Multilingual Progress
- BLOOM: A 176B-Parameter Open-Access Multilingual Language Model
- The Llama 3 Herd of Models
- Cross-Lingual Ability of Multilingual BERT: An Empirical Study
- PolyIPA -- Multilingual Phoneme-to-Grapheme Conversion Model
- LLM-based phoneme-to-grapheme for phoneme-based speech recognition
- XPhoneBERT: A Pre-trained Multilingual Model for Phoneme Representations for Text-to-Speech
- No Language Left Behind: Scaling Human-Centered Machine Translation
- FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
- Qwen2.5 Technical Report
- Qwen3 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering