Why Pretraining Fails to Share Cross-Lingual Knowledge
cs.CL, cs.AI, cs.LG
Submitted: 2026-09-16
Updated: 2026-09-16
Code: https://github.com/AdamJaber03/torchtitan-mulitlingual
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large Language Models (LLMs) have made remarkable progress in the processing and modeling of many languages.
Terminology
Abstract
Large Language Models (LLMs) have made remarkable progress in the processing and modeling of many languages. Yet, unlike human multilinguals, they exhibit surprisingly limited cross-lingual knowledge transfer. While this limitation is well documented, its origins during multilingual training remain unclear. We pretrain 360M- and 7B-parameter LLMs and show that poor cross-lingual knowledge generalization emerges during pretraining and persists under standard interventions. To isolate its cause, we employ a controlled bilingual pretraining setting using two copies of the same language, sharing identical text and token segmentation, but mapped to disjoint token spaces. We find that disjoint tokens alone are enough to induce knowledge compartmentalization, even between identical copies of the same language, establishing disjoint token spaces as a fundamental barrier to cross-lingual knowledge generalization. Guided by this understanding, we suggest mapping languages into a shared token space by simple word-wise translation and find it substantially improves cross-lingual knowledge generalization, recovering up to 12.6% of native-language learning efficiency --- 14 times the baseline.
Sources
- SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
- Fineweb-Edu-Ar: Machine-translated Corpus to Support Arabic Small Language Models
- Large Reasoning Models Struggle to Transfer Parametric Knowledge Across Scripts
- Empty Shelves or Lost Keys? Recall Is the Bottleneck for Parametric Factuality
- Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures
- Crosslingual Capabilities and Knowledge Barriers in Multilingual Large Language Models
- ECLeKTic: a Novel Challenge Set for Evaluation of Cross-Lingual Knowledge Transfer
- The Llama 3 Herd of Models
- Language models struggle with compartmentalization
- Dynamic Chunking for End-to-End Hierarchical Sequence Modeling
- Scripts Through Time: A Survey of the Evolving Role of Transliteration in NLP
- Tracing Multilingual Factual Knowledge Acquisition in Pretraining
- Inference-Time Steering for Cross-Lingual Factual Consistency in LLMs
- Qwen2.5 Technical Report
- The Role of Language Imbalance in Cross-lingual Generalisation: Insights from Cloned Language Experiments
- Representation Learning with Contrastive Predictive Coding
- PolyFact: Comparing Consistency-Driven Post-training Methods for Cross-Lingual Factual Recall
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering