To Each Language Its Tokenizer: Modular Tokenizers for Efficient Multilingual LLMs
cs.CL
Submitted: 2026-09-14
Updated: 2026-09-14
Code: https://github.com/kyutai-labs/modular-tokenization
License: http://creativecommons.org/licenses/by/4.0/
The gist: Multilingual Large Language Models (LLMs) traditionally rely on a single vocabulary shared by all supported languages, which can lead to uneven compression across them.
Terminology
Abstract
Multilingual Large Language Models (LLMs) traditionally rely on a single vocabulary shared by all supported languages, which can lead to uneven compression across them. Moreover, their large embedding and output matrices increase memory usage and slow inference, notably for small-scale models. It is also wasteful as models are often used for only a subset of languages. To address these issues, we introduce a modular framework for multilingual model training. First, we propose methods to learn large modular BPE and Unigram tokenizers that enable extraction of subtokenizers tailored to any language subset. These subtokenizers achieve compression on par with monolingual tokenizers and improve cross-lingual fairness. Second, we design a pretraining strategy that samples subtokenizers to form batches, restricting predictions to the relevant vocabulary subset and allowing efficient training despite a large vocabulary. This supports efficient inference with any combination of language-specific vocabularies. Therefore, it reduces memory usage and speeds up inference in models without sacrificing performance.
Sources
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Efficient and Effective Text Encoding for Chinese LLaMA and Alpaca
- The Llama 3 Herd of Models
- MaLA-500: Massive Language Adaptation of Large Language Models
- ZeRO-Offload: Democratizing Billion-Scale Model Training
- Tiny Aya: Bridging Scale and Multilingual Depth
- Qwen3 Technical Report
- Gemma 3 Technical Report
- Gemma 2: Improving Open Language Models at a Practical Size
- Towards Multilingual LLM Evaluation for European Languages
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering