Universal Byte-Level Encoding: UTF-8/UTF-16 Routing to Reduce Cross-Script Token-Budget Disparities
cs.CL, cs.LG
Submitted: 2026-10-01
Updated: 2026-10-02
Code: https://github.com/meta-llama/llama-models
Terminology
Sources
- One Tokenizer To Rule Them All: Emergent Language Plasticity via Multilingual Tokenizers
- Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
- Stable LM 2 1.6B Technical Report
- BLOOM: A 176B-Parameter Open-Access Multilingual Language Model
- Reducing Tokenization Premiums for Low-Resource Languages
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Separate Before You Compress: The WWHO Tokenization Architecture
- The Script Tax: Measuring Tokenization-Driven Efficiency and Latency Disparities in Multilingual Language Models
- The Llama 3 Herd of Models
- Bolmo: Byteifying the Next Generation of Language Models
- Bit-level BPE: Below the byte boundary
- Qwen2.5 Technical Report
- MUTANT: A Recipe for Multilingual Tokenizer Design
- The Roots of Performance Disparity in Multilingual Language Models: Intrinsic Modeling Difficulty or Design Choices?
- Fast Transformer Decoding: One Write-Head is All You Need
- The Art of Breaking Words: Rethinking Multilingual Tokenizer Design
- TinyLlama: An Open-Source Small Language Model
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering