What Language is This? Ask Your Tokenizer
cs.CL
Submitted: 2026-02-19
Updated: 2026-09-10
Comments: In Proceedings of ICML 2026
Code: https://github.com/Ahmetcanyvz/UNILID
Project page: https://cimeister.github.io
License: http://creativecommons.org/publicdomain/zero/1.0/
The gist: Language Identification (LID) is an important component of many multilingual natural language processing pipelines, where it facilitates corpus curation, training data analysis, and cross-lingual
Terminology
Abstract
Language Identification (LID) is an important component of many multilingual natural language processing pipelines, where it facilitates corpus curation, training data analysis, and cross-lingual evaluation of large language models. Despite near-perfect performance on high-resource languages, existing systems remain brittle in low-resource and closely related language settings. We introduce UniLID, a simple and efficient LID method based on the UnigramLM tokenization algorithm. In short, to predict a string's language label, we simply ask: under which language's unigram distribution is this string most likely? Our formulation is data- and compute-efficient, supports incremental addition of new languages without retraining existing models, and can naturally be integrated into existing language model tokenization pipelines. Empirical evaluations against widely used baselines, including fasttext, GlotLID-M, and CLD3, show that UniLID achieves competitive performance on standard benchmarks, reaches 69% accuracy with five labeled samples per language and 89% with 25, and delivers large gains on fine-grained dialect identification.
Sources
- No Language Left Behind: Scaling Human-Centered Machine Translation
- ConLID: Supervised Contrastive Learning for Low-Resource Language Identification
- Mistral 7B
- The WiLI benchmark dataset for written language identification
- Which Pieces Does Unigram Tokenization Really Need?
- DIVERS-Bench: Evaluating Language Identification Across Domain Shifts and Code-Switching
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering