Fewer Words, Not Fewer Tokens: Measuring the Sanskrit Tokenization Penalty per Proposition

arXiv:2609.12960 · cs.CL · Submitted 2026-09-11 · Read on arXiv

cs.CL

Submitted: 2026-09-11

Updated: 2026-09-11

Comments: 20 pages, of which 8 are the body; 4 figures, 18 tables. Code, data pipeline and the full results snapshot: https://github.com/DS436/sanskrit-token

Code: https://github.com/DS436/sanskrit-token

License: http://creativecommons.org/licenses/by/4.0/

The gist: Sanskrit fuses case, number, person and tense into word endings and chains clauses into compounds, so it is information-dense per word.

Terminology

Abstract

Sanskrit fuses case, number, person and tense into word endings and chains clauses into compounds, so it is information-dense per word. Whether that density survives subword tokenization is a separate question, to be asked per unit of meaning rather than per word. On identical FLORES-200 devtest content, Sanskrit costs 1.774-2.187 times the English tokens under deployed tokenizers with vocabularies of 200,019 ids or more, but only 1.325-1.353 times the Hindi tokens. Against a deployed English tokenizer, Sanskrit-trained BPE arms then look cheaper per proposition than English on contemporary prose (0.887). Against a matched English control, the same algorithm and vocabulary trained on the English side of the same corpus, that flip disappears: at 32,000 and 64,000 pieces all 8 matched pairs, each size-matched arm against both a pair-matched and a byte-matched control, sit above 1.0 on prose with 95% intervals excluding it. The gap closes as the vocabulary grows: at 128,000 pieces the BPE pair reads 0.983 in domain while staying above parity out of domain (1.025) and on FLORES (1.116). The ratio factorises into a character-length ratio and a tokens-per-character ratio, the second near 1 throughout: what survives matched tokenization is character-level length, which Sanskrit prose lacks over English in SLP1 (1.028) and Sanskrit verse has (0.596). The robust statement is about deployed practice: on contemporary prose and on FLORES, with the Sanskrit side in SLP1 against the deployed o200k English pivot, Sanskrit costs 1.831-2.899 English tokens per proposition under the tokenizers people actually ship. Code, the results snapshot and every table here are public.

Related papers