CyrillicQA: The Influence of Phonetically Encoded Secret Language on LLM Performance

arXiv:2608.21462 · cs.CL, cs.AI, cs.LG · Submitted 2026-08-20 · Read on arXiv

cs.CL, cs.AI, cs.LG

Submitted: 2026-08-20

Updated: 2026-08-20

Comments: 4 pages, in English; 4 pages, in German (original); German version originally published in: Rüdian, S. (2026). Prompt-Engineering in Education (2nd ed., pp. 39-42). Humboldt-Universität zu Berlin. https://doi.org/10.5281/zenodo.21413892

License: http://creativecommons.org/licenses/by-sa/4.0/

The gist: Due to the selection of their training data, large language models (LLMs) perform best on standard-language inputs from languages using the Latin alphabet with large speaker populations, while

Terminology

Abstract

Due to the selection of their training data, large language models (LLMs) perform best on standard-language inputs from languages using the Latin alphabet with large speaker populations, while disadvantaging other language varieties. Nevertheless, they can also be a versatile tool for preserving precisely such endangered languages. But do they also possess the necessary creativity and capacity for abstraction to decode phonetically encoded language the same way humans do?

Related papers