CyrillicQA: The Influence of Phonetically Encoded Secret Language on LLM Performance
cs.CL, cs.AI, cs.LG
Submitted: 2026-08-20
Updated: 2026-08-20
Comments: 4 pages, in English; 4 pages, in German (original); German version originally published in: Rüdian, S. (2026). Prompt-Engineering in Education (2nd ed., pp. 39-42). Humboldt-Universität zu Berlin. https://doi.org/10.5281/zenodo.21413892
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: Due to the selection of their training data, large language models (LLMs) perform best on standard-language inputs from languages using the Latin alphabet with large speaker populations, while
Terminology
Abstract
Due to the selection of their training data, large language models (LLMs) perform best on standard-language inputs from languages using the Latin alphabet with large speaker populations, while disadvantaging other language varieties. Nevertheless, they can also be a versatile tool for preserving precisely such endangered languages. But do they also possess the necessary creativity and capacity for abstraction to decode phonetically encoded language the same way humans do?
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering