Skip to results
MLSift
← Feed
routineNLP & Language ModelsLLM2608.21462

CyrillicQA: The Influence of Phonetically Encoded Secret Language on LLM Performance

Erik Thureck

cs.CL cs.AI cs.LG

Abstract

Due to the selection of their training data, large language models (LLMs) perform best on standard-language inputs from languages using the Latin alphabet with large speaker populations, while disadvantaging other language varieties. Nevertheless, they can also be a versatile tool for preserving precisely such endangered languages. But do they also possess the necessary creativity and capacity for abstraction to decode phonetically encoded language the same way humans do?

Topics

Classified with taxonomy v2 on Sat, 5 Sept 2026.

The PDF is 1–3 MB. Open it in your browser's viewer, or load it here.

Open PDF