Moral Semantics Survive Machine Translation: Cross-Lingual Evidence from Moral Foundations Corpora

arXiv:2605.22660 · cs.CL, cs.AI · Submitted 2026-05-21 · Read on arXiv

cs.CL, cs.AI

Submitted: 2026-05-21

Updated: 2026-09-11

Comments: Published at ACM GoodIT'26

DOI: 10.1145/3794786.3830756

License: http://creativecommons.org/licenses/by/4.0/

The gist: Moral language is subtle and culturally variable, making it difficult to translate faithfully across languages.

Terminology

Abstract

Moral language is subtle and culturally variable, making it difficult to translate faithfully across languages. Idiomatic expressions, slang, and cultural references introduce hard-to-avoid translation artifacts. Yet automated moral values classification depends on language-specific annotated corpora that exist almost exclusively in English. We investigate whether LLM-based translation can bridge this gap, taking Polish as a test case. Using 50k morally-annotated social media posts from a diverse range of topics, we apply a principled four-method validation pipeline: LaBSE cross-lingual embedding similarity, Centered Kernel Alignment (CKA), LLM-as-judge evaluation, and deep learning classifier parity tests. We show that despite shortcomings in handling slang, vulgarity, and culturally-loaded expressions, direct translation preserves subtle moral cues well enough to be harvested by cross-lingual machine learning - with a mean cosine similarity of 0.89 and classification accuracy gaps of 0.01-0.02 AUROC across foundations. These results demonstrate that machine translation is a practical and cost-effective path to moral research in languages currently under-resourced in this domain. We demonstrate this for Polish as a representative Slavic language, with expected generalization to related languages.

Sources

Related papers