Beyond English: Evaluating Automated Measurement of Moral Foundations in Non-English Discourse with a Chinese Case Study
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Beyond English: Evaluating Automated Measurement of Moral Foundations in Non-English Discourse with a Chinese Case Study".
Jane: The paper was written by Calvin Yixiang Cheng and Scott A. Hale from Oxford Internet Institute, University of Oxford.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone! Today we're digging into a paper that's got a mouthful of a title: "Beyond English: Evaluating Automated Measurement of Moral Foundations in Non-English Discourse with a Chinese Case Study." Jane, I have to say, just reading that title got me excited.
Jane: Oh, me too, Tom. And for our listeners who might be tuning in fresh, let's break down what that title actually means. The paper is about how we automatically detect moral values in text—things like care, fairness, loyalty—but doing it in languages other than English. And they use Chinese as their test case.
Tom: Right, and that's a big deal because most of the tools we have for this kind of analysis were built for English. So the authors, Calvin Yixiang Cheng and Scott A. Hale from Oxford, are basically asking: can we take these English-centric tools and make them work for the rest of the world?
Jane: Exactly. And the title "Beyond English" really captures that ambition. They're not just tweaking an existing tool; they're evaluating four completely different approaches to see which one actually works when you step outside the English-speaking bubble.
Tom: And that's what got me hooked. I mean, we all know that language shapes how we express ideas, but morality? That feels deeply cultural. So the question becomes, can a model trained on American tweets really understand what a Chinese person means when they talk about loyalty or authority?
Jane: That's the million-dollar question, Tom. And the paper doesn't just speculate—they actually test it. They use three different Chinese datasets, including real-world news sentences, and they compare machine translation, local dictionaries, and two types of language models. It's a proper head-to-head.
Tom: And spoiler alert, the results are not what you'd expect. I mean, you'd think translating everything to English and using a proven tool would work fine, right? But the paper shows that approach has some serious blind spots.
Jane: Yeah, and that's the part I can't wait to unpack. But before we get into the weeds, let's just appreciate the scope of this. The authors are trying to solve a problem that affects every non-English speaker doing social science research. If we can't measure moral values in local languages, we're missing out on understanding entire cultures.
Tom: Absolutely. And the implications go way beyond academia. Think about social media moderation, political analysis, even marketing—if you can't read the moral temperature of a conversation in its original language, you're flying blind.
Jane: So stick with us, because in the next segment we're going to look at what the paper actually found. And trust me, there's a twist involving large language models that you won't want to miss.
Summary: Tom: Alright, Jane, we've set the stage. Now let's talk about what this paper actually discovered. The summary is pretty wild—they tested four approaches and the results really shake up the field.
Jane: They do. So, the four approaches were: machine translation into English, using a local Chinese moral dictionary, fine-tuning a multilingual encoder model, and using a large language model like Llama. And the headline finding is that the old-school methods—translation and dictionaries—just don't cut it.
Tom: Yeah, I mean, the dictionary approach, which is basically counting words that match a pre-built list, scored pretty poorly. On their real-world Chinese news dataset, it only got an F1 score of around zero point four three. That's barely better than flipping a coin for some categories.
Jane: And machine translation wasn't much better. They took Chinese text, translated it to English, and ran it through a state-of-the-art English model called Mformer. It did okay on some values, but on "loyalty" it actually performed worse than random guessing. Can you believe that?
Tom: That's the part that blew my mind. Random guessing gets a score of zero point two three, and Mformer got zero point two one. So the translation process was actively destroying the cultural signals that tell you something is about loyalty. It's like the model was looking at the text and seeing nothing.
Jane: Right, and that's because translation loses so much context. Idioms, political euphemisms, even just the way people talk about family—it all gets flattened into generic English. The paper gives a great example where a Chinese phrase about faking an accident for insurance money gets translated literally, and the moral meaning just vanishes.
Tom: So then we get to the language models, and this is where it gets interesting. The multilingual encoder model, XLM-T, did better. It scored around zero point six three on the real-world dataset. But the real star was the large language model, Llama3 point 1.
Jane: Absolutely. Just using the base model with a clever prompt, Llama got a zero point four eight F1 score. But then they fine-tuned it with English data and machine-translated Chinese data, and boom—it jumped to zero point seven four. That's a massive improvement, and it didn't require any manually annotated Chinese data.
Tom: And that's the kicker, Jane. The LLM approach was not only more accurate, but it was way more data-efficient. The encoder model needed over two thousand locally annotated examples to get strong, but the LLM got there with just English data that was automatically translated.
Jane: So the summary is: if you want to measure moral values in a non-English language, skip the dictionaries and the translation pipelines. Go straight to a large language model and fine-tune it smartly. That's the recipe for success.
Tom: And that's a huge deal for researchers who don't have the budget to manually label thousands of local language texts. But there's a catch, right? The paper also found that even the best models miss cultural nuances.
Jane: Oh, definitely. We'll get into that in a bit, but first, let's talk about the specific improvements the authors suggest. Because it's not just about picking the right model—it's about how you use it.
Improvements: Tom: So Jane, we've established that LLMs are the way to go, but the paper doesn't just stop at "use a big model." They actually lay out a specific recipe for getting the best results. What did they find?
Jane: Well, the first improvement is about prompting. They found that using English prompts actually works better than Chinese prompts, even when analyzing Chinese text. At least initially. The base model scored zero point five four with English prompts but only zero point four eight with Chinese prompts on the real-world dataset.
Tom: That's counterintuitive, right? You'd think the model would understand the text better if you ask in the same language. But the paper explains that these models are just more fluent in English, so they reason better when prompted in English.
Jane: Right, but here's the clever part. Once they fine-tune the model, that gap shrinks. And if they fine-tune with both English and machine-translated Chinese data, the difference almost disappears. So the improvement is about using data augmentation to bridge that language gap.
Tom: And that data augmentation is the second big improvement. They took all the English annotated data, translated it into Chinese, and fed it back into the model. That alone boosted the F1 score from zero point six five to zero point seven four on the CCV dataset. No new human annotation needed.
Jane: Exactly. And they also recommend using a binary classification approach—one model per moral value—instead of a single multi-class model. That's a detail that might sound technical, but it really helps because each value has its own nuances.
Tom: And the third improvement is about the model size. They tested Llama3 point 1-8b and Llama3 point 1-70b, and the bigger model performed better, especially when prompted in Chinese. So if you have the compute, bigger is better.
Jane: But here's the thing, Tom. The paper is also very honest about the limitations. They found that fine-tuning with local language data can sometimes backfire. When they added small batches of Chinese annotated data to the LLM, the performance actually dropped for some values like "care" and "fairness" before it recovered.
Tom: That's a warning shot for anyone thinking "more data is always better." The paper suggests that small amounts of noisy local data can confuse the model, especially if it's already been fine-tuned on English.
Jane: Right, and that leads to their final recommendation: human-in-the-loop validation. Even the best model misses cultural nuances, especially for values like "authority" and "sanctity" that have very different meanings in Chinese versus Western contexts.
Tom: So the improvements are: use English prompts, augment with translated data, use binary classifiers, and always have a human double-check the results. That's a practical playbook for anyone doing cross-cultural research.
Jane: And it's a playbook that could really change how social science is done. But we're not done yet—let's look at the actual first page of the paper to see how they frame this whole problem.
First Page: Tom: Alright, Jane, let's zoom in on the very first page of "Beyond English: Evaluating Automated Measurement of Moral Foundations in Non-English Discourse with a Chinese Case Study." What's the setup?
Jane: So the first page lays out the core problem: moral foundation theory is this hugely influential framework in psychology that says there are five universal moral values—care, fairness, loyalty, authority, and sanctity. And it's supposed to apply to everyone, everywhere.
Tom: But the catch is, the tools to measure these values in text were all built for English. So if you're a researcher in China or Brazil or anywhere else, you're stuck. You either translate your data and lose meaning, or you build your own tools from scratch, which is expensive and slow.
Jane: Exactly. And the authors point out that this English-only focus is holding back cross-cultural research. They mention that moral foundation theory has been used to study political polarization, climate change attitudes, even terrorism. But all of that work is limited to English-speaking populations.
Tom: And that's a real problem because morality is deeply cultural. What counts as "authority" in a Confucian society might be totally different from what it means in a Western liberal democracy. If we can't measure that, we're missing the whole picture.
Jane: Right. And the first page also introduces the four approaches they're going to test. They're not just picking one method and hoping for the best—they're systematically comparing machine translation, local dictionaries, multilingual encoder models, and large language models.
Tom: And they're doing it with three different Chinese datasets. One is a set of moral vignettes, which are like little stories about right and wrong. Another is a collection of scenarios written by native Chinese speakers. And the third is real-world news data. So they're covering everything from controlled experiments to messy reality.
Jane: That's what I love about this paper—it's thorough. They're not just testing on one dataset and declaring victory. They're stress-testing across different types of text to see where each method fails.
Tom: And the first page already hints at their main conclusion: language models, especially LLMs, are the future. But they also warn that even the best models have blind spots. There's this line about "cultural misalignment" that really stuck with me.
Jane: Yeah, because it's not just about getting the right label. It's about understanding why a label is right. And the paper argues that LLMs, for all their power, still don't fully grasp the cultural context that native speakers take for granted.
Tom: So the first page sets up a really compelling research question: can we trust automated tools to measure morality across languages? And the answer, as we've seen, is a qualified yes—but only if you use the right tools and keep a human in the loop.
Jane: And that's the perfect setup for our conclusion, where we'll wrap up the whole paper and talk about what it means for the future of research.
Conclusion: Tom: Well, Jane, we've covered a lot of ground on "Beyond English: Evaluating Automated Measurement of Moral Foundations in Non-English Discourse with a Chinese Case Study." Let's bring it all together.
Jane: Let's do it. So the big takeaway is that measuring moral values in non-English text is possible, but you have to be smart about it. Machine translation and local dictionaries are not enough—they lose too much cultural information.
Tom: Right, and the paper shows that language models are the way forward. Multilingual encoder models work if you have enough local data, but large language models are more efficient and more accurate, especially when you fine-tune them with translated English data.
Jane: And the practical recipe is: use English prompts, augment your training data with machine translation, use binary classifiers for each value, and always have humans validate the results. That's the formula for getting reliable moral measurements in languages like Chinese.
Tom: But we also have to remember the caveats. Even the best LLMs miss cultural nuances, especially for values like "authority" and "sanctity" that are deeply tied to specific cultural traditions. And fine-tuning with small amounts of local data can sometimes hurt performance.
Jane: So the paper isn't saying "trust the machine." It's saying "use the machine wisely." And that's a really important distinction for researchers who want to do cross-cultural work.
Tom: Absolutely. And the implications go beyond just moral foundations. This approach could be used for any kind of deductive coding in non-English languages—hate speech detection, sentiment analysis, you name it. The methods they've tested here are generalizable.
Jane: And that's what makes this paper so exciting. It's not just a technical contribution; it's a methodological roadmap for making social science more global. If we can measure moral values in Chinese, we can do it in Hindi, Arabic, Swahili—any language.
Tom: So as we say goodbye to this paper, let's give a round of applause to Cheng and Hale for tackling a problem that's been holding back the field. And to our listeners, if you're doing research outside of English, this paper is a must-read.
Jane: Definitely. And on that note, we're wrapping up this discussion. Thanks for joining us, and we'll see you in the next episode with another exciting paper.
Tom: Take care, everyone!
Calvin Yixiang Cheng, Scott A. Hale
Oxford Internet Institute, University of Oxford
cs.CL, cs.SI
Submitted: 2025-07-22
Comments: 12 pages, 2 figures, 6 tables
DOI: 10.1609/icwsm.v20i1.42650
Code: https://github.com/calvinchengyx/cross-lan-mft-measure
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 64/100
The gist: This study explores computational approaches for measuring moral foundations (MFs) in non-English corpora.
Key concepts
- Moral Foundation Theory
- This psychological framework suggests five universal moral values—care, fairness, loyalty, authority, and sanctity—that should apply across all cultures. The challenge is that existing tools to measure these values are primarily built for English.
- Automated Measurement
- The process of using AI models to detect specific moral values within text without human intervention. The study compares four methods: machine translation, local dictionaries, multilingual encoders, and large language models.
- Large Language Models (LLMs)
- Advanced AI models like Llama3.1 that the study found to be highly effective for measuring moral foundations. They perform better than older methods because they can be fine-tuned using translated English data, even without extensive local training.
Terminology
Summary
This study explores computational approaches for measuring moral foundations (MFs) in non-English corpora. Since most resources are developed primarily for English, crosslinguistic applications of moral foundation theory remain limited. Using Chinese as a case study, this paper evaluates the effectiveness of applying English resources to machine translated text, local language lexicons, multilingual encoder-only language models, and decoder-only large language models (LLMs) in measuring MFs in non-English texts. The results indicate that machine translation and local lexicon approaches are insufficient for complex moral assessments, frequently resulting in a substantial loss of cultural information. In contrast, language models demonstrate reliable cross-language performance with transfer learning, with LLMs excelling in terms of data efficiency. Importantly, this study also underscores the need for human-in-the-loop validation of automated MF assessment, as even the most advanced models may overlook cultural nuances and face potential risks in cultural misalignment. The findings highlight the potential of LLMs for cross-language MF measurements and other complex multilingual deductive coding tasks.
Moral intuitions have long fascinated social scientists, as they help explain a wide range of cognitive and behavioral phenomena across individuals and groups (Effron and Helgason 2022). Moral foundation theory (MFT) is among the most prominent psychology frameworks for understanding the origin and development of human morality (Graham et al. 2013). Rooted in moral nativism, MFT argues there are five universal moral foundations—care/harm, fairness/cheating, authority/subversion, loyalty/betrayal, and sanctity/degradation—that transcend languages and cultures and underlie people’s moral judgments and decision-making processes (Graham et al. 2013). While some scholars propose other foundations (Haidt 2012; Atari et al. 2023a), these five foundations have received the most empirical validation across domains, languages, and cultures (Iurino and Saucier 2020).
A growing body of literature employs MFT to investigate online social behaviors. For example, MFT provides a framework for understanding rising political polarization. Individuals who prioritize loyalty, authority, and sanctity foundations are more likely to endorse conservative views and engage in polarized political discourse, while those affiliated with liberal ideologies tend to value all MFs more evenly (Koleva et al. 2012; Haidt and Graham 2007). MFT also sheds light on online cultural clashes. Atran (2007) identified differing valuations of the sanctity foundation as a key factor in many religious and ideological conflicts. Beyond polarization, MFT has been applied to a range of social issues, including climate change (Markowitz and Shariff 2012), vaccine hesitancy (Amin et al. 2017), anti-abortion views (Koleva et al. 2012), nationalism (Kertzer et al. 2014), collective violence (Nussio 2023), and terrorism (Tamborini et al. 2020).
Given its broad relevance, measuring MF values in online discourse is essential; yet, automated extraction of MF values from large-scale texts remains challenging, particularly for non-English corpora. Like other latent human values, MFs are often conveyed through abstract narratives that vary across languages. Also, most computational resources for the measurement of MFs are designed for English content (Hoover et al. 2020; Trager et al. 2022). This reliance on English hinders cross-cultural comparative research on MFT and limits theoretical advancement from non-English data (Cheng and Zhang 2023). Although MFT is intended to apply across languages and cultures, the limited availability of non-English resources severely restricts its research scope and further development of the theoretical framework (Graham et al. 2013).
In this work, we investigate various computational approaches for cross-language measurement of MFs with a particular focus on data-efficiency. We use Chinese as an example and find that (1) MF local language lexicons yield suboptimal performance. They are worse than machine translation approaches that utilize established English measurements such as Mformer. (2) Multilingual encoder-only models can achieve moderate success when trained on annotated English data along with some local language labeled data. However, this strategy is less data-efficient in measuring MFs than in other deductive coding tasks, such as hate speech detection (Röttger et al. 2022). (3) Decoder-only LLMs outperform other approaches in cross-language MF measurements in both accuracy and data efficiency. Simply fine-tuning and augmenting with English annotated data can achieve strong performance on non-English corpora; (4) nevertheless, this performance is inconsistent on different MF values, as LLMs may overlook cultural nuances in cross-language measurements, particularly for culturally distinct values.
The unique link between word usage and the expressed moral values provides a theoretical foundation for automated MF measurement from online texts (Brady, Crockett, and Van Bavel 2020; Gantman and Van Bavel 2014, 2016). Scholars have explored different computational approaches, including dictionaries (Graham, Haidt, and Nosek 2009; Hopp et al. 2021), word embeddings (Kwak et al. 2021; Araque, Gatti, and Kalimeri 2020), machine learning (Lan and Paraboni 2022), deep learning language models (Preniqi et al. 2024; Nguyen et al. 2024) and LLMs (Rathje et al. 2024). These computational methods demonstrate great advantages on scalability and labor intensity compared to traditional human annotations.
However, these approaches are primarily developed for English, and are not directly applicable to non-English texts due to several major concerns: differences in cultural contexts, the lack of annotated datasets, and limitations in domain generalizability. To address the cross-language challenges and bridge knowledge gaps in MF measurements, various computational approaches have been proposed, which can be broadly categorized into two paths: machine translation to English and the development of cross-language measurement tools (Zhuang et al. 2020).
Translation is a widely used technique in cross-language MF measurement. For instance, the MFT survey has been translated into over 20 languages for cross-lingual studies (Yilmaz et al. 2016; Nilsson and Erlandsson 2015). For large scale text analysis, the development of multilingual neural machine translation has significantly improved translation quality compared to earlier statistical methods, enhancing context understanding, ambiguity resolution, and fluency (Stasimioti et al. 2020). This advancement enables a machine-translated approach to cross-language MF measurement by translating target languages into English and applying established English-based methods (Artetxe, Labaka, and Agirre 2020).
Established methods for automatically measuring MFs in English include dictionaries (Graham, Haidt, and Nosek 2009; Frimer et al. 2019; Hopp et al. 2021), word embeddings (Kwak et al. 2021; Araque, Gatti, and Kalimeri 2020), machine learning and deep learning models (Preniqi et al. 2024; Nguyen et al. 2024) trained on annotated English-language social media data (Hoover et al. 2020; Trager et al. 2022).
Moral foundation dictionaries (MFDs) Word-count methods with crafted English moral lexicons are common. There are four common English MFDs. The original MFD is an expert-crafted dictionary containing a list of 600 words across five foundation values (Graham, Haidt, and Nosek 2009). Frimer et al. (2019) then expanded this vocabulary to MFD2 to over 2,000 words by automatically identifying similar words with word2vec word embeddings (Mikolov et al. 2013). Similarly, Araque, Gatti, and Kalimeri (2020) extended the original MFD to a MoralStrength dictionary with approximately 1,000 English lemmas based on WordNet synsets (Princeton 2010). Compared to MFD2, MoralStrength added a round of crowd-sourced ratings on the expanded lemmas. Hopp et al. (2021), however, curated a fully crowd-sourced dictionary named eMFD. It is different from previous expert-curated dictionaries for its layperson focus, contextual annotations, probability labeling and large vocabulary with 3,200 words.
Moral word embeddings Semantic similarity methods using embeddings are another approach. To address the limitations of word-count methods, such as context insensitivity and vocabulary coverage (Nguyen et al. 2024), scholars have introduced semantic similarity methods. For example, Kwak et al. (2021) proposed an embedding framework FrameAxis. It predefines a vector space of micro moral frames with two sets of opposing seed words. Target documents are then converted into vectors using word embedding models, and their MF values are determined by comparing to the micro-frames. This method has been used to extract MF values from various online texts (e.g., Mokhberian et al. 2020; Jing and Ahn 2021).
Moral language models Supervised classification models with annotated English-language training data have also been used. With advancements in language models and efforts to create human-labeled MF training datasets (Hoover et al. 2020; Trager et al. 2022), recent studies have demonstrated the potential of fine-tuning language models. For example, Preniqi et al. (2024) fine-tuned a BERT-based classifier MoralBert with large-scale annotated English data and achieved state-of-the-art performance. To address generalizability limitations in out-of-domain datasets (Liscio et al. 2022), Nguyen et al. (2024) proposed another language model Mformer, and reported superior performance compared to other established English methods in evaluations.
Although machine translation offers several advantages in cross-language MF measurement, including interpretability, scalability, efficiency and accessibility, it also faces significant limitations. First, translation quality varies across languages and domains (Ranathunga et al. 2023). In some low-resource languages like Tamil, machine translation often makes errors in translating domain terms, polysemous words, and contains repetitions for semantically similar terms (Ramesh et al. 2021). Second, it often fails to retain non-propositional information, such as emotional nuances. This can lead to a loss of emotions, toning down, or amplification across languages, introducing bias in subsequent analyses (Troiano, Klinger, and Pado 2020). Third, machine translation struggles to capture cultural elements, which is a major concern for cross-cultural and comparative research (Haidt 2012). It often shows limited performance with rare or culture-specific words, idiomatic phrases, and metaphor recognition (Dorothy, Keiko, and Vladimir 2019). Thus, it remains unclear whether machine translation is a reliable method for measuring cross-language MFs.
A second path is to develop cross-language tools, where scholars create computational MF resources tailored to local languages. Common cross-language measurements include local language dictionaries, task-specific encoder-only language models, and LLMs.
Local language dictionaries Due to the efficiency at scale and multilingual capabilities, local language dictionaries are widely used to estimate MF values from non-English texts (Hopp et al. 2021). Developing local language dictionaries generally involves three steps: (1) translating English dictionaries to target languages; (2) adding culturally specific and non-translatable vocabulary; and (3) validating with native speakers and local language corpora. Several extensive non-English MFDs have been developed and validated in Turkish (Alper et al. 2020), Japanese (Matsuo et al. 2019), Portuguese (Carvalho et al. 2020), and Chinese (Cheng and Zhang 2023). Despite the abovementioned advantages, the dictionary approach still faces the inherent limitations of general bag-of-words methods. In this paper, we use CMFD2—a Chinese MFD—as a cross-language tool to evaluate a local language dictionary approach. We also test a semantic similarity approach using the FrameAxis architecture and cross-language word embedding models.
Multilingual encoder-only models To overcome the limitations of bag-of-words methods, literature has suggested machine learning and deep learning approaches. A primary challenge with these methods is the scarcity of annotated data in local languages for model training (e.g., Ji et al. 2024; Nguyen et al. 2024). Therefore, scholars have adopted transfer learning techniques that leverage English-annotated resources for cross-language classifier development. Two major transfer learning strategies are commonly proposed: (1) machine-translating annotated English-language data into local languages to train monolingual models (Schuster et al. 2019), or (2) using annotated English-language data to train multilingual encoder-only models (Barriere and Balahur 2020). Multilingual models have demonstrated strong performance in deductive coding tasks, such as sentiment analysis (e.g., Barriere and Balahur 2020) and hate speech detection (e.g., Röttger et al. 2022). Nevertheless, they also have shown language bias in morality classification tasks, as pretrained multilingual encoder-only language models often display distinct moral directions across languages (Hämmerl et al. 2023). This paper focuses on the second transfer learning strategy and evaluates multilingual encoder-only models for cross-language MF measurement.
Large language models (LLMs) The rise of decoder-only LLMs provides an alternative for cross-language MF measurement. LLMs show exceptional zero/few-shot learning capability, enabling them to directly label human values out-of-the-box, which is particularly valuable for tasks with limited human annotated data (Ziems et al. 2024). They also have demonstrated strong performance in measuring various human values, including moral reasoning tasks (Ziems et al. 2024; Agarwal et al. 2024). Notably, LLMs sometimes are not as good as specialized fine-tuned language models (Amin, Cambria, and Schuller 2023; Preniqi et al. 2024), which may be due to the lack of explicit, colloquial definitions of the target human values (Ziems et al. 2024). MFT’s well-established conceptual framework may help address this limitation. Not only are LLMs pre-trained on rich MFT literature (e.g., Abdulhai et al. 2023), but MFT also offers clear guidance for crafting clear and effective prompts. Additionally, LLMs trained on vast multilingual data exhibit promising capabilities to handle cross-language measurements (Ahuja et al. 2023).
Despite the strengths in accessibility, efficiency, multilingualism, and reasoning, there are also some concerns in using LLMs’ in cross-language MF measurement. First, LLMs exhibit a substantial degree of subordinate multilingualism, displaying proficiency in some languages but not others (Zhang et al. 2023), which has a strong correlation with the proportion of those languages in the pre-training corpus (Li et al. 2024). Second, there are potential language biases, particularly in human-value relevant coding tasks (Kirk et al. 2024; Yu et al. 2024). For example, non-English prompts are more likely to generate malicious responses compared to English prompts (Shen et al. 2024). Third, different LLMs exhibit varying baseline moral tendencies (Ji et al. 2024). For instance, GPT-3’s MF preferences align more closely with politically conservative individuals when minimal prompt engineering is used in zero-shot learning (Abdulhai et al. 2023). These moral tendencies, however, are sensitive to prompting, with different prompting strategies significantly influencing classification outcomes in MF measurements (Abdulhai et al. 2023; Agarwal et al. 2024). Thus, it is unclear how LLMs performs in cross-language MF measurement tasks. This paper selects a cutting-edge, open-source model—Llama3.1 (Meta 2024) to evaluate on LLM approach.
We used three human annotated datasets—moral foundation vignettes (MFV), Chinese moral scenarios (CCS) and Chinese core values (CCV), to benchmark the performance of different cross-language MF measurement approaches. We also used three English annotated MFT datasets for model training and fine-tuning.
MFV is a list of social behaviors constructed by psychologists based on MFT, describing the violations of specific moral values from a third-party perspective (Clifford et al. 2015). It has been widely validated and used to assess measures of moral judgment (e.g., Graham, Haidt, and Nosek 2009; Kivikangas et al. 2021; Ji et al. 2024). We used the MFV as an expert-crafted benchmark stimulus to evaluate the baseline performance of different moral foundation measurement approaches. Since the original MFV is in English, we followed a careful translation process to ensure cultural neutrality in Chinese. First, a native Chinese speaker translated the vignettes with minor modifications to preserve cultural appropriateness. Then, two additional native speakers reviewed the translations to assess whether the scenarios remained representative and meaningful in Chinese cultural contexts. This ensured that the translated vignettes could serve as a culturally neutral benchmark for cross-method comparability.
CCS is list of moral scenarios written by 202 native Chinese to describe their intuitive understandings of MF values (Cheng and Zhang 2023). This reverse-annotation method is commonly used for validating MF measurements (e.g., Cheng and Zhang 2023; Frimer et al. 2019; Matsuo et al. 2019). It incorporates culturally specific content, reflecting native speakers’ natural and intuitive understanding of MFT in a real-word context.
CCV is a human-annotated real-world dataset, including 6,994 sentences collected from four local news websites. The dataset is annotated by three native Chinese speakers based on the Chinese core socialist moral value coding scheme, which is highly correlate with the five universal moral foundation values (Liu et al. 2022). The original CCV dataset includes eight labels, we employed five Chinese native speakers with postgraduate degrees to re-label the action categories in CCV to five moral foundation values, excluding vice and virtue. The mapped values were determined by majority vote. We sample 20% of the CCV as the primary benchmarking dataset to evaluate the performance across approaches stratifying on the values. The remainder is used as training data to test the data-efficiency fine-tuning language models.
We use Google Translate as an example of a machine translation approach due to its accessibility and consistent performance across domains. First, we machine translate benchmarking datasets except MFV to Chinese using the Google Cloud API—Basic Translation service. Then we estimate the MF values from translated documents with established English measurements, including lexicons MFD (Graham, Haidt, and Nosek 2009), MFD2 (Frimer et al. 2019), eMFD (Hopp et al. 2021) and MoralStrength (Araque, Gatti, and Kalimeri 2020); word embeddings with FrameAxis (Kwak et al. 2021); and specialized-fine-tuned language models MoralBert (Preniqi et al. 2024) and Mformer (Nguyen et al. 2024).
For MFD, MFD2, and eMFD, we calculate word frequencies using the eMFDscore Python package (Hopp et al. 2021). For MFD and MFD 1.0, each document’s MF value is determined by the most frequent MF class in the respective dictionary. A document is mapped to multiple classes if there is an equal number of class matches. If there are no matching words, no class is assigned to the document. eMFD, however, assigns probabilities to its vocabulary, representing their likelihood of being associated with certain MF classes. We sum the probabilities and label the document with the class that has the highest sum.
We use the moralstrength package (Araque, Gatti, and Kalimeri 2020) for the MoralStrength dictionary. We first test its performance of the lexicon features alone with bag-of-word methods; then we train a Support Vector Machine (SVM) model and combine its lexicon features. Since English training sets contain many non-moral labels, while the benchmarking dataset contains only moral labels, we train two SVM models: one with the full training data and another with only moral-labeled training data.
For word embedding methods, we use the FrameAxis Python package (Kwak et al. 2021) to compute anchor micro-frames based on different MFDs with the word2vec embedding model (Mikolov et al. 2013). For each MFD, we generate the corresponding micro moral frames from the vocabulary in its class. We then compute and aggregate word contributions to each microframe in the document, and label the document’s moral class by identifying significant microframes through comparison with a null model.
For language models, we apply pre-trained language models MoralBert and Mformers from HuggingFace with their default settings and no additional fine-tuning.
We apply C-MFD2 (Cheng and Zhang 2023) to evaluate the performance of the local language dictionary approach. We test two techniques for locally-developed MF lexicons: word counts and embeddings. For the word embedding method, we test two approaches with the fastText model (Grave et al. 2018). One involves measure simple semantic similarity. Words in C-MFD2 are grouped into five pseudo-documents based on their MF labels, each serving as anchor frames. MF values are then determined by calculating the semantic distance between the text to be classified and these anchor frames. The other uses FrameAxis to construct anchor frames. As C-MFD2 does not have the virtue/vice dimension, which is essential to calculate micro-frames in FrameAxis, we automatically assign this dimension using the RoBERTa-based Chinese sentiment model c2-roberta-base-finetuned-dianping-chinese from HuggingFace.
We select XLM-T as the base model to test transfer learning on the multilingual encoder-only model approach (Barbieri, Anke, and Camacho-Collados 2022). Fine-tuned on 198 million multilingual tweets on the XLM-RoBERTa architecture—originally trained on 2.5 TB of Common Crawl data (Conneau et al. 2019)—this model is particularly well-suited for analyzing online content. Also, previous research indicates that it has strong performance in cross-language deductive coding tasks such as hate speech detection, compared to other encoder-only models (Röttger et al. 2022).
We follow Nguyen et al. (2024)’s experience on fine-tuning Mformer with some tweaks. First, we replace the base architecture from RoBERTa-base (Liu et al. 2019) with twitter-xlm-roberta-base and tokenization is handled by the model’s built-in tokenizer (Barbieri, Anke, and Camacho-Collados 2022), with token sequences truncated to a maximum of 512. Second, we set the learning rate, epochs, and batch size to 2e − 5, 3, and 16 respectively following Röttger et al. (2022). Third, we opt to fine-tune five binary classifiers rather than a single multi-label classifier because binary models generally outperform multi-label models in English MF measurements (Nguyen et al. 2024). Fourth, we adopt a conservative under-sampling strategy in the English annotated training dataset to address class imbalance, establishing a baseline for future improvements. After fine-tuning with English annotated data, we further fine-tune each base model with the 80% of CCV dataset in order to test the data-efficiency as in Röttger et al. (2022). We incrementally train models with batches of additional CCV data, with each batch containing 100 annotated records.
For the decoder-only LLMs approach, we select the Llama3.1-8b instruct model, an open-source LLM developed by Meta with a reasonable balance between model performance and computational cost. Compared to closed-source models like GPT-4, Llama3.1 offers greater control, flexibility, transparency, and reproducibility—all of which are important in human-value measurement tasks.
We first test LLMs with prompt-engineering and few-shot learning. Next, we apply the same fine-tuning process used for XLM-T to Llama3.1-8b. Notably, an additional round of data augmentation is performed afterwards, where all English annotations are machine-translated into Chinese using the Google Translate API. Fine-tuning is conducted with the unsloth package using 4-bit quantization on a single NVIDIA L40S GPU. Additionally, we test data efficiency with 20 batches of Chinese annotated items from the CCV dataset. A conservative under-sampling strategy is applied as well with each batch containing 50 records evenly distributed across the five classes.
We conducted qualitative analysis to gain in-depth understanding of the cultural loss in moral foundation measurements across different approaches. By randomly sampling 100 mislabeled records, we compared predicted labels to ground truth (i.e., human annotated labels), focusing on MFs that are known with cultural distinctions between English and Chinese (i.e., authority, loyalty, and sanctity).
As shown in Table 2, the performance of machine translation generally fall short in evaluation. With the MFV benchmarking dataset, the lexicon method MFD2 shows the best performance with a weighted F1 score of 0.60. In the reverse-annotated CCS dataset, a simple SVM model with lexicon features outperforms other measurements (F1 = 0.74), but the model coverage is relatively low at only 23%. In the real-word CCV dataset, the deep learning model Mformer exhibits the strongest performance (F1 = 0.47) and maintains comparable results across the other two benchmark datasets. Note that although some MF classes in Mformer displayed good performance (i.e., care/harm, F1 = 0.72), some fine-grained MF measurements are very poor. For example, Mformer’s prediction on “loyalty” (F1 = 0.21) is worse than random guessing baseline (F1 = 0.23) in the cross-language evaluation setting.
Local language lexicon approaches demonstrate similarly moderate performance with word-count methods. Table 3 shows that across all benchmarking datasets, C-MFD2 consistently outperforms other methods although its performance is generally still poor. In the real-world CCV dataset, C-MFD2 achieves the best performance (F1 = 0.43), which is comparable to the machine translation approach with Mformer. We find multilingual word embedding model fastText with FrameAxis framework do not improve the cross-language measurement performance in this task.
The multilingual encoder-only model XLM-T, trained on the same English annotated data, shows a moderate but reduced performance compared to the monolingual English model Mformer. In Table 4, XLM-T shows moderate performance across all benchmarking datasets (F1MFV = 0.71, F1CCS = 0.71, F1CCV = 0.63) outperforming both machine translation and local lexicon approaches. In addition, its performance is generally consistent across the five foundation values, showing a better reliability in the fine-grained MF measurements. This consistency is likely due to training five separate classifiers. In the real-world CCV dataset, the fine-tuned XLM-T model has moderate performance in four out of five foundation values with an average F1 score of 0.63. The “authority” model has lower performance with F1 = 0.46.
In the follow-up evaluation of batch training with local language data, we observe improved model performance as the volume of local language training data increases. As shown in Figure 1, the fine-tuned XLM-T models achieve higher F1 scores with more locally labeled data. For example, the F1 score of the “loyalty” model rose from 0.62 to 0.80 with 22 batches of local-language data. Similar increasing patterns are observed across all five models.
However, the amount of annotated non-English data required to reach reliable classification thresholds for MFs is much more than for hate speech detection (Röttger et al. 2022). On average, 16 batches are needed to reach the F1 score of 0.70 and 49 batches to reach 0.80. The “care” model, which is the best performing model of the five, requires six additional batches to reach an F1 score of 0.70 and 27 batches to reach 0.80. There are 10 batches of data available for the “sanctity” model, and it shows no significant improvement even once all 10 batches are used for training.
It is worth noting that limited local language annotated data may also negatively impact the model’s performance. For example, the “fairness” model’s F1 score initially drops from 0.50 to 0.46 when fine-tuned with local-language data. It only begins to stabilize and improve after 11 batches.
In summary, while moderate performance can be achieved with local-language annotations, considerably more data is required to attain robust classification performance with the multilingual encoder-only model approach in cross-language MF measurements.
Table 5 shows the results of the LLMs approach. With only prompt engineering, Llama3.1-8b already performs better than the machine translation and local language lexicon approaches across all benchmarking datasets. After fine-tuning with annotated English-language data, it immediately achieves strong performance on the MFV (F1 = 0.81) and CCS (F1 = 0.82) datasets, and moderate performance on the CCV (F1 = 0.65) dataset. This performance exceeds that of XLM-T with the same English-language training data (F1 = 0.63).
The model is fine-tuned with a data augmentation strategy where the English labeled data are machine translated into Chinese and fed into the LLM again. This step significantly improves the model performance and results in the best performance for all three benchmarking dataset (MFV F1 = 0.86; CCS F1 = 0.82; CCV F1 = 0.74). These F1 scores exceed all other approaches tested in this paper.
Additionally, we notice that English prompts generally outperform non-English prompts on Llama3.1-8b, even when analyzing non-English documents. This difference is more pronounced in the base instruct model but becomes less significant with progressive fine-tuning using annotated data. After fine-tuning with English data, the difference in the F1 scores for English and Chinese prompts on the CCV dataset decreases from 0.11 to 0.08, further reducing to 0.03 with data augmentation and 0.02 when fine-tuned with both English and translated Chinese data.
Moreover, a better LLM base model may further improve the model performance. Table 6 shows that under the same condition, Llama3.1-70b outperforms Llama3.1-8b when prompting in local languages. In the CCV dataset, under a few-shot learning setup with Chinese prompting, the larger LLM (F1 = 0.60) is better than both its smaller counterpart (F1 = 0.42) and the English prompting counterpart (F1 = 0.51).
We selected MFormer and fine-tuned Llama3.1-8b-instruct with English and translated Chinese data as models to analyze cultural loss in translation and non-translation approaches. We identified three common types of information loss during the translation process (See Table 4 in the Appendix). First, idioms and slang are often mistranslated, as in Example 1, where the Chinese phrase meaning “faked or staged accident for compensation” is simply translated to its superficial meaning. Second, contextual meaning. As illustrated in Example 2, where the English translation “look towards” misses the cultural connotation of “favoring someone,” resulting in different expressed moral foundations. Third, political euphemisms lose their cultural context, as in Example 3, where the English translation fails to convey that the text refers to civil servant behavior. These information loss in the translation process often result in mislabeling in MF measurement.
We further analyzed the non-translational approach to understand how LLMs process cultural nuances. Compared to the translational approach, LLMs better recognize slang and idioms (Examples 2 and 13), but demonstrate more subtle cultural misalignment. Our analysis revealed several distinct patterns. Most notably, LLMs failed to distinguish certain Confucian morality concepts such as filial piety, which encompass family responsibilities to respect and care for elders (Examples 6, 8, 9, and 10). Native Chinese speakers tend to labeled these instances as “authority,” whereas language models often only categorized them as “care.” Additionally, we observed misalignment regarding government and local political entities. Human annotators tended to label cases involving government (Examples 3, 4) and officials (Example 11) as “loyalty,” while LLMs classified them as “authority” or “fairness.” Finally, examining misalignment in “sanctity” revealed that Chinese annotators frequently associated sanctity with corrupt politicians or role model civil servants (Examples 11 and 12), a distinction rarely captured by LLMs.
We find traditional approaches, such as machine translation and local language lexicons, may not be the most effective solutions for cross-language moral foundation measurements. Instead, fine-tuning multilingual encoder-only models and leveraging LLMs through transfer learning emerge as promising alternatives. Notably, LLMs demonstrate strong performance and data efficiency, making them particularly well-suited for addressing the challenges of cross-language MF analysis.
Above all, simple machine translation approach has proven suboptimal in cross-language MF measurement. This issue is particularly concerning for measuring culturally significant values. Machine translation often introduces bias in the rendering of slangs, contextual meanings, and political euphemisms into local languages. Moreover, when relying on pretrained English classifiers such as MFormer—which is predominantly trained on English-annotated data—this cultural information loss may be further amplified. For instance, MFormer’s classification performance on the “loyalty” foundation in translated data (F1 = 0.21) falls below random chance (F1 = 0.23), indicating a substantial loss of cultural nuance. Nevertheless, it is important to note that machine translation models are rapidly evolving. Current deficiencies—such as loss of contextual nuance or culturally embedded moral cues—limit their effectiveness in this domain. Advances in context-aware and domain-adapted translation models (Jin et al. 2023; Saunders 2022), particularly those fine-tuned on moral discourse, may mitigate some of these issues, which deserves future exploration.
Researchers should also be cautious with local language lexicons—they perform worse than the machine translation for cross-language MF measurement in some cases. Given the inherent limitations of lexicon-based methods and the extensive resources required to develop them, they are inefficient for cross-language MF measurement, particularly at the document level where semantic complexity is higher. However, culturally specific moral lexicons may still provide valuable insights at the vocabulary level for cross-cultural research, though broader applications should be approached with careful validations.
Fine-tuning LLMs is more data-efficient than fine-tuning multilingual encoder-only models for cross-language MF measurement. In the CCV benchmarking dataset, XLM-T requires on average more than 2,000 local-language annotated records (20 batches) to achieve strong performance (F1 = 0.75). In contrast, Llama3.1-8b can reach comparable performance with only English data machine-translated to Chinese and thus no local language labeling. Our findings also suggest that larger LLMs such as Llama3.1-70b may perform better for measuring MFs in non-English corpora. That is, updating the base LLM to a more power model with enhanced multilingual capabilities could further improve the performance of the LLM approach.
Nevertheless, several limitations of using LLMs for cross-language MF measurement should be acknowledged. Firstly, the performance on fine-grained MF values can vary. Although the measurement on some culturally sensitive values like “loyalty” is relatively reliable, others like “authority” and “sanctity” still miss significant cultural nuances.
Second, fine-tuning LLMs might backfire. Fine-tuning with limited local-language data may degrades model performance. While training with local language annotated data is commonly used to incorporate cultural specificity and enhance transfer learning, this approach appears less effective for LLMs in MF measurements task when data volume is small. We fine-tuned Llama3.1-8b using 20 batches of CCV data with the same strategy used in XLM-T. Each batch containing 50 MF labels evenly distributed across five classes. As shown in Figure 2, the initial drop in model performance struggles to recover within the available data for ”care”, ”fairness”, and ”sanctity” values.
In addition, fine-tuning with English-annotated data may introduce extra cultural bias to LLMs. We replicated the same finetuning strategy to an Italian dataset “moralConvITA” (Stranisci et al. 2021), and observed similar significant performance drops in culturally-nuanced values like ”loyalty” and ”sanctity” (see Appendix). This signals the risk of cultural bias in English-centric LLMs. Finetuning such LLMs with English annotation data may exacerbate their bias against non-English languages, resulting in more cultural information loss in the measurement task, which echoes concerns about LLMs being shaped by values associated with WEIRD populations (Atari et al. 2023b). These findings highlight the need for culturally-sensitive fine-tuning approaches that preserve cross-cultural moral nuances.
Third, we note that our experiments were conducted in a single non-English language, which may limit the generalizability of our findings. Nevertheless, we hope this work can serve as a foundation for future research that extends to a broader range of non-English languages on MF measurement, particularly low-resource ones, where such methodological advancements would be especially valuable.
Finally, we highlight the essential role of human validation in measuring cross-language MFs with LLMs. Given the limitations discussed, human validation is strongly recommended to ensure the reliability of LLM-based cross-language MF measurements. LLMs can effectively assist in this process by providing rationales alongside their classification outputs, supporting human evaluators and facilitating a more efficient result assessment.
This study examined four computational approaches—machine translation, dictionary, multilingual encoder-only models and decoder-only LLMs—for automatically measuring moral foundation values in non-English texts. It uses Chinese as a case study and leverages established English resources for this cross-language deductive coding task. It first highlights the limitations of dictionary and machine translation approaches: while local language dictionaries can support lexicon-level analysis, they often lack the depth needed for complex semantic assessments. Notably, advanced English-based tools applied to machine-translated data outperform local lexicon-based methods, which underscores the limitations of lexicon approaches in this task. The study then explores the potential of transfer learning in language models, showing that both multilingual encoder-only models and LLMs demonstrate strong performance, with LLMs performing better and being more data-efficient. It is recommended to select approaches based on the availability of local-language annotated data. When sufficient data is available, a smaller multilingual language model generally yields satisfactory results. Otherwise, LLMs can serve as a reliable tool.
We recommend the following steps for applying LLMs to cross-language MF measurement: (1) start with multilingual LLMs and use culturally specific prompt engineering; (2) adopt a binary classification approach for each moral foundation; (3) fine-tune models using available English annotations alongside carefully translated local-language data, while curating out English-centric value cases; and (4) incorporate human validation, especially for culturally distinctive values.
Findings in this paper provide valuable insights for cross-cultural MF research and shed light on future applications of LLM-assisted deductive coding in multilingual tasks.
Improvements for AI systems
Based on the scientific paper, here are the specific improvements I can make to AI systems and what the improved system can do:
Improvement: Implement a fine-tuned Llama3.1-8b (or larger) model with a dual-language fine-tuning strategy (English + machine-translated local language data) for moral foundation classification in non-English texts.
What the improved system can do:
-
Classify Chinese, Italian, and other non-English texts into five moral foundations (care, fairness, loyalty, authority, sanctity) with F1 scores of 0.74–0.86, outperforming machine translation + English-only models by 15–20%.
-
Use culturally-aware prompting (e.g.,
consider the Chinese cultural context
) to reduce misclassification of Confucian concepts like filial piety (which should map toauthority
but often gets labeledcare
). -
Distinguish between government loyalty vs. authority in political texts—a nuance that current models miss.
The improved AI system can:
-
Measure moral foundations in non-English texts with near-human accuracy (F1 0.74–0.86) without requiring large local-language annotated datasets.
-
Preserve cultural nuances that machine translation and English-only models lose, particularly for Confucian, political, and religious concepts.
-
Self-audit for cultural bias and flag high-risk predictions for human review.
-
Deploy rapidly to new languages with minimal annotation effort, making cross-cultural moral research scalable and accessible.
Abstract
This study explores computational approaches for measuring moral foundations (MFs) in non-English corpora. Since most resources are developed primarily for English, cross-linguistic applications of moral foundation theory remain limited. Using Chinese as a case study, this paper evaluates the effectiveness of applying English resources to machine translated text, local language lexicons, multilingual language models, and large language models (LLMs) in measuring MFs in non-English texts. The results indicate that machine translation and local lexicon approaches are insufficient for complex moral assessments, frequently resulting in a substantial loss of cultural information. In contrast, multilingual models and LLMs demonstrate reliable cross-language performance with transfer learning, with LLMs excelling in terms of data efficiency. Importantly, this study also underscores the need for human-in-the-loop validation of automated MF assessment, as the most advanced models may overlook cultural nuances in cross-language measurements. The findings highlight the potential of LLMs for cross-language MF measurements and other complex multilingual deductive coding tasks.
Sources
- Moral Foundations of Large Language Models
- Unsupervised Cross-lingual Representation Learning at Scale
- The Llama 3 Herd of Models
- The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models
- Language Ranker: A Metric for Quantifying LLM Performance Across High and Low-Resource Languages
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- The Language Barrier: Dissecting Safety Challenges of LLMs in Multilingual Contexts
- The Moral Foundations Reddit Corpus
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering