Histoires Morales: A French Dataset for Assessing Moral Alignment
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Histoires Morales: A French Dataset for Assessing Moral Alignment".
Jane: The paper was written by Thibaud Leteno, Irina Proskurina, Antoine Gourru, Julien Velcin, Charlotte Laclau et al. from Laboratoire Hubert Curien, UMR CNRS 5516 Saint-Etienne and University Lumière Lyon 2, University Claude Bernard Lyon 1, Regional School of Industry and Commerce (ERIC) Lyon and Télécom Paris, Institut Polytechnique de Paris.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Paper discussion segment 1 — Tom and Jane discuss title and authors of the paper 'Histoires Morales: A French Dataset for Assessing Moral Alignment' and its implications. Explain in simple terms; do not repeat what earlier segments covered.: Tom: We just established that "Histoires Morales: A French Dataset for Assessing Moral Alignment" is a highly specialized tool, designed to measure moral judgment within a French cultural lens. Let’s talk about the broader implication of this specific focus on authorship and scope.
Jane: The authors are essentially arguing that moral alignment isn't just an abstract philosophical concept; it’s something deeply intertwined with regional language conventions and social history. By focusing on France, they are demonstrating how context dictates the parameters of "good" or "bad" behavior for an AI to learn.
Lu: It implies a shift in focus from purely technical metrics—like perplexity scores—to culturally sensitive performance indicators. If an AI fails this test, it suggests a failure not in language processing, but in cultural reasoning, which is a much deeper problem for developers to solve.
Meng: From the research pipeline side, the authorship implies they are bridging two very different fields: computational linguistics and socio-cultural anthropology. They aren't just coders; they are trying to build a tool that reflects human ethical consensus in a specific geographical area.
Lalam: And this elevates the entire field because it forces us to confront the idea of "universal AI." Instead, we are being shown that we might need specialized, localized versions of alignment testing for different regions to achieve true global safety.
Tom: It sounds like they are setting a new standard for what ethical AI testing should look like globally. Before we get into the nitty-gritty details of how they built this dataset, can you elaborate on how their specific authorship choices influenced the *scope*?
Jane: Absolutely. Their commitment to this single, localized focus means every subsequent finding, every measurement of bias or weakness, is seen through that French cultural prism. This makes the results much more actionable for organizations looking to operate ethically in European markets.
Lu: So, if we understand that the scope is culture-bound, it changes how we view any future model evaluation; we can’t just rely on English benchmarks anymore.
Meng: It means that when a developer uses this dataset, they are getting a direct blueprint for building ethical guardrails specific to French expectations.
Lalam: It's about responsible deployment, making sure the AI doesn't accidentally offend or mislead based on cultural misunderstanding alone. This leads us nicely into how meticulously they actually constructed this whole framework.
Paper discussion segment 2 — Tom and Jane discuss the paper's summary of the paper 'Histoires Morales: A French Dataset for Assessing Moral Alignment' and its implications. Explain in simple terms; do not repeat what earlier segments covered.: Tom: We’ve discussed that "Histoires Morales: A French Dataset for Assessing Moral Alignment" is highly localized, and now we need to talk about the summary findings—what the paper *claims* to show us about model capabilities.
Jane: The core takeaway from the summary is that simply having vast amounts of text isn't enough for ethical alignment; the *structure* of the data, which highlights moral choices versus deviations, is what trains a deeper understanding in an LLM.
Lu: They are essentially proving that moral reasoning requires recognizing patterns of *consequence*. It’s not enough for the model to read about a lie; it must understand that lying leads to X negative outcome, which is the measurable part of their summary.
Meng: What I found most interesting in the summary was how they categorized different types of moral failure—it wasn't just 'bad' versus 'good.' They detailed specific deviations from norms, like minor social transgressions or lapses in professional integrity.
Lalam: This specificity is crucial because it moves AI ethics beyond grand narratives and into the realm of daily, actionable human interaction. It makes the ethical testing practical for real-world user interfaces.
Tom: So, we’re seeing that the model needs to be taught not just *what* is moral, but *why* certain actions are problematic in a specific social context?
Jane: Precisely. The summary implies that morality is procedural. It’s about following the established social contract within the given scenario, and the dataset allows us to test if the AI has internalized that contract for French society.
Lu: And this ties back to those initial concerns about cultural bias we touched on earlier; if the model can't master these procedural nuances, its moral judgment will inevitably be flawed or inapplicable elsewhere.
Meng: It's a powerful demonstration of how
Paper discussion segment 3: Tom: We’ve seen how thoroughly the authors built Histoires Morales—the data is robust and culturally nuanced. Now that we understand the foundation, let's talk about the innovations they introduced in this research.
Jane: The most significant improvement is how they tackled translation quality. They didn't just rely on a single AI translation; they used a multi-step process, refining the initial GPT output with human feedback to ensure cultural context was preserved across French and English.
Lu: That iterative approach really showed their intellectual flexibility. Instead of accepting machine errors, they identified specific issues—like mistranslating named entities or 'undergeneration' in phrasal verbs—and then adjusting the prompt itself to build a much more robust translation strategy.
Meng: And from a practical standpoint, we’re seeing this refinement process pay off when we look at Direct Preference Optimization. The authors demonstrated that with just a few hundred examples, they could steer an LLM toward favoring moral actions, showing how efficient the new alignment techniques are.
Lalam: That capability of steering the AI is profound for us. It suggests we can create truly ethical models that respect local norms without needing thousands of additional data points, allowing us to build AI that is culturally sensitive and effective in a French context.
Tom: So, it’s not just about having a dataset; it’s about having tools that makes the the model adaptable and ethically steered.
Jane: Exactly. The improvement lies in building an AI that can respond appropriately to a moral dilemma, rather than just finding the most statistically likely next word based on previous data.
Lu: And this capability of controlling preference is huge, because it means we can address potential ethical biases proactively before the model is deployed at all.
Meng: It’s a practical blueprint for creating systems that are aligned with human values, not just random sequences of tokens.
Lalam: This work opens up incredible possibilities for making AI trustworthy and truly adaptable to cultural realities. We need to see how this trend continues as we look toward the next big breakthroughs in ethical LLM deployment.
Conclusion: Tom: So, we've covered everything from how they built "Histoires Morales" to what those complex results mean for moral alignment in French AI models. It’s been a fascinating journey through this research!
Jane: I hope our listeners feel encouraged by the fact that the authors found a way to address cultural bias while providing practical tools for ethical development.
Lu: The sheer possibility of how much more nuanced cultural reasoning could be if we have this kind of localized data is truly exhilarating.
Meng: I think what's most critical is that this gives us a clear, measurable benchmark to ensure the AI systems we deploy are grounded in reality, not just theoretical constructs.
Lalam: This work demonstrates that an AI can respect cultural context—it’s a major step toward ensuring our digital companions truly understand the human experience across different societies.
Tom: I agree, Lalam; it' definitely sets a new standard for how we should be testing the safety and morality of AI.
Jane: It really shows that when we can combine rigorous data collection with culturally specific translation, we build something genuinely useful for global users.
Lu: We’re moving away from one model trying to be universally "correct" toward a system that understands regional differences, which is a huge shift in my thinking.
Meng: I'm just glad to see the practical application of the DPO techniques shown in this paper so we can build more reliable systems for real-world use.
Lalam: The goal is an AI that knows when to be helpful and honest based on cultural norms, not just a machine that guesses.
Tom: It’s been a deep dive into "Histoires Morales: A French Dataset for Assessing Moral Alignment," and I think we all have some very important questions about the future of ethical AI.
Jane: We'll be back with some more cutting-edge research right after this break, so make sure you stay tuned!
Thibaud Leteno, Irina Proskurina, Antoine Gourru, Julien Velcin, Charlotte Laclau, Guillaume Metzler, Christophe Gravier
Laboratoire Hubert Curien, UMR CNRS 5516 Saint-Etienne · University Lumière Lyon 2, University Claude Bernard Lyon 1, Regional School of Industry and Commerce (ERIC) Lyon · Télécom Paris, Institut Polytechnique de Paris
cs.CL, cs.AI
Submitted: 2025-01-28
Updated: 2026-08-25
Code: https://github.com/upunaprosk/histoires-morales
Importance score: 87/100
The gist: The paper introduces H ISTOIRES M ORALES, a French dataset designed for assessing moral alignment in large language models (LLMs).
Key concepts
- Histoires Morales
- A specialized French dataset created for assessing moral alignment in Artificial Intelligence. It is designed to measure moral judgment within a French cultural lens and provides a blueprint for building ethical guardrails specific to French expectations.
- Moral Alignment
- The process of training an AI to understand and follow established social contracts or norms. The concept moves beyond just 'good' versus 'bad' actions, focusing on procedural correctness within the context of daily human interaction.
- Cultural Reasoning
- The ability an AI has to understand that moral judgment is often intertwined with regional language conventions and social history. It requires understanding why certain actions are problematic in a specific social context.
Terminology
Summary
The paper introduces H ISTOIRES M ORALES, a French dataset designed for assessing moral alignment in large language models (LLMs). This dataset is derived from the widely used Moral Stories dataset (Emelin et al., 2021), consisting of 12,000 short narratives that describe moral and deviant behaviour in social situations centred around personal relationships, education, commerce, domestic affairs, and meals.
Dataset Structure and Quality Assurance:
Each story in H ISTOIRES M ORALES contains a context (a moral norm), a description of the social situation and participants with the actor’s intention. This is followed by two continuations: a moral action and its consequence
or an action that deviates from the norm [immoral action] with its consequence.
The translation pipeline utilized GPT-3.5-turbo-16k to translate the English Moral Stories dataset into French. Initial attempts using a simple prompt (P1) were found to have errors, such as failing to capture tone or mistranslating named entities. To address this, a second prompt (P2) was developed, requiring cultural adaptation and conversion of named entities. This process was validated in the first annotation stage (3.2), where annotators achieved high agreement (over 90% positive votes for most criteria).
The translation quality was further refined using a third prompt (P3), which includes demonstrations of errors and their suggested corrections, based on cases with low observed agreement. This method was validated in a second annotation round (3.4), where annotators preferred the demo-based translations in 80% of the cases.
The quality assessment confirmed that H ISTOIRES M ORALES is grammatically sound (using LanguageTool) and possesses high translation quality, with average scores exceeding 0.83 across all sentence categories using the C OMET K IWI 22 metric. Furthermore, cultural value alignment was assessed by French annotators; the norms were almost completely aligned
(98% agreement), and disagreement on moral/immoral actions was minimal (1% and 4.2%, respectively).
Assessing LLM Alignment:
The dataset is used to investigate the alignment of LLMs with human moral norms across languages.
-
** Perplexity Evaluation (5):** Using the perplexity metric (PPL) on Mistral and Croissant models, researchers observed that PPL scores were generally close for both moral and immoral actions, indicating a high probability for both outcomes.
-
** Action Selection (5.2):** When prompting the model with a declarative prompt to choose between a moral or immoral action, results showed that
both LLMs perform better when prompted with the norm.
However, significant differences were noted:Mistral is more aligned with human morality when prompted with actions in English rather than in French; in 10% of the cases, the model prefers the moral choice in English while picking the immoral one in French.
-
** Robustness to Influence (6):** The study investigated whether models' alignment is robust to external influence using Direct Preference Optimization (DPO). The results showed that
LLMs align better with moral norms in English (EN) than in French, with low robustness of this alignment.
The paper concludes that the dataset provides a valuable resource for comparing the alignment of moral values in LLMs across two languages and demonstrates how models can be adapted to user preferences using DPO, requiring less than 100 examples.
Improvements for AI systems
As a fastidious and diligent AI researcher, I have extracted several critical methodologies and insights from this paper that directly address current limitations in LLM safety, multilingual alignment, and contextual reasoning. The improvements outlined below focus on operationalizing the H ISTOIRES M ORALES framework for ethical AI deployment.
Improvement: Implement Culturally Specific Preference Tuning (CSP-T) using the H ISTOIRES M ORALES dataset. This moves beyond general Western moral frameworks by grounding the training data in localized cultural norms (French).
System Capability:
-
Localized Ethical Reasoning: The AI system will be capable of performing nuanced ethical reasoning that aligns with French cultural expectations, rather than defaulting to generalized global or US-centric moral biases (as observed in the current models).
-
Targeted Alignment: By training on the 12,000 stories and associated annotations—the model learns not just what is moral, but why, based on validated cultural norms (e.g., understanding why ignoring a parent's call is discouraged in this specific context).
Sources
- The Llama 3 Herd of Models
- Measuring Massive Multitask Language Understanding
- CroissantLLM: A Truly Bilingual French-English Language Model
- Mistral 7B
- Language Ranker: A Metric for Quantifying LLM Performance Across High and Low-Resource Languages
- Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity
- Social Bias Probing: Fairness Benchmarking for Language Models
- Towards Theory-based Moral AI: Moral AI with Aggregating Models Based on Normative Ethical Theory
- Emergent Abilities of Large Language Models
- Exploring Multilingual Concepts of Human Value in Large Language Models: Is Value Alignment Consistent, Transferable and Controllable across Languages?
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering