Morality is Contextual: Learning Interpretable Moral Contexts from Human Data with Probabilistic Clustering and Large Language Models

arXiv:2512.21439 · cs.CL, cs.AI, cs.LG · Submitted 2025-12-24 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Morality is Contextual".

Jane: Moral actions are judged by their context,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, let's talk about who put this together. The paper introduces COMETH—Contextual Organization of Moral Evaluation from Textual Human inputs—and they use a combination of probabilistic learning and LLMs to figure out these moral contexts. The authors are from the Morlat Institute of Intelligent Systems and Robotics at Sorbonne University in Paris, which suggests a strong background in the intersection of AI and complex systems.

Jane: It’s interesting that they brought together human moral evaluations with LLM abstraction; it shows they recognize that we need both real-world data and powerful language models to get a complete picture of how context shapes judgment.

Lu: The authors clearly focused on making the framework interpretable, which is crucial when you are dealing with something as sensitive as morality in an AI context; transparency matters a lot.

Meng: I’m curious about the scale of their human data input; if they only used a small set of judgments, how robust would that clustering be when trying to apply it to novel situations?

Lalam: The authors' goal seems to be creating a system that can learn these contexts autonomously, which suggests they are aiming for something adaptive rather than just applying a static set of rules.

The paper's summary: Tom: In terms of the core research, the paper describes how they took three hundred scenarios involving actions like violating 'Do not kill' or 'Do not deceive,' and used those to build a framework that identifies moral contexts based on human responses, which are judged as blame, neutral, or support <ref:2512.21439#pg0>. They then use a probabilistic context learner with adding and merging modules to group scenarios into distinct moral contexts.

Jane: That process of grouping scenarios isn't just about putting them in buckets; it’s about the learner autonomously refining those groups online by looking at how the human judgments distribute across different contexts, which is quite sophisticated.

Lu: The paper highlights that they use an LLM-based module to extract descriptive contextual features from these clusters and then binarize them into feature vectors, which is a clever way to make the context more machine-readable while keeping the original human judgment distributions intact.

Meng: I see the methodology focusing heavily on those modules; specifically how they use Kullback-Leibler divergence for assignment and semi-weighted Jensen-Shannon divergence for merging, which shows a very deliberate design choice for managing redundancy.

Lalam: What strikes me is that they are not just classifying actions; they are modeling the *distributions* of moral judgments within these contexts, which gives the AI a much richer understanding of the uncertainty involved in moral decisions.

The paper's improvements: Tom: The main improvement they highlight is how COMETH moves beyond just looking at outcomes to explicitly modeling how context shapes acceptability, and it achieves this by integrating the probabilistic context learner with semantic abstraction and human evaluations. This gives us a much better picture of moral reasoning.

Jane: It’s about making AI systems better at contextual moral reasoning; instead of just saying 'lying is bad,' they can figure out *why* lying might be acceptable in one specific situation versus another, which is vital for real-world application.

Lu: The use of LLM-based feature extraction to create interpretable binary features, and then learning the importance weights via a likelihood-based model, offers a way to get these complex contextual relationships into something that we can actually understand and trace back.

Meng: I think the improvement in alignment is significant; they show roughly double the alignment with majority human judgments compared to using end-to-end LLM prompting alone, which suggests their structured approach yields much more reliable moral predictions.

Lalam: That improved reliability is key because it means we are building systems that respect human moral intuition better, which speaks directly to the value alignment problem Russell and Norvig talked about.

Conclusion: Tom: So, to wrap things up on "Morality is Contextual: Learning Interpretable Moral Contexts from Human Data with Probabilistic Clustering and Large Language Models," COMETH provides a concrete framework for modeling moral contexts by combining probabilistic learning with LLM features to get interpretable results. It really shows that context is the key variable in moral judgment.

Jane: Exactly, Tom. The implication is that we can start building AI agents capable of nuanced decision-making in complex social or legal situations where simple rules just don't cut it anymore because morality is so situational.

Lu: This work opens the door for more creative ways to structure how we feed moral data into models, showing that combining clustering with semantic understanding can lead to novel forms of context modeling.

Meng: For practical impact, the ability to get higher alignment rates means these systems are more trustworthy when deployed in areas where moral judgment is critical, like autonomous decision support tools.

Lalam: I feel the biggest win here is creating a mechanism that allows us to see *why* an AI made a certain judgment by showing us which contextual features were most important for that specific moral context.

Tom: And with that, we wrap up our discussion on this fascinating paper today. We’ll be right back after the break with more exciting research from arXiv!

Geoffroy Morlat, Marceau Nahon, Augustin Chartouny, Raja Chatila, Ismael T. Freire*, Mehdi Khamassi*

Institute of Intelligent Systems and Robotics, Sorbonne University

cs.CL, cs.AI, cs.LG

Submitted: 2025-12-24

Updated: 2026-10-01

Importance score: 82/100

The gist: Moral actions are judged by their context, and this framework models how context shapes the acceptability of ambiguous actions by integrating a probabilistic context learner with LLM-based semantic

Key concepts

Ternary Human Moral Evaluations
Participants are asked to judge scenarios using three labels: Blame, Neutral, or Support. This data is used as the ground truth to train the model on how different situations elicit varying moral responses. These judgments are crucial for defining what constitutes a specific moral context.
Probabilistic Context Learner
This component autonomously infers and refines 'moral contexts' by comparing new scenarios against existing models using mathematical divergence measures like KL divergence. It assigns new situations to the closest established context or merges similar ones, dynamically building a map of moral situations.
Generalization Module
This module takes the identified clusters and uses an LLM to extract concise, non-evaluative binary features that describe each context. By training this module to predict true cluster assignments, it learns which contextual features are most important for making a final moral prediction.
Kullback-Leibler (KL) Divergence
KL divergence is a mathematical tool used to measure the difference between two probability distributions—in this case, the distribution of human judgments. It helps the Probabilistic Context Learner determine how much a new scenario's outcome distribution differs from an existing moral context model.

Terminology

Summary

Moral actions are judged by their context, and this framework models how context shapes the acceptability of ambiguous actions by integrating a probabilistic context learner with LLM-based semantic abstraction and human moral evaluations.

The gist: COMETH is a novel framework that integrates empirical moral judgment data with a Probabilistic RL architecture designed to infer context-specific reward models from ternary human moral evaluations (blame, neutral, support) to model how context shapes the acceptability of morally ambiguous actions.

Data Collection and Preprocessing

The research curated an empirically grounded dataset of 300 scenarios across six core actions (violating Do not kill, Do not deceive, and Do not break the law) and collected ternary judgments (Blame/Neutral/Support) from N=101 participants. To standardize actions, a preprocessing pipeline standardizes actions via an LLM filter to extract the principal action in a uniform “to + verb + complement” format. These representations were then embedded using allMiniLM-L6-v2 Sentence Transformer and clustered with K-means, producing robust, reproducible coreaction clusters.

Probabilistic Context Learning

The Probabilistic Context Learner's objective is to autonomously infer and refine clusters online—referred to as moral contexts—by identifying patterns in ternary outcome distributions that reflect human moral judgments. This process uses two main components: the adding module and the merging module. The adding module compares a new scenario’s reward distribution to existing context models using Kullback-Leibler (KL) divergence; if the minimal KL divergence is below a threshold ∆a, the scenario is assigned to the closest context. The merging module uses semi-weighted Jensen-Shannon divergence (swJS) to assess similarity between models and merges them if their divergence falls below a threshold ∆m, thereby reducing redundancy while preserving a diverse representation of moral scenarios.

Generalization and Interpretability

The Generalization module is introduced to extract concise, non-evaluative binary contextual features from the clusters. It uses an LLM to infer descriptive contextual features that characterize each cluster. These are encoded as binary statements, and the module trains by minimizing the negative log-likelihood of the true cluster assignments, learning feature importance weights. This allows for a prediction of moral judgment by selecting the most probable label from the assigned cluster distribution, providing an interpretable alternative to end-to-end LLMs.

Empirical Performance and Comparison

The empirical evaluation demonstrated that COMETH roughly doubles alignment with majority human judgments relative to end-to-end LLM prompting (≈ 60% vs. ≈ 30% on average). This improvement is attributed to the structured feature representations. The pipeline showed that Mistral 8B and Qwen 80B reach the highest mean alignment rates (0.63), confirming that grounding predictions in structured action representations substantially improves generalization compared to end-to-end LLM approaches. Furthermore, the system provides interpretability by assigning weights to features within each cluster, clarifying how individual features contribute to human moral judgments.

Ablation and Robustness

A series of ablation tests assessed pipeline robustness. One study examined the impact of fixing the number of clusters in K-means, showing that while fixed-k setting yielded stable results, it resulted in a coarser representation compared to the Probabilistic Context Learner's dynamic clustering. Another test confirmed that LLM-based reformulation is not strictly required to recover the underlying core actions for this dataset, suggesting robustness across different structural patterns. The framework was found to be robust to variations in both cluster-number selection and the use of LLM-based reformulation.

Conclusion and Limitations

COMETH offers a concrete and practical path toward building AI systems that can better distinguish and model moral contexts in a structured, transparent, and human-aligned manner. Limitations include reliance on probabilistic representations of moral preference distributions, requirements for human survey data, and evaluation using a majority-label decision rule. Future work aims to scale the framework by integrating mechanisms to automatically select and weight the most informative features or adding active learning components.

How it works

  1. Data Collection and Preprocessing: The researchers curated an empirically grounded dataset of 300 scenarios across six core actions and used an LLM filter to extract the principal action into a uniform format, followed by K-means clustering on embeddings to produce Core Action categories.

Improvements for AI systems

Here are specific improvements to AI systems based on the COMETH framework, along with what those improved systems can achieve:


  1. The proposed system integrates a probabilistic context learner (COMETH) with LLM semantic abstraction and human moral evaluations. This allows for the creation of models that explicitly capture how context shapes the acceptability of ambiguous actions, moving beyond rigid, outcome-only morality.

  2. The system can perform contextual moral reasoning. Specifically, it can distinguish between two scenarios where the action is identical (e.g., to lie) but judged differently based on context (e.g., lying to support vs. lying for self-interest). This capability enables AI agents to make nuanced decisions in complex social or legal situations where contextual nuances are critical for ethical compliance.

  3. The system provides a transparent, interpretable framework through a Generalization module that extracts concise, non-evaluative binary contextual features and learns feature weights. This allows the AI to provide explainable moral reasoning by explicitly stating which situational factors (e.g., presence of armed attacker, explicit consent given) drive its moral prediction for a specific action/context combination.

  4. The system can achieve significantly higher alignment with majority human judgments (approximately doubling performance compared to end-to-end LLM prompting). This means the AI's moral decisions are more reliable and trustworthy because they are grounded in empirically validated human preferences rather than opaque model outputs.

  5. The system offers improved robustness across different LLMs, demonstrating that structured feature representations can mitigate scale disparities between smaller and larger models, making robust moral alignment feasible even with lighter AI architectures.

  6. The system supports dynamic context modeling through online learning (adding and merging modules). This allows the AI to continuously refine its understanding of moral contexts as it encounters new scenarios, ensuring its moral framework adapts to evolving social norms or emerging dilemmas without requiring complete retraining.

  7. The system enables better scenario classification and risk assessment by identifying stable, fine-grained structures in human moral judgments that traditional distance-based clustering (like K-means) misses. This allows for more precise categorization of morally ambiguous situations, leading to higher accuracy in predicting the appropriate moral context.

Sources

Related papers