Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments
summary
In short
The episode discusses the paper "Agreement Is Not Alignment," which argues that simply matching human answers on moral questions does not prove a model understands ethics. The authors found that while LLMs achieved high label agreement, their underlying moral reasoning (or 'moral grounds') systematically diverged from human thought, necessitating rationale-aware evaluation for trust and safety.
Key concepts
- Agreement vs. Alignment
- Agreement means a model gives the same final answer as a human on an ethical judgment. Alignment requires more; it means the model is responsive to the same kinds of moral considerations and reasoning processes that humans recognize, even if they arrive at the same label.
- Structured Rationale/Moral Grounds
- This refers to collecting structured reasons for a judgment (e.g., 'harm' or 'respectfulness') rather than just a final yes/no answer. The paper used this method to show that models' explanations differed significantly from human explanations.
- Deontology
- Deontology is one of the five ethical domains tested in the benchmark. In this area, humans make precise judgments about whether an excuse is relevant to a specific duty. The episode noted that LLMs struggled with this distinction, often over-moralizing situations.
Terminology used across episodes
This episode discusses
- Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments · Paper Radio
- Moral Foundations of Large Language Models
- Faithfulness vs. Plausibility: On the (Un)Reliability of Explanations from Large Language Models
- A General Language Assistant as a Laboratory for Alignment
- Constitutional AI: Harmlessness from AI Feedback
- e-SNLI: Natural Language Inference with Natural Language Explanations
- Deep reinforcement learning from human preferences
- The Capacity for Moral Self-Correction in Large Language Models
- Can Machines Learn Morality? The Delphi Experiment
- When to Make Exceptions: Exploring Language Models as Accounts of Human Moral Judgment
- Measuring Faithfulness in Chain-of-Thought Reasoning
- Training language models to follow instructions with human feedback
- Knowledge of cultural moral norms in large language models
- Learning to summarize from human feedback
- The Moral Foundations Reddit Corpus
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting
The paper
Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments · Read on arXiv
Octavian M. Machidon, Alina L. Machidon, Vojko Strahovnik, Mateja Centa Strahovnik, Jonas Miklavčič, Marko Robnik Šikonja
University of Ljubljana
Agreement with human judgments is a common proxy for evaluating the alignment of large language models (LLMs). Yet agreement in final labels does not show that human annotators and models rely on the same moral grounds. Two agents may reach the same judgment while appealing to different principles, contextual assumptions, or interpretations of the situation. We test this distinction using a curated 500-item ETHICS-derived benchmark spanning five domains of moral judgment, with new human annotator and LLM annotations of both final labels and supporting rationales. Across frontier and open model families, agreement with human annotator majority labels is often high. However, rationale-level analysis reveals systematic divergence in the moral grounds expressed by human annotators and models. In particular, models redistribute attention across categories such as harm, respect, promise-keeping, justice, desert, and excuse relevance, even when their final labels match the human annotator majority. Our results show that agreement should not be treated as equivalent to alignment. Label-based evaluation can therefore be misleadingly reassuring unless complemented by analysis of the reasons, principles, and moral priorities expressed in model judgments.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments".
Jane: The paper was written by Octavian M. Machidon, Alina L. Machidon, Vojko Strahovnik, Mateja Centa Strahovnik, Jonas Miklavčič et al. from University of Ljubljana.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the channel, everyone. Today we're looking at a paper that's been making the rounds in the AI ethics world, and the title alone is worth pausing on. It's called "Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments." Jane, when you first saw that title, what went through your head?
Jane: Honestly, Tom, I thought it was one of those titles that sounds obvious once you hear it, but then you realize nobody's actually tested it properly. The authors are saying that just because a language model gives the same answer as a human on a moral question, that doesn't mean they're thinking about it the same way. And that's a big deal.
Tom: Exactly. The team is from the University of Ljubljana — Machidon, Strahovnik, Miklavčič, Robnik Šikonja, and colleagues. They took the ETHICS benchmark, which is this standard set of moral judgment tasks, and they didn't just ask models to pick the right answer. They also asked them to explain why, using the same structured categories that human annotators used.
Jane: And that's where things get interesting. The models agreed with human majority labels around eighty-eight to eighty-nine percent of the time. That sounds great, right? But when you look at the reasons they gave, the moral grounds, they were systematically different from what humans said.
Tom: So you can have a model that says "yes, this action is wrong" and a human who says "yes, this action is wrong," but the model thinks it's wrong because it's disrespectful, while the human thinks it's wrong because it breaks a promise. Same verdict, totally different moral reasoning.
Jane: And the paper's point is that this matters. If you're deploying a model to moderate content or give ethical advice, you need to know not just what it decides, but why it decides that way. Because those reasons will shape how it handles new situations it hasn't seen before.
Tom: Right, and that's the gap between agreement and alignment. Agreement is just matching the final label. Alignment, real alignment, means the model is responsive to the same kinds of moral considerations that humans recognize.
Jane: And that's what makes this paper so important. It's giving us a way to measure that gap, not just talk about it.
Tom: Lu, you've been nodding along — what's your take on the title and the framing?
Lu: I think the title is a genuine contribution in itself. It names a problem that a lot of us in the field have felt but haven't articulated clearly. We've been using agreement as a proxy for alignment because it's easy to compute, but this paper shows that proxy is misleading. And once you see the data, you can't unsee it.
Tom: So we're going to dig into that data in a moment. But first, Jane, what would you say to a listener who thinks "well, if the model gives the right answer, why should I care about the reasons?"
Jane: I'd say imagine you have a doctor who always prescribes the right medication, but for the wrong reasons. One day, you give them a patient who's slightly different, and suddenly their reasoning fails because they never understood why the medication worked in the first place. That's the risk with these models.
Tom: Great analogy. And that's exactly what the paper's authors are worried about — models that look aligned on the surface but are brittle underneath. We'll get into the specifics next.
Summary: Tom: So we're back with "Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments," and now we need to talk about what the paper actually did. Jane, can you walk us through the setup?
Jane: Sure. They curated a five hundred-item benchmark from ETHICS, with one hundred items from each of five domains: commonsense morality, deontology, justice, utilitarianism, and virtue ethics. Then they had three human annotators per item, drawn from a pool of five, and they also asked seven different language models to answer the same questions under the same protocol.
Tom: And the key innovation was that they didn't just collect final labels. They also collected rationales — structured reasons for the judgments. For commonsense morality, that meant picking from a list of ten moral principles like harm, veracity, promise-keeping, respectfulness, and so on.
Jane: Right. And for deontology, annotators had to classify whether an excuse was relevant to a duty or not. That turned out to be where the biggest divergence appeared.
Tom: Let's get into that, because it's the most striking finding. In deontology, human annotators most often classified weak excuses as "irrelevant" to the duty. But both frontier and open models most often classified the whole situation as "morally wrong." That's a huge difference in how they're framing the task.
Jane: It really is. The humans are making a precise judgment about the relationship between an excuse and a duty. The models are responding to the moral charge of the scenario overall. So you can have a case where someone says "I can't pay the toll because I stole this car," and everyone agrees the excuse is bad. But humans say it's irrelevant to the duty of paying the toll, while models say the whole situation is morally wrong.
Lu: And that's not just a semantic quibble. It shows the models are over-moralizing. They're treating everything as a moral violation, whereas humans are making finer distinctions about what kind of violation it is, or whether it's even a violation of the specific duty in question.
Tom: So the models are like someone who always says "that's wrong" without being able to say why it's wrong in this particular context.
Jane: Exactly. And the paper quantifies this. In the commonsense domain, both humans and models picked "harm" as the most frequent principle, but models over-selected "respectfulness" while humans more often picked "promissory fidelity" and "veracity." So models are reaching for a different moral vocabulary.
Meng: I want to jump in here, because from an engineering standpoint, this has practical implications. If a model is explaining its judgment to a user, and it says "this is wrong because it's disrespectful," but the user thinks "this is wrong because it broke a promise," the explanation is going to feel off. It might even erode trust in the system.
Tom: That's a really good point, Meng. And the paper shows this isn't a rare edge case. The human–LLM overlap on deontology principles was only zero point two three five on the Jaccard scale. That's low.
Jane: Meanwhile, the frontier models were very consistent with each other — their internal agreement was much higher than their agreement with humans. So they're not noisy; they're confidently wrong in a different direction.
Lu: That's the scary part. The models have converged on their own way of moralizing, and it's not the same as how humans do it. If we only look at label agreement, we'd never notice.
Tom: And that's the core message of "Agreement Is Not Alignment." The paper gives us a concrete way to see the gap. But what do they suggest we do about it? That's next.
Improvements: Tom: We're back with "Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments," and now we need to talk about what the authors think we should do differently. Jane, what's their main recommendation?
Jane: They're not saying we should throw away label agreement metrics. They're saying we need to add a rationale-aware layer on top. Instead of just asking "did the model pick the same answer as the human majority," we should also ask "does the model express a similar distribution of moral grounds?"
Tom: And they're careful to say this isn't about forcing models to imitate every human rationale. Human moral judgment is itself diverse, and some disagreement is legitimate. The goal is to check whether models are responsive to a broad enough range of human-recognizable moral categories.
Lu: That's an important nuance. The paper isn't proposing a single "correct" moral theory that models must follow. It's proposing a diagnostic. You want to see whether the model's moral vocabulary overlaps with the vocabulary humans actually use when they reason about these cases.
Meng: So practically, how would this work? Would you need to build a new benchmark for every deployment?
Jane: Not necessarily. The paper shows that the structured rationale approach can be applied to existing benchmarks. They used ETHICS, but the same annotation protocol — final label plus structured rationale categories — could be applied to other moral judgment datasets.
Tom: And they're also honest about the limitations. The human annotator pool is small and culturally specific — five annotators from a Central European background. So the human reference isn't universal. It's a comparison point, not moral ground truth.
Meng: That's good to hear, because if you're building a system for a global user base, you need to know that the reference distribution itself might shift across cultures.
Jane: Exactly. And the paper also notes that the free-text rationales from utilitarianism and virtue ethics weren't included in the structured analysis. That's future work — they'd need qualitative coding to compare those domains at the same level of detail.
Lu: I think the bigger point is that this changes how we evaluate alignment. Right now, a model can score high on a benchmark and be considered aligned. This paper shows that score can be misleading. You need to look under the hood at the reasons.
Tom: And that has implications beyond benchmarks. Think about content moderation, AI tutoring, or any system that gives ethical advice. If the model's reasons diverge from human reasons, it might generalize poorly to new cases, or explain itself in ways that don't resonate with users.
Meng: It also affects debugging. If a model is making errors in deployment, you need to know whether it's a reasoning problem or a vocabulary problem. Rationale-aware evaluation gives you that signal.
Jane: And the authors frame it as a transparency issue. If we want models whose judgments are inspectable and contestable by humans, we need to know what moral grounds they're actually expressing. Not just what they decide.
Tom: So the improvement is really about adding a second dimension to evaluation — not just what, but why. And that's a shift in how we think about alignment.
Lu: It's a shift that's long overdue. The paper gives us the tools to make it happen.
Tom: We'll wrap up with our final thoughts next.
Conclusion: Tom: Alright, we're at the end of our discussion of "Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments." Jane, what's the one thing you want listeners to remember?
Jane: That agreement and alignment are not the same thing. A model can match human judgments on the surface while relying on completely different moral reasons underneath. And if we only measure the surface, we're missing the part that actually matters for trust and safety.
Tom: The paper showed that on a curated five hundred-item ETHICS-derived benchmark, models hit around eighty-eight percent agreement with human majority labels. But their rationale profiles diverged systematically — especially in deontology, where models moralized situations that humans judged as irrelevant to the specific duty.
Lu: And that divergence isn't random noise. It's a consistent pattern across model families. The models have their own way of moralizing, and it's not the same as how humans do it.
Meng: From a practical standpoint, that means we need rationale-aware evaluation if we want models we can actually trust in real-world applications. You can't just look at accuracy.
Jane: The authors are careful to say this isn't about forcing models to think like humans. It's about making sure their judgments are grounded in reasons that humans can recognize, inspect, and challenge.
Tom: And that's the real contribution of this paper. It gives us a way to measure the gap between agreement and alignment, and it makes a strong case that we need to pay attention to it.
Lu: I'd add that this is going to be even more important as models get more capable. The better they get at matching human labels, the easier it is to be fooled into thinking they're aligned. This paper gives us a warning and a tool.
Jane: And with that, we're saying goodbye to "Agreement Is Not Alignment." It's been a great conversation, and I think this paper is going to influence how a lot of people evaluate AI ethics going forward.
Tom: Absolutely. Thanks to everyone who listened, and we'll see you next time with another paper from the arXiv. Take care, everyone.
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language