Moral Entropy: Auditing Bias and Uncertainty in Moral Judgment

arXiv:2609.21992 · cs.CL, cs.CY, stat.ML · Submitted 2026-09-18 · Read on arXiv

cs.CL, cs.CY, stat.ML

Submitted: 2026-09-18

Updated: 2026-09-18

Comments: accepted to UncertaiNLP @ EMNLP 2026

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

The gist: Most work in computational ethics treats annotator disagreement on moral content as noise to be voted away, collapsed into majority vote or the more permissive any-annotator rule the moment a single

Terminology

Abstract

Most work in computational ethics treats annotator disagreement on moral content as noise to be voted away, collapsed into majority vote or the more permissive any-annotator rule the moment a single annotator flags an item. We argue this uncertainty should instead be modeled and learned from. We introduce Moral Entropy, a Bayesian framework that keeps a full posterior over the true label and decomposes its entropy into aleatoric uncertainty (irreducible disagreement about the moral content) and epistemic uncertainty (from insufficient or noisy annotation) -- and lets any heuristic consensus rule be audited against a calibrated ground truth via entropy methods such as cross-entropy/KL, Brier score, and expected calibration error. Across three corpora and fifteen discourse domains, auditing the standard aggregation rules against this posterior reveals bias that no current pipeline reports: the any-annotator rule disagrees with the calibrated posterior on roughly 30% of items -- pooled, almost entirely false positives, though the errors invert at the foundation level (19.9%/38.9% mean FPR/FNR on MFTC) -- while the stricter majority and two-vote rules miss 63-83% of true positives.

Related papers