Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations

arXiv:2609.20779 · cs.CL, cs.AI · Submitted 2026-09-17 · Read on arXiv

cs.CL, cs.AI

Submitted: 2026-09-17

Updated: 2026-09-17

Comments: Accepted at EMNLP 26 Main Conference

Code: https://github.com/unitaryai/detoxify

License: http://creativecommons.org/licenses/by/4.0/

The gist: Safety evaluations for large language models rely on surface-form classifiers that report declining harm scores across model generations.

Terminology

Abstract

Safety evaluations for large language models rely on surface-form classifiers that report declining harm scores across model generations. We provide evidence that this methodology is systematically incomplete: explicit discriminatory content is transformed rather than removed. We call this harm laundering. Analysing 450,000 gender-directed completions across 15 models spanning GPT-2 through to GPT-5 (OpenAI GPT lineage; three demographic conditions), we show that sexual violence clusters prevalent in GPT-2 women-directed output disappear by GPT-4, while men-directed completions gain positive representational territory (caregiving, emotional range, ally identity) that women-directed completions do not. The pattern is most visible at GPT-5: Topic 5 (1,997 documents) frames breast cancer as a men's rights debate, while zero equivalent clusters appear in women-directed output. Three independent classifiers score this content as non-toxic. Sentiment scores invert at GPT-4: early models demean women; later models over-correct. Topic diversity in women-directed completions falls 36% relative to men at the GPT-4 alignment boundary (W/M = 0.58, from 0.91 at GPT-2). REGARD representational harm disparity correlates with release date (ρ= +0.55, p =.034) while Detoxify does not (ρ= -0.23, p =.42): toxicity scores fall as representational harm grows. We formalise harm laundering as a three-criteria test and provide a three-stage detection protocol applicable to any generative model. Within the OpenAI GPT lineage, toxicity score reduction is not a sufficient proxy for harm reduction.

Sources

Related papers