Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations
Summary
This arXiv paper analyzes how safety-focused generations of GPT models may transform discriminatory content rather than remove it, a process termed harm laundering. It examines 450,000 gender-directed outputs across GPT-2 to GPT-5 and finds that female-directed content loses some harmful clusters while male-directed content gains representational stigma, with toxicity metrics not aligning with representational harm. The authors formalize a three-criteria test and propose a three-stage detection protocol to audit safety across generative models.