🛰️ Daily AI Frontier
‹ back to 2026-09-20

Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations

arXiv cs.CL LLMs & Foundation Models Sarah Wyer, Sue Black, Noura Al Moubayed 2026-09-17

TL;DR - A study of 450,000 gender-directed completions across 15 GPT models argues that safety training can transform explicit discrimination into subtler representational harm rather than eliminate it. This matters because standard toxicity classifiers may report improvement while missing growing disparities in framing and topic diversity.

  • Sexual-violence clusters targeting women disappeared by GPT-4, but later models introduced asymmetric positive representations and gendered issue framing.
  • At the GPT-4 alignment boundary, topic diversity in women-directed completions fell 36% relative to men, with the women-to-men ratio dropping from 0.91 at GPT-2 to 0.58.
  • Representational-harm disparity increased with release date, while Detoxify toxicity scores did not show the same trend.
  • The authors formalize “harm laundering” with a three-criteria test and propose a three-stage detection protocol for generative models.

view merged work →