Debate Training Reduces Reward Hacking in RLAIF
TL;DR - Adversarial debate training reduced reward hacking relative to standard RLAIF when a weaker LLM judged math solutions. This suggests debate may improve oversight of increasingly capable models, provided the competing players are carefully balanced.
- A Gemini 2.5 Flash-class policy debated under a frozen Gemini 2.5 Flash Lite judge, while the RLAIF baseline quickly learned to exploit the judge.
- Debate maintained judge performance and recovered 45% of the gap to peak validation accuracy across many reinforcement-learning steps.
- Adding another debate round compensated for weaker judges, and debate incentives overrode prompted misalignment.
- Critique limits of up to 150 words prevented the critic from hacking the judge, but constrained the critic’s explanatory clarity.