ImpossibleRubrics: Stress-Testing Generated Rubrics as Reward Signals
TL;DR - ImpossibleRubrics is a 169-task benchmark testing whether LLM-generated rubrics reward honest acknowledgment of impossible requests over adversarially fabricated answers. It reveals that rubric specificity can worsen reward hacking by signaling which false claims attackers should make.
- Eleven rubric generators were exploited on 8–26% of tasks in the unbiased evaluation cut.
- On a deliberately selected stress cut, the strongest tested generator was exploited 36% of the time, versus 0% for a certificate-faithful rubric.
- Seven of eleven tailored-rubric generators performed worse than a generic “be decisive, penalize hedging” rubric, which had a 64% exploitation rate.
- Each task includes a verifiable oracle certificate defining permissible claims, enabling systematic detection of certificate-violating answers.