🛰️ Daily AI Frontier
‹ back to 2026-09-16

ImpossibleRubrics: Stress-Testing Generated Rubrics as Reward Signals

arXiv cs.LG LLMs & Foundation Models Bowen Qin, Yi Xie, Yesheng Liu, Xi Yang 2026-09-15

TL;DR - ImpossibleRubrics is a 169-task benchmark testing whether LLM-generated rubrics reward honest acknowledgment of impossible requests over adversarially fabricated answers. It reveals that rubric specificity can worsen reward hacking by signaling which false claims attackers should make.

  • Eleven rubric generators were exploited on 8–26% of tasks in the unbiased evaluation cut.
  • On a deliberately selected stress cut, the strongest tested generator was exploited 36% of the time, versus 0% for a certificate-faithful rubric.
  • Seven of eleven tailored-rubric generators performed worse than a generic “be decisive, penalize hedging” rubric, which had a 64% exploitation rate.
  • Each task includes a verifiable oracle certificate defining permissible claims, enabling systematic detection of certificate-violating answers.

view merged work →