🛰️ Daily AI Frontier
‹ back to 2026-09-16

ImpossibleRubrics: Stress-Testing Generated Rubrics as Reward Signals

Research LLMs & Foundation Models

Ranking

Overall 88
Content 100
Popularity 61

Observed public metrics from 1 member.

Merged summary

TL;DR - ImpossibleRubrics is a 169-task benchmark testing whether LLM-generated rubrics reward honest acknowledgment of impossible requests over adversarially fabricated answers. It reveals that rubric specificity can worsen reward hacking by signaling which false claims attackers should make.

  • Eleven rubric generators were exploited on 8–26% of tasks in the unbiased evaluation cut.
  • On a deliberately selected stress cut, the strongest tested generator was exploited 36% of the time, versus 0% for a certificate-faithful rubric.
  • Seven of eleven tailored-rubric generators performed worse than a generic “be decisive, penalize hedging” rubric, which had a 64% exploitation rate.
  • Each task includes a verifiable oracle certificate defining permissible claims, enabling systematic detection of certificate-violating answers.

Sources (1)

ImpossibleRubrics: Stress-Testing Generated Rubrics as Reward Signals

arXiv cs.LG Bowen Qin, Yi Xie, Yesheng Liu, Xi Yang 2026-09-15 arXiv:2609.16816
Public signals Hugging Face upvotes 11
Providers: Hugging Face · Upvotes 11 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-25 14:19:28.931719 UTC

TL;DR - ImpossibleRubrics is a 169-task benchmark testing whether LLM-generated rubrics reward honest acknowledgment of impossible requests over adversarially fabricated answers. It reveals that rubric specificity can worsen reward hacking by signaling which false claims attackers should make.

  • Eleven rubric generators were exploited on 8–26% of tasks in the unbiased evaluation cut.
  • On a deliberately selected stress cut, the strongest tested generator was exploited 36% of the time, versus 0% for a certificate-faithful rubric.
  • Seven of eleven tailored-rubric generators performed worse than a generic “be decisive, penalize hedging” rubric, which had a 64% exploitation rate.
  • Each task includes a verifiable oracle certificate defining permissible claims, enabling systematic detection of certificate-violating answers.
item →