Rubrics on Trial: Evolving Rubrics from a Single Query via Synthetic Pairwise Evidence
TL;DR - "Rubrics on Trial" is a query-only framework that automatically evolves validated evaluation rubrics for LLMs without human annotations or model training, addressing the difficulty of building reliable query-specific rubrics.
- Grows a rubric set from empty using only synthetic rubric-conditioned response pairs — no human rubrics, preference data, or sampled responses required.
- Validates each candidate rubric before adding it, screening out non-discriminative, over-specific, and style-only rubrics that don't reflect true answer quality.
- Evaluated on five preference benchmark suites, achieving the best average accuracy and leading on six of seven evaluation sets.
- Targets a known weakness of direct query-to-rubric generation: plausible rubrics may reward optional style or penalize valid alternative strategies without an explicit usefulness check.