Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness
Merged summary
TL;DR - GAMUT is a benchmark and two-level rubric framework for measuring factual completeness—not just correctness—in long-form generation. It exposes substantial room for improvement, with the best of 14 evaluated models scoring 58.7%.
- Structured meta-rubrics encode required content, organization, and importance, then compile into machine-gradable binary checklists.
- The benchmark includes 1,813 questions across 10 domains, grounded in wearable imagery with expert-verified, evidence-backed rubrics.
- Its modality-agnostic design supports both multimodal and text-only evaluation.
- Results indicate GAMUT is challenging, discriminative, and robust to the choice of LLM judge.
Sources (1)
Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness
TL;DR - GAMUT is a benchmark and two-level rubric framework for measuring factual completeness—not just correctness—in long-form generation. It exposes substantial room for improvement, with the best of 14 evaluated models scoring 58.7%.
- Structured meta-rubrics encode required content, organization, and importance, then compile into machine-gradable binary checklists.
- The benchmark includes 1,813 questions across 10 domains, grounded in wearable imagery with expert-verified, evidence-backed rubrics.
- Its modality-agnostic design supports both multimodal and text-only evaluation.
- Results indicate GAMUT is challenging, discriminative, and robust to the choice of LLM judge.