🛰️ Daily AI Frontier
‹ back to 2026-07-22

Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness

arXiv cs.CL LLMs & Foundation Models Xilun Chen, Zhaleh Feizollahi, Ross Goodwin, Seungwhan Moon, Scott Yih, Pinar Donmez, Babak Damavandi, Luna Dong 2026-07-21

TL;DR - GAMUT is a benchmark and two-level rubric framework for measuring factual completeness—not just correctness—in long-form generation. It exposes substantial room for improvement, with the best of 14 evaluated models scoring 58.7%.

  • Structured meta-rubrics encode required content, organization, and importance, then compile into machine-gradable binary checklists.
  • The benchmark includes 1,813 questions across 10 domains, grounded in wearable imagery with expert-verified, evidence-backed rubrics.
  • Its modality-agnostic design supports both multimodal and text-only evaluation.
  • Results indicate GAMUT is challenging, discriminative, and robust to the choice of LLM judge.

view merged work →