🛰️ Daily AI Frontier
‹ back to 2026-07-22

Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness

Research LLMs & Foundation Models

Ranking

Overall 84
Content 90
Popularity 69

Observed public metrics from 1 member.

Merged summary

TL;DR - GAMUT is a benchmark and two-level rubric framework for measuring factual completeness—not just correctness—in long-form generation. It exposes substantial room for improvement, with the best of 14 evaluated models scoring 58.7%.

  • Structured meta-rubrics encode required content, organization, and importance, then compile into machine-gradable binary checklists.
  • The benchmark includes 1,813 questions across 10 domains, grounded in wearable imagery with expert-verified, evidence-backed rubrics.
  • Its modality-agnostic design supports both multimodal and text-only evaluation.
  • Results indicate GAMUT is challenging, discriminative, and robust to the choice of LLM judge.

Sources (1)

Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness

arXiv cs.CL Xilun Chen, Zhaleh Feizollahi, Ross Goodwin, Seungwhan Moon, Scott Yih, Pinar Donmez, Babak Damavandi, Luna Dong 2026-07-21 arXiv:2607.19322
Public signals Hugging Face upvotes 11
Providers: Hugging Face · Upvotes 11 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-08-21 14:39:05.040653 UTC

TL;DR - GAMUT is a benchmark and two-level rubric framework for measuring factual completeness—not just correctness—in long-form generation. It exposes substantial room for improvement, with the best of 14 evaluated models scoring 58.7%.

  • Structured meta-rubrics encode required content, organization, and importance, then compile into machine-gradable binary checklists.
  • The benchmark includes 1,813 questions across 10 domains, grounded in wearable imagery with expert-verified, evidence-backed rubrics.
  • Its modality-agnostic design supports both multimodal and text-only evaluation.
  • Results indicate GAMUT is challenging, discriminative, and robust to the choice of LLM judge.
item →