🛰️ Daily AI Frontier
‹ back to 2026-09-08

Beyond Aggregate Scores: Behavioral Correctness Assumptions for Assessing Reference-Based Automatic Evaluation Methods

arXiv cs.AI LLM Evaluation Maria Mahbub, Ashley Rice, Michael R. Munroe, Amidu Kamara, Amir Sadovnik 2026-09-04
Representative image for Beyond Aggregate Scores: Behavioral Correctness Assumptions for Assessing Reference-Based Automatic Evaluation Methods

TL;DR - This paper proposes behavioral correctness assumptions for diagnosing how reference-based natural language generation evaluators respond to controlled, correctness-preserving or correctness-altering transformations. It reveals behavioral differences and trade-offs that aggregate agreement scores can obscure.

  • The framework defines a taxonomy of expected evaluator behaviors and tests them using controlled response transformations.
  • It covers lexical, character-level, semantic, LLM-based, and hybrid evaluation methods.
  • Evaluators are analyzed for stability, sensitivity, repeat-run variability, configuration sensitivity, and reproducibility.
  • No evaluator satisfies every proposed assumption, and evaluators with similar aggregate performance can have substantially different behavioral profiles.

view merged work →