Beyond Aggregate Scores: Behavioral Correctness Assumptions for Assessing Reference-Based Automatic Evaluation Methods
TL;DR - This paper proposes behavioral correctness assumptions for diagnosing how reference-based natural language generation evaluators respond to controlled, correctness-preserving or correctness-altering transformations. It reveals behavioral differences and trade-offs that aggregate agreement scores can obscure.
- The framework defines a taxonomy of expected evaluator behaviors and tests them using controlled response transformations.
- It covers lexical, character-level, semantic, LLM-based, and hybrid evaluation methods.
- Evaluators are analyzed for stability, sensitivity, repeat-run variability, configuration sensitivity, and reproducibility.
- No evaluator satisfies every proposed assumption, and evaluators with similar aggregate performance can have substantially different behavioral profiles.