🛰️ Daily AI Frontier
‹ back to 2026-09-08

Beyond Aggregate Scores: Behavioral Correctness Assumptions for Assessing Reference-Based Automatic Evaluation Methods

Research LLM Evaluation

Ranking

Overall 78
Content 90
Popularity N/A

No observed public metrics; popularity remains neutral/archived.

Representative image for Beyond Aggregate Scores: Behavioral Correctness Assumptions for Assessing Reference-Based Automatic Evaluation Methods

Merged summary

TL;DR - This paper proposes behavioral correctness assumptions for diagnosing how reference-based natural language generation evaluators respond to controlled, correctness-preserving or correctness-altering transformations. It reveals behavioral differences and trade-offs that aggregate agreement scores can obscure.

  • The framework defines a taxonomy of expected evaluator behaviors and tests them using controlled response transformations.
  • It covers lexical, character-level, semantic, LLM-based, and hybrid evaluation methods.
  • Evaluators are analyzed for stability, sensitivity, repeat-run variability, configuration sensitivity, and reproducibility.
  • No evaluator satisfies every proposed assumption, and evaluators with similar aggregate performance can have substantially different behavioral profiles.

Sources (1)

Beyond Aggregate Scores: Behavioral Correctness Assumptions for Assessing Reference-Based Automatic Evaluation Methods

arXiv cs.AI Maria Mahbub, Ashley Rice, Michael R. Munroe, Amidu Kamara, Amir Sadovnik 2026-09-04 arXiv:2609.05289
Public signals N/A
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-26 14:16:45.980060 UTC

TL;DR - This paper proposes behavioral correctness assumptions for diagnosing how reference-based natural language generation evaluators respond to controlled, correctness-preserving or correctness-altering transformations. It reveals behavioral differences and trade-offs that aggregate agreement scores can obscure.

  • The framework defines a taxonomy of expected evaluator behaviors and tests them using controlled response transformations.
  • It covers lexical, character-level, semantic, LLM-based, and hybrid evaluation methods.
  • Evaluators are analyzed for stability, sensitivity, repeat-run variability, configuration sensitivity, and reproducibility.
  • No evaluator satisfies every proposed assumption, and evaluators with similar aggregate performance can have substantially different behavioral profiles.
item →