🛰️ Daily AI Frontier
‹ back to 2026-08-26

A Judge Should Know What Changed:Construct Validity for LLM-as-a-Judge Evaluation

arXiv cs.AI LLM Evaluation Jianlin Chen, Wenhui Chen, Ziyao Lin, Chi Man Vong 2026-08-25

TL;DR - This paper introduces a two-dimensional construct-validity framework for LLM judges, measuring both invariance to irrelevant edits and sensitivity to meaningful ones. Across seven judges, high invariance coexisted with poor sensitivity, showing that agreement and robustness alone can overstate evaluator quality.

  • The framework separates invariance (S) under construct-preserving edits from sensitivity (R) under minimal construct-changing edits; the two are independent and cannot be faithfully collapsed into one score.
  • At matched invariance of at least 0.90, judges averaged S = 0.945 but only R = 0.319.
  • Judges were more sensitive to scope changes than strength changes: R_scope = 0.383 versus R_strength = 0.262, consistently across all seven judges.
  • Surface-only predictors reproduced 55%–67% of labels in five public datasets, including 67.4% of MT-Bench human votes, indicating potential validation-set artifacts.

view merged work →