A Judge Should Know What Changed:Construct Validity for LLM-as-a-Judge Evaluation
TL;DR - This paper introduces a two-dimensional construct-validity framework for LLM judges, measuring both invariance to irrelevant edits and sensitivity to meaningful ones. Across seven judges, high invariance coexisted with poor sensitivity, showing that agreement and robustness alone can overstate evaluator quality.
- The framework separates invariance (S) under construct-preserving edits from sensitivity (R) under minimal construct-changing edits; the two are independent and cannot be faithfully collapsed into one score.
- At matched invariance of at least 0.90, judges averaged S = 0.945 but only R = 0.319.
- Judges were more sensitive to scope changes than strength changes: R_scope = 0.383 versus R_strength = 0.262, consistently across all seven judges.
- Surface-only predictors reproduced 55%–67% of labels in five public datasets, including 67.4% of MT-Bench human votes, indicating potential validation-set artifacts.