Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints
TL;DR - A preregistered audit finds that black-box LLM judges on shared endpoints are too unstable to serve as reliable measurement instruments. Repeated identical requests produced substantially inconsistent rankings, undermining fixed evaluation gates and leaderboard comparisons.
- Same-window repeat rankings achieved Spearman 0.400 versus the preregistered 0.90 requirement; byte-identical next-day replays reached 0.78 versus 0.99.
- Instability arose from label-to-meaning bias, candidate differences far below the observer’s noise floor, and identical inputs yielding different rankings.
- Changing metrics, sampling strategies, waiting, or switching among four providers did not resolve the tested reliability failures.
- Self-hosting with batch-invariant kernels helped only under quiet-server conditions; the authors recommend validating instrument reliability before freezing evaluation thresholds.