🛰️ Daily AI Frontier
‹ back to 2026-09-04

Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints

arXiv cs.AI LLM Evaluation Haoyaun Zhu, Jie Zhang 2026-09-03

TL;DR - A preregistered audit finds that black-box LLM judges on shared endpoints are too unstable to serve as reliable measurement instruments. Repeated identical requests produced substantially inconsistent rankings, undermining fixed evaluation gates and leaderboard comparisons.

  • Same-window repeat rankings achieved Spearman 0.400 versus the preregistered 0.90 requirement; byte-identical next-day replays reached 0.78 versus 0.99.
  • Instability arose from label-to-meaning bias, candidate differences far below the observer’s noise floor, and identical inputs yielding different rankings.
  • Changing metrics, sampling strategies, waiting, or switching among four providers did not resolve the tested reliability failures.
  • Self-hosting with batch-invariant kernels helped only under quiet-server conditions; the authors recommend validating instrument reliability before freezing evaluation thresholds.

view merged work →