🛰️ Daily AI Frontier
‹ back to 2026-09-04

Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints

Research LLM Evaluation

Ranking

Overall 81
Content 100
Popularity 37

Observed public metrics from 1 member.

Merged summary

TL;DR - A preregistered audit finds that black-box LLM judges on shared endpoints are too unstable to serve as reliable measurement instruments. Repeated identical requests produced substantially inconsistent rankings, undermining fixed evaluation gates and leaderboard comparisons.

  • Same-window repeat rankings achieved Spearman 0.400 versus the preregistered 0.90 requirement; byte-identical next-day replays reached 0.78 versus 0.99.
  • Instability arose from label-to-meaning bias, candidate differences far below the observer’s noise floor, and identical inputs yielding different rankings.
  • Changing metrics, sampling strategies, waiting, or switching among four providers did not resolve the tested reliability failures.
  • Self-hosting with batch-invariant kernels helped only under quiet-server conditions; the authors recommend validating instrument reliability before freezing evaluation thresholds.

Sources (1)

Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints

arXiv cs.AI Haoyaun Zhu, Jie Zhang 2026-09-03 arXiv:2609.04198
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-24 14:26:54.833706 UTC

TL;DR - A preregistered audit finds that black-box LLM judges on shared endpoints are too unstable to serve as reliable measurement instruments. Repeated identical requests produced substantially inconsistent rankings, undermining fixed evaluation gates and leaderboard comparisons.

  • Same-window repeat rankings achieved Spearman 0.400 versus the preregistered 0.90 requirement; byte-identical next-day replays reached 0.78 versus 0.99.
  • Instability arose from label-to-meaning bias, candidate differences far below the observer’s noise floor, and identical inputs yielding different rankings.
  • Changing metrics, sampling strategies, waiting, or switching among four providers did not resolve the tested reliability failures.
  • Self-hosting with batch-invariant kernels helped only under quiet-server conditions; the authors recommend validating instrument reliability before freezing evaluation thresholds.
item →