Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints
Ranking
Overall
81
Content
100
Popularity
37
Observed public metrics from 1 member.
Merged summary
TL;DR - A preregistered audit finds that black-box LLM judges on shared endpoints are too unstable to serve as reliable measurement instruments. Repeated identical requests produced substantially inconsistent rankings, undermining fixed evaluation gates and leaderboard comparisons.
- Same-window repeat rankings achieved Spearman 0.400 versus the preregistered 0.90 requirement; byte-identical next-day replays reached 0.78 versus 0.99.
- Instability arose from label-to-meaning bias, candidate differences far below the observer’s noise floor, and identical inputs yielding different rankings.
- Changing metrics, sampling strategies, waiting, or switching among four providers did not resolve the tested reliability failures.
- Self-hosting with batch-invariant kernels helped only under quiet-server conditions; the authors recommend validating instrument reliability before freezing evaluation thresholds.
Sources (1)
Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - A preregistered audit finds that black-box LLM judges on shared endpoints are too unstable to serve as reliable measurement instruments. Repeated identical requests produced substantially inconsistent rankings, undermining fixed evaluation gates and leaderboard comparisons.
- Same-window repeat rankings achieved Spearman 0.400 versus the preregistered 0.90 requirement; byte-identical next-day replays reached 0.78 versus 0.99.
- Instability arose from label-to-meaning bias, candidate differences far below the observer’s noise floor, and identical inputs yielding different rankings.
- Changing metrics, sampling strategies, waiting, or switching among four providers did not resolve the tested reliability failures.
- Self-hosting with batch-invariant kernels helped only under quiet-server conditions; the authors recommend validating instrument reliability before freezing evaluation thresholds.