How Sensitive Are LLM Leaderboard Claims to Hidden Model Selection?
TL;DR - This paper introduces sensitivity curves for assessing whether LLM leaderboard advantages remain statistically credible when providers may have privately selected among multiple model variants. An audit of 394 adjacent-rank claims found that 391 lacked statistical support even before correcting for hidden selection.
- The method estimates the maximum hidden variant count consistent with a statistically supported advantage, given a lower bound on within-family correlation.
- Correlation estimates depend strongly on the ranking score and resampling model: reported values ranged from 0.46 to 0.92 for composite scores.
- For claims that pass an uncorrected test, certification may still hinge on assumptions about correlation among hidden variants.
- The curves expose these assumptions without requiring researchers to estimate the provider’s unobserved search size.