How Sensitive Are LLM Leaderboard Claims to Hidden Model Selection?
Ranking
Overall
78
Content
100
Popularity
27
Observed public metrics from 1 member.
Merged summary
TL;DR - This paper introduces sensitivity curves for assessing whether LLM leaderboard advantages remain statistically credible when providers may have privately selected among multiple model variants. An audit of 394 adjacent-rank claims found that 391 lacked statistical support even before correcting for hidden selection.
- The method estimates the maximum hidden variant count consistent with a statistically supported advantage, given a lower bound on within-family correlation.
- Correlation estimates depend strongly on the ranking score and resampling model: reported values ranged from 0.46 to 0.92 for composite scores.
- For claims that pass an uncorrected test, certification may still hinge on assumptions about correlation among hidden variants.
- The curves expose these assumptions without requiring researchers to estimate the provider’s unobserved search size.
Sources (1)
How Sensitive Are LLM Leaderboard Claims to Hidden Model Selection?
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - This paper introduces sensitivity curves for assessing whether LLM leaderboard advantages remain statistically credible when providers may have privately selected among multiple model variants. An audit of 394 adjacent-rank claims found that 391 lacked statistical support even before correcting for hidden selection.
- The method estimates the maximum hidden variant count consistent with a statistically supported advantage, given a lower bound on within-family correlation.
- Correlation estimates depend strongly on the ranking score and resampling model: reported values ranged from 0.46 to 0.92 for composite scores.
- For claims that pass an uncorrected test, certification may still hinge on assumptions about correlation among hidden variants.
- The curves expose these assumptions without requiring researchers to estimate the provider’s unobserved search size.