🛰️ Daily AI Frontier
‹ back to 2026-09-24

How Sensitive Are LLM Leaderboard Claims to Hidden Model Selection?

Research LLMs & Foundation Models

Ranking

Overall 78
Content 100
Popularity 27

Observed public metrics from 1 member.

Merged summary

TL;DR - This paper introduces sensitivity curves for assessing whether LLM leaderboard advantages remain statistically credible when providers may have privately selected among multiple model variants. An audit of 394 adjacent-rank claims found that 391 lacked statistical support even before correcting for hidden selection.

  • The method estimates the maximum hidden variant count consistent with a statistically supported advantage, given a lower bound on within-family correlation.
  • Correlation estimates depend strongly on the ranking score and resampling model: reported values ranged from 0.46 to 0.92 for composite scores.
  • For claims that pass an uncorrected test, certification may still hinge on assumptions about correlation among hidden variants.
  • The curves expose these assumptions without requiring researchers to estimate the provider’s unobserved search size.

Sources (1)

How Sensitive Are LLM Leaderboard Claims to Hidden Model Selection?

arXiv stat.ML Chen Yang, Xianyang Zhang, Jun Chen 2026-09-23 arXiv:2609.28177
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-24 14:15:38.007968 UTC

TL;DR - This paper introduces sensitivity curves for assessing whether LLM leaderboard advantages remain statistically credible when providers may have privately selected among multiple model variants. An audit of 394 adjacent-rank claims found that 391 lacked statistical support even before correcting for hidden selection.

  • The method estimates the maximum hidden variant count consistent with a statistically supported advantage, given a lower bound on within-family correlation.
  • Correlation estimates depend strongly on the ranking score and resampling model: reported values ranged from 0.46 to 0.92 for composite scores.
  • For claims that pass an uncorrected test, certification may still hinge on assumptions about correlation among hidden variants.
  • The curves expose these assumptions without requiring researchers to estimate the provider’s unobserved search size.
item →