🛰️ Daily AI Frontier
‹ back to 2026-09-24

How Sensitive Are LLM Leaderboard Claims to Hidden Model Selection?

arXiv stat.ML LLMs & Foundation Models Chen Yang, Xianyang Zhang, Jun Chen 2026-09-23

TL;DR - This paper introduces sensitivity curves for assessing whether LLM leaderboard advantages remain statistically credible when providers may have privately selected among multiple model variants. An audit of 394 adjacent-rank claims found that 391 lacked statistical support even before correcting for hidden selection.

  • The method estimates the maximum hidden variant count consistent with a statistically supported advantage, given a lower bound on within-family correlation.
  • Correlation estimates depend strongly on the ranking score and resampling model: reported values ranged from 0.46 to 0.92 for composite scores.
  • For claims that pass an uncorrected test, certification may still hinge on assumptions about correlation among hidden variants.
  • The curves expose these assumptions without requiring researchers to estimate the provider’s unobserved search size.

view merged work →