🛰️ Daily AI Frontier
‹ back to 2026-08-13

Who Thinks Best Depends on How Long You Let Them: Budget-Dependent Rankings in LLM Evaluation

arXiv cs.AI LLMs & Foundation Models Rodrigo Guedes de Souza, Alison R. Panisson 2026-08-12

TL;DR - LLM rankings can change or reverse as generation-token budgets vary, undermining evaluations performed at a single inference budget. Budget-conditioned evaluation and model routing may better reflect accuracy, efficiency, and model complementarity.

  • Across 56,476 inferences, 3–19% of items became less accurate with larger budgets.
  • Model rankings reversed across budgets on all three reasoning benchmarks.
  • Oracle model selection improved performance by up to 27.8 percentage points, especially under constrained budgets.
  • A budget-aware router captured 14.1% of the cross-domain oracle gap, but budget features transferred poorly between domains.

view merged work →