🛰️ Daily AI Frontier
‹ back to 2026-09-16

Coding Agents Have Converged: Why the SWE-bench Leaderboard Can No Longer Order Its Top Entries, and What to Measure Instead

arXiv cs.SE LLM Agents Fengshuo Liu, Ying Liu, Ruize Sun, Lie Luo, Siyuan Guo 2026-09-15
Representative image for Coding Agents Have Converged: Why the SWE-bench Leaderboard Can No Longer Order Its Top Entries, and What to Measure Instead

TL;DR - An audit of 254 SWE-bench submissions finds that small score gaps cannot reliably rank leading coding agents because their successes overlap heavily and results depend on model-scaffold pairings. The authors propose reporting statistical resolution, grouping sensitivity, and model-scaffold provenance instead of treating aggregate scores as definitive rankings.

  • The top two SWE-bench Verified entries both solve 396 of 500 tasks; none of the 29 adjacent top-30 pairs differ significantly under exact paired McNemar tests at α=0.05.
  • Top-ten systems share 285 successes and 51 failures, with a median solution-set nesting of 0.935 versus a score-implied baseline of 0.774.
  • Within-model scaffold performance ranges reach 29.8 percentage points, exceeding the top-30 score spread of 8.8 points, though the observational analysis does not establish causality.
  • The larger Test split distinguishes 14 of 23 adjacent pairs, indicating that evaluation-set size materially affects leaderboard resolution.

view merged work →