代码榜刷满分,AI科研却撞墙?140道真实研究题撕开「自进化」真相
TL;DR - MLS-Bench evaluates AI agents on 140 real machine-learning research tasks across 12 fields. Frontier models can optimize and recombine existing methods but still fail to reliably surpass human SOTA through genuine algorithmic innovation.
- Models received SOTA implementations and could iteratively edit code and run experiments, yet their overall results did not exceed human SOTA.
- Gains mainly came from tuning, search, and recombining known components; prompts to optimize existing methods worked better than prompts to invent new ones.
- Each task tests transfer across at least three environments while controlling data pipelines, compute protocols, and model capacity to exclude shortcuts.
- More sampling, iterations, compute, web search, and theoretical context produced limited gains, highlighting scientific judgment—not information access alone—as the bottleneck.