🛰️ Daily AI Frontier
‹ back to 2026-08-15

代码榜刷满分,AI科研却撞墙?140道真实研究题撕开「自进化」真相

WeChat: 新智元 LLM Agents 2026-08-12
Representative image for 代码榜刷满分,AI科研却撞墙?140道真实研究题撕开「自进化」真相

TL;DR - MLS-Bench evaluates AI agents on 140 real machine-learning research tasks across 12 fields. Frontier models can optimize and recombine existing methods but still fail to reliably surpass human SOTA through genuine algorithmic innovation.

  • Models received SOTA implementations and could iteratively edit code and run experiments, yet their overall results did not exceed human SOTA.
  • Gains mainly came from tuning, search, and recombining known components; prompts to optimize existing methods worked better than prompts to invent new ones.
  • Each task tests transfer across at least three environments while controlling data pipelines, compute protocols, and model capacity to exclude shortcuts.
  • More sampling, iterations, compute, web search, and theoretical context produced limited gains, highlighting scientific judgment—not information access alone—as the bottleneck.

view merged work →