🛰️ Daily AI Frontier
‹ back to 2026-08-15

代码榜刷满分,AI科研却撞墙?140道真实研究题撕开「自进化」真相

Research LLM Agents

Ranking

Overall 78
Content 85
Popularity 61

Observed public metrics from 1 member.

Representative image for 代码榜刷满分,AI科研却撞墙?140道真实研究题撕开「自进化」真相

Merged summary

TL;DR - MLS-Bench evaluates AI agents on 140 real machine-learning research tasks across 12 fields. Frontier models can optimize and recombine existing methods but still fail to reliably surpass human SOTA through genuine algorithmic innovation.

  • Models received SOTA implementations and could iteratively edit code and run experiments, yet their overall results did not exceed human SOTA.
  • Gains mainly came from tuning, search, and recombining known components; prompts to optimize existing methods worked better than prompts to invent new ones.
  • Each task tests transfer across at least three environments while controlling data pipelines, compute protocols, and model capacity to exclude shortcuts.
  • More sampling, iterations, compute, web search, and theoretical context produced limited gains, highlighting scientific judgment—not information access alone—as the bottleneck.

Sources (1)

代码榜刷满分,AI科研却撞墙?140道真实研究题撕开「自进化」真相

WeChat: 新智元 2026-08-12 arXiv:2605.08678
Public signals Hugging Face upvotes 9
Providers: Hugging Face · Upvotes 9 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-14 14:23:30.213350 UTC

TL;DR - MLS-Bench evaluates AI agents on 140 real machine-learning research tasks across 12 fields. Frontier models can optimize and recombine existing methods but still fail to reliably surpass human SOTA through genuine algorithmic innovation.

  • Models received SOTA implementations and could iteratively edit code and run experiments, yet their overall results did not exceed human SOTA.
  • Gains mainly came from tuning, search, and recombining known components; prompts to optimize existing methods worked better than prompts to invent new ones.
  • Each task tests transfer across at least three environments while controlling data pipelines, compute protocols, and model capacity to exclude shortcuts.
  • More sampling, iterations, compute, web search, and theoretical context produced limited gains, highlighting scientific judgment—not information access alone—as the bottleneck.
item →