🛰️ Daily AI Frontier
‹ back to 2026-08-15

Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development

arXiv cs.AI LLM Agents Yiwei Li, Wanli Yang, Hexiang Tan, Xiangzhou Huang, Zhengyu Chen, Ziran Li, Borun Chen, Shanglin Lei, Huaisheng Zhu, Hao Tian, Fei Sun, Xunliang Cai, Jingang Wang 2026-08-13

TL;DR - A framework evaluating seven frontier models on 36 long-horizon R&D tasks finds that current agents behave more like engineering optimizers than autonomous researchers. They implement practical solutions but remain inconsistent and rarely produce genuinely novel methods.

  • Rule-based metrics examine Solution Framing, Execution, and Feedback Control within each run.
  • Controlled comparisons assess whether agents reuse experience effectively within and across tasks.
  • Similar final scores can conceal different process bottlenecks, while prior experience may help or mislead later decisions.
  • Agent stability depends partly on harness design, suggesting improvements to training, inference strategies, and experience management.

view merged work →