🛰️ Daily AI Frontier
‹ back to 2026-08-19

StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows

arXiv cs.AI LLM Agents Liya Zhu, Xin Ma, Tao Liu, Haodong Wang, Ge Zhang, Jingzhe Ding, Qingshui Gu, Yongjie Zhong, Jinxiang Meng, Yuan Gao, Yunqiu Zhou, Hao Zhu, Jifeng He, Yongzhi Liao, Xinyi Zhang, Chaoxin Li, Yi Zhu, Xi Lin, Duju Zeng, Xiang Gao, Wen Zhang, Yunyang Wang, Duo Wang, Huan Zhou, Zuo Wang, Jin Chen, Kaiyuan Zhang, Chuqian Yu, Tianhao Yu, Longxiang Liu, Jianbo Xue, Huimin Che, Jiahao Wang, Yujia Qin, Jiaheng Liu, Shen Yan, Xiaolong Chang, Wenhao Huang 2026-08-18
Representative image for StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows

TL;DR - StartupBench evaluates general-purpose agents on end-to-end workflows derived from AI startup products with demonstrated market adoption. The strongest evaluated model completes only about 30% of tasks, indicating a substantial gap between partial progress and reliable real-world delivery.

  • Tasks span diverse professional domains and require complete, deliverable-oriented outputs.
  • Fine-grained rubrics assess the complex requirements of each workflow under a unified agent harness.
  • Agents often make meaningful partial progress but fail to complete workflows successfully.
  • Complex instruction following and domain-specific expertise are identified as major failure sources.

view merged work →