🛰️ Daily AI Frontier
‹ back to 2026-07-31

How Benchmarks Mis-Score Computer-Use Agents

arXiv cs.AI LLM Agents Zihan Dong, Zhiyuan Ma, Zekun Wang, Yunqing Li, Zirou Liu, Ruixuan Deng, Qishi Zhan, Rui Qian 2026-07-30

TL;DR - An audit of computer-use agent benchmarks finds that 15.3% of sampled FAIL verdicts are incorrect, showing that brittle evaluators and broken tasks can significantly distort reported performance.

  • Audits 150 failure-scored trajectories across five web, enterprise-workflow, and desktop-control benchmarks.
  • Attributes 10.7% of FAIL verdicts to evaluator false negatives and 4.7% to broken tasks.
  • Finds verification/feedback and planning failures dominate execution/grounding errors among genuine failures.
  • Proposes a reliability framework and stage-specific evaluation rules covering task construction, trajectory observation, scoring, and reporting.

view merged work →