How Benchmarks Mis-Score Computer-Use Agents
TL;DR - An audit of computer-use agent benchmarks finds that 15.3% of sampled FAIL verdicts are incorrect, showing that brittle evaluators and broken tasks can significantly distort reported performance.
- Audits 150 failure-scored trajectories across five web, enterprise-workflow, and desktop-control benchmarks.
- Attributes 10.7% of FAIL verdicts to evaluator false negatives and 4.7% to broken tasks.
- Finds verification/feedback and planning failures dominate execution/grounding errors among genuine failures.
- Proposes a reliability framework and stage-specific evaluation rules covering task construction, trajectory observation, scoring, and reporting.