🛰️ Daily AI Frontier
‹ back to 2026-09-22

OSWorld-Pro: Process-based Evaluation for Computer Use Agents

Research LLM Agents

Ranking

Overall 82
Content 95
Popularity N/A

No observed public metrics; popularity remains neutral/archived.

Representative image for OSWorld-Pro: Process-based Evaluation for Computer Use Agents

Merged summary

TL;DR - OSWorld-Pro is a process-based benchmark for evaluating computer-use agents across intermediate subgoals rather than only final task outcomes. It exposes specific failure modes that end-state evaluation can obscure, helping target improvements in agent reliability and efficiency.

  • Includes 300+ tasks, 2,800+ subgoals, and over 67,000 human annotations.
  • Uses human-aligned LLM judges to assess progress through sequentially dependent subgoals.
  • Claude Opus 5, the top reported performer, scored 75.7% on OSWorld-Pro versus 83.4% on OSWorld.
  • Process-level analysis identifies failures such as subgoal-irrelevant actions, keyboard-input errors, and imprecise GUI clicks.

Sources (1)

OSWorld-Pro: Process-based Evaluation for Computer Use Agents

arXiv cs.CL Zhilin Wang, Shaokun Zhang, Yifan Zhang, Hao Zhang, Jin Xu, Binfeng Xu, Jian Hu, Yunheng Zou, Karan Sapra, Andrew Tao, Jan Kautz, Yi Dong 2026-09-21 arXiv:2609.24890
Public signals N/A
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-26 14:14:42.060144 UTC

TL;DR - OSWorld-Pro is a process-based benchmark for evaluating computer-use agents across intermediate subgoals rather than only final task outcomes. It exposes specific failure modes that end-state evaluation can obscure, helping target improvements in agent reliability and efficiency.

  • Includes 300+ tasks, 2,800+ subgoals, and over 67,000 human annotations.
  • Uses human-aligned LLM judges to assess progress through sequentially dependent subgoals.
  • Claude Opus 5, the top reported performer, scored 75.7% on OSWorld-Pro versus 83.4% on OSWorld.
  • Process-level analysis identifies failures such as subgoal-irrelevant actions, keyboard-input errors, and imprecise GUI clicks.
item →