🛰️ Daily AI Frontier
‹ back to 2026-07-30

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

arXiv cs.AI LLM Agents Qiushi Sun, Kanzhi Cheng, Yian Wang, Bowen Yang, Hang Yan, Liheng Chen, Fangzhi Xu, Zichen Ding, Nuo Chen, Jialin Cao, Xingdong Gong, Zehao Li, Kaiming Jin, Xinfeng Yuan, Zhoumianze Liu, Jingyang Gong, Zhangyue Yin, Jiahui Gao, Zhiyong Wu, Tianbao Xie, Jianbing Zhang, Ben Kao, Lingpeng Kong 2026-07-30
Representative image for OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

TL;DR - OSReward is a benchmark suite for evaluating vision-language reward models that judge computer-use agent trajectories. It reveals widespread leniency toward failed runs and introduces open OS-Shepherd models designed to provide reliable judgments at lower cost.

  • Includes realistic, human-verified trajectories across platforms, plus hard-case and fine-grained evaluation variants.
  • Finds that state-of-the-art judges often misclassify failed trajectories as successful.
  • Releases OS-Shepherd-100K, an open corpus of reasoning-annotated trajectory judgments.
  • OS-Shepherd 9B and 35B reportedly match commercial judges at 30–60% lower cost than frontier alternatives.

view merged work →