🛰️ Daily AI Frontier
‹ back to 2026-07-30

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

Research LLM Agents

Ranking

Overall 87
Content 95
Popularity 70

Observed public metrics from 1 member.

Representative image for OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

Merged summary

TL;DR - OSReward is a benchmark suite for evaluating vision-language reward models that judge computer-use agent trajectories. It reveals widespread leniency toward failed runs and introduces open OS-Shepherd models designed to provide reliable judgments at lower cost.

  • Includes realistic, human-verified trajectories across platforms, plus hard-case and fine-grained evaluation variants.
  • Finds that state-of-the-art judges often misclassify failed trajectories as successful.
  • Releases OS-Shepherd-100K, an open corpus of reasoning-annotated trajectory judgments.
  • OS-Shepherd 9B and 35B reportedly match commercial judges at 30–60% lower cost than frontier alternatives.

Sources (1)

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

arXiv cs.AI Qiushi Sun, Kanzhi Cheng, Yian Wang, Bowen Yang, Hang Yan, Liheng Chen, Fangzhi Xu, Zichen Ding, Nuo Chen, Jialin Cao, Xingdong Gong, Zehao Li, Kaiming Jin, Xinfeng Yuan, Zhoumianze Liu, Jingyang Gong, Zhangyue Yin, Jiahui Gao, Zhiyong Wu, Tianbao Xie, Jianbing Zhang, Ben Kao, Lingpeng Kong 2026-07-30 arXiv:2607.28609
Public signals Hugging Face upvotes 73
Providers: Hugging Face · Upvotes 73 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-08-28 14:31:40.336984 UTC

TL;DR - OSReward is a benchmark suite for evaluating vision-language reward models that judge computer-use agent trajectories. It reveals widespread leniency toward failed runs and introduces open OS-Shepherd models designed to provide reliable judgments at lower cost.

  • Includes realistic, human-verified trajectories across platforms, plus hard-case and fine-grained evaluation variants.
  • Finds that state-of-the-art judges often misclassify failed trajectories as successful.
  • Releases OS-Shepherd-100K, an open corpus of reasoning-annotated trajectory judgments.
  • OS-Shepherd 9B and 35B reportedly match commercial judges at 30–60% lower cost than frontier alternatives.
item →