OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
Ranking
Overall
87
Content
95
Popularity
70
Observed public metrics from 1 member.
Merged summary
TL;DR - OSReward is a benchmark suite for evaluating vision-language reward models that judge computer-use agent trajectories. It reveals widespread leniency toward failed runs and introduces open OS-Shepherd models designed to provide reliable judgments at lower cost.
- Includes realistic, human-verified trajectories across platforms, plus hard-case and fine-grained evaluation variants.
- Finds that state-of-the-art judges often misclassify failed trajectories as successful.
- Releases OS-Shepherd-100K, an open corpus of reasoning-annotated trajectory judgments.
- OS-Shepherd 9B and 35B reportedly match commercial judges at 30–60% lower cost than frontier alternatives.
Sources (1)
OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
Public signals
Hugging Face upvotes 73
TL;DR - OSReward is a benchmark suite for evaluating vision-language reward models that judge computer-use agent trajectories. It reveals widespread leniency toward failed runs and introduces open OS-Shepherd models designed to provide reliable judgments at lower cost.
- Includes realistic, human-verified trajectories across platforms, plus hard-case and fine-grained evaluation variants.
- Finds that state-of-the-art judges often misclassify failed trajectories as successful.
- Releases OS-Shepherd-100K, an open corpus of reasoning-annotated trajectory judgments.
- OS-Shepherd 9B and 35B reportedly match commercial judges at 30–60% lower cost than frontier alternatives.