RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents
Ranking
Overall
87
Content
95
Popularity
67
Observed public metrics from 1 member.
Merged summary
TL;DR - RecreationWorld is a cross-platform framework for training and evaluating hybrid computer-use agents that combine GUI exploration, coding, execution, and visual verification. Its results show meaningful transfer from recreation-based training, but also reveal substantial gaps in reproducing interactive behavior and computed outputs.
- RecreationWorld supports reproducible tasks across Ubuntu, macOS, Windows, Android, and Web through unified GUI-control and coding tools.
- Training on trajectories generated from open-source applications improved performance across five out-of-distribution coding and hybrid computer-use benchmarks.
- RecreationBench contains 250 held-out tasks with reference-validated programmatic and visual assertions spanning multiple interaction depths.
- GPT-6 Astra scored 58.1% overall but passed every programmatic test on only 2.8% of tasks; agents handled static interfaces better than interactions and computed outputs.
Sources (1)
RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents
Public signals
Hugging Face upvotes 78
TL;DR - RecreationWorld is a cross-platform framework for training and evaluating hybrid computer-use agents that combine GUI exploration, coding, execution, and visual verification. Its results show meaningful transfer from recreation-based training, but also reveal substantial gaps in reproducing interactive behavior and computed outputs.
- RecreationWorld supports reproducible tasks across Ubuntu, macOS, Windows, Android, and Web through unified GUI-control and coding tools.
- Training on trajectories generated from open-source applications improved performance across five out-of-distribution coding and hybrid computer-use benchmarks.
- RecreationBench contains 250 held-out tasks with reference-validated programmatic and visual assertions spanning multiple interaction depths.
- GPT-6 Astra scored 58.1% overall but passed every programmatic test on only 2.8% of tasks; agents handled static interfaces better than interactions and computed outputs.