五大高校联手发榜!首份机器人三视角世界模型评测结果出炉,榜单持续更新中
Ranking
Overall
47
Content
45
Popularity
N/A
No observed public metrics; popularity remains neutral/archived.
Merged summary
TL;DR - Five Chinese universities (Peking, Tsinghua, Beihang, SJTU, USTC) launched TriWorldBench, the first leaderboard for robot three-view world models, and published its inaugural weekly results. It shifts world-model evaluation from single-view visual fidelity toward cross-view consistency and physical/task understanding.
- Benchmark covers head, left-wrist, and right-wrist views over 500 synchronized tri-view episodes spanning 50 robot manipulation tasks; 19 signals across six dimensions (tri-view consistency, task alignment, physics/3D consistency, motion quality, temporal consistency, visual quality) roll up into a single TWB-Score.
- Metrics are routed to the most reliable camera — head view judges instruction compliance, arm trajectory, and final outcome; wrist views judge contact, grasp stability, and slippage — so wrist occlusion doesn't distort global task scoring.
- STATE annotations derived from reference robot trajectories mark the action phase, active arm, and whether each view should be moving or static, penalizing both frozen wrist views and spurious camera motion; visual/aesthetic scores are task-constrained so a "pretty but wrong" video can't win.
- Week-one top three: WoVR_Plus (CASIA-DRL), BetaBWM (TONGJI Spatial Intelligence Team), and Fysiverse-Video (Fysics AI); 14 teams registered, 10k+ site visits, 200+ toolkit downloads, with the leaderboard and GitHub repo open globally and updating continuously.
Sources (1)
五大高校联手发榜!首份机器人三视角世界模型评测结果出炉,榜单持续更新中
Public signals
N/A
TL;DR - Five Chinese universities (Peking, Tsinghua, Beihang, SJTU, USTC) launched TriWorldBench, the first leaderboard for robot three-view world models, and published its inaugural weekly results. It shifts world-model evaluation from single-view visual fidelity toward cross-view consistency and physical/task understanding.
- Benchmark covers head, left-wrist, and right-wrist views over 500 synchronized tri-view episodes spanning 50 robot manipulation tasks; 19 signals across six dimensions (tri-view consistency, task alignment, physics/3D consistency, motion quality, temporal consistency, visual quality) roll up into a single TWB-Score.
- Metrics are routed to the most reliable camera — head view judges instruction compliance, arm trajectory, and final outcome; wrist views judge contact, grasp stability, and slippage — so wrist occlusion doesn't distort global task scoring.
- STATE annotations derived from reference robot trajectories mark the action phase, active arm, and whether each view should be moving or static, penalizing both frozen wrist views and spurious camera motion; visual/aesthetic scores are task-constrained so a "pretty but wrong" video can't win.
- Week-one top three: WoVR_Plus (CASIA-DRL), BetaBWM (TONGJI Spatial Intelligence Team), and Fysiverse-Video (Fysics AI); 14 teams registered, 10k+ site visits, 200+ toolkit downloads, with the leaderboard and GitHub repo open globally and updating continuously.