🛰️ Daily AI Frontier
‹ back to 2026-08-16

H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models

arXiv cs.RO Multimodal & Generative Dingyi Rong, Yue Shi, Chaofan Ma, Jiezhang Cao, Zongrui Wang, Zeyu Zhang, Yao Mu, Guangtao Zhai, Ning Liu 2026-08-13
Representative image for H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models

TL;DR - H2R-Bench evaluates whether video world models can translate egocentric human demonstrations into robot-manipulation videos under embodiment constraints. Results show current models struggle with consistent robot embodiment, functional interactions, and successful task execution.

  • Each instance includes a human video, target embodiment constraints, and annotations for goals, actions, contacts, and object responses.
  • Evaluation covers goal and action completion, functional-contact transfer, embodiment correctness, and overall video quality.
  • Eleven video-generation models are benchmarked across six manipulation families and two robot embodiments.
  • The benchmark targets scalable generation of robot-centric training data from abundant human demonstrations.

view merged work →