🛰️ Daily AI Frontier
‹ back to 2026-08-16

H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models

Research Multimodal & Generative

Ranking

Overall 78
Content 85
Popularity 63

Observed public metrics from 1 member.

Representative image for H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models

Merged summary

TL;DR - H2R-Bench evaluates whether video world models can translate egocentric human demonstrations into robot-manipulation videos under embodiment constraints. Results show current models struggle with consistent robot embodiment, functional interactions, and successful task execution.

  • Each instance includes a human video, target embodiment constraints, and annotations for goals, actions, contacts, and object responses.
  • Evaluation covers goal and action completion, functional-contact transfer, embodiment correctness, and overall video quality.
  • Eleven video-generation models are benchmarked across six manipulation families and two robot embodiments.
  • The benchmark targets scalable generation of robot-centric training data from abundant human demonstrations.

Sources (1)

H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models

arXiv cs.RO Dingyi Rong, Yue Shi, Chaofan Ma, Jiezhang Cao, Zongrui Wang, Zeyu Zhang, Yao Mu, Guangtao Zhai, Ning Liu 2026-08-13 arXiv:2608.13049
Public signals Hugging Face upvotes 20
Providers: Hugging Face · Upvotes 20 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-15 14:32:52.809050 UTC

TL;DR - H2R-Bench evaluates whether video world models can translate egocentric human demonstrations into robot-manipulation videos under embodiment constraints. Results show current models struggle with consistent robot embodiment, functional interactions, and successful task execution.

  • Each instance includes a human video, target embodiment constraints, and annotations for goals, actions, contacts, and object responses.
  • Evaluation covers goal and action completion, functional-contact transfer, embodiment correctness, and overall video quality.
  • Eleven video-generation models are benchmarked across six manipulation families and two robot embodiments.
  • The benchmark targets scalable generation of robot-centric training data from abundant human demonstrations.
item →