IJCAI 2026 专访:机器人想学会人类动作,还差一座桥 | GAIR Paper 118
Ranking
No observed public metrics; popularity remains neutral/archived.
Merged summary
TL;DR - An interview with Tsinghua PhD student Feng Zhiyuan on an IJCAI 2026 survey (Tsinghua / HKUST / MSRA) arguing that all methods for teaching robots from human video are really building the same "representation bridge" between unlabeled human video and robot actions. It matters because human video (HowTo100M's 136M clips, Ego4D's 3600+ hours) is the only data source cheap enough to scale past robot datasets like Open X-Embodiment and DROID.
- Four routes are unified as a choice of where to place the supervision layer: latent actions (LAPA), explicit 2D cues (ATM point tracks, Magma's Trace-of-Mark), explicit 3D hand trajectories via MANO (EgoVLA, H-RDT, Being-H0, VITRA), and world models (GR-1/GR-2, and the newer WAM adding an action expert on a world-model backbone).
- The routes are complementary, not competing — current practice treats 3D trajectory supervision as near-mandatory in VLA training, 2D as auxiliary, and world models as stackable with pixel- or latent-level losses.
- Three open challenges: semantic/interaction-based video segmentation instead of fixed time windows; simultaneous embodiment gap (human hand → low-DoF gripper is underconstrained) and viewpoint gap; and benchmarks (LIBERO, CALVIN, SIMPLER) that may not predict real deployment, with proposed fixes around transfer efficiency under matched robot-data and compute budgets.
- Candid assessment: world models are more research-coherent but underperform theory, industry's best demos still come from VLA+RL without true generalization, and robotics lacks a converged "recipe" (unlike Transformer pretraining for LLMs or DiT for video) — so an embodied "GPT-3.5 moment" is judged still far off.
Sources (1)
IJCAI 2026 专访:机器人想学会人类动作,还差一座桥 | GAIR Paper 118
TL;DR - An interview with Tsinghua PhD student Feng Zhiyuan on an IJCAI 2026 survey (Tsinghua / HKUST / MSRA) arguing that all methods for teaching robots from human video are really building the same "representation bridge" between unlabeled human video and robot actions. It matters because human video (HowTo100M's 136M clips, Ego4D's 3600+ hours) is the only data source cheap enough to scale past robot datasets like Open X-Embodiment and DROID.
- Four routes are unified as a choice of where to place the supervision layer: latent actions (LAPA), explicit 2D cues (ATM point tracks, Magma's Trace-of-Mark), explicit 3D hand trajectories via MANO (EgoVLA, H-RDT, Being-H0, VITRA), and world models (GR-1/GR-2, and the newer WAM adding an action expert on a world-model backbone).
- The routes are complementary, not competing — current practice treats 3D trajectory supervision as near-mandatory in VLA training, 2D as auxiliary, and world models as stackable with pixel- or latent-level losses.
- Three open challenges: semantic/interaction-based video segmentation instead of fixed time windows; simultaneous embodiment gap (human hand → low-DoF gripper is underconstrained) and viewpoint gap; and benchmarks (LIBERO, CALVIN, SIMPLER) that may not predict real deployment, with proposed fixes around transfer efficiency under matched robot-data and compute budgets.
- Candid assessment: world models are more research-coherent but underperform theory, industry's best demos still come from VLA+RL without true generalization, and robotics lacks a converged "recipe" (unlike Transformer pretraining for LLMs or DiT for video) — so an embodied "GPT-3.5 moment" is judged still far off.