🛰️ Daily AI Frontier
‹ back to 2026-08-27

CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators

arXiv cs.RO Multimodal & Generative Kechen Liu, Ola Shorinwa 2026-08-27

TL;DR - CLAP is a cross-embodiment, action-conditioned video world model that learns shared physical dynamics from heterogeneous human and robot videos. It enables zero-shot simulation across robot platforms while matching or exceeding leading single-embodiment models in environments such as DROID.

  • Unifies disparate action spaces through end-effector poses, language instructions, and learned latent actions.
  • Uses curriculum learning to first acquire physical priors from unlabeled videos, then ground them in robot action spaces for zero-shot deployment.
  • Supports multiple robot morphologies, including DROID, Bridge, bimanual YAM robots, and G1 humanoids.
  • Few-shot adaptation further improves single-embodiment performance; the authors open-source the code and models.

view merged work →