🛰️ Daily AI Frontier
‹ back to 2026-08-27

CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators

Research Multimodal & Generative

Ranking

Overall 85
Content 95
Popularity 61

Observed public metrics from 1 member.

Merged summary

TL;DR - CLAP is a cross-embodiment, action-conditioned video world model that learns shared physical dynamics from heterogeneous human and robot videos. It enables zero-shot simulation across robot platforms while matching or exceeding leading single-embodiment models in environments such as DROID.

  • Unifies disparate action spaces through end-effector poses, language instructions, and learned latent actions.
  • Uses curriculum learning to first acquire physical priors from unlabeled videos, then ground them in robot action spaces for zero-shot deployment.
  • Supports multiple robot morphologies, including DROID, Bridge, bimanual YAM robots, and G1 humanoids.
  • Few-shot adaptation further improves single-embodiment performance; the authors open-source the code and models.

Sources (1)

CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators

arXiv cs.RO Kechen Liu, Ola Shorinwa 2026-08-27 arXiv:2608.27406
Public signals Hugging Face upvotes 0 · Semantic Scholar citations 1 · Semantic Scholar influential citations 0
Providers: Hugging Face · Upvotes 0 OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 1 · Influential citations 0 X · N/A Fetched 2026-09-25 14:27:35.813732 UTC

TL;DR - CLAP is a cross-embodiment, action-conditioned video world model that learns shared physical dynamics from heterogeneous human and robot videos. It enables zero-shot simulation across robot platforms while matching or exceeding leading single-embodiment models in environments such as DROID.

  • Unifies disparate action spaces through end-effector poses, language instructions, and learned latent actions.
  • Uses curriculum learning to first acquire physical priors from unlabeled videos, then ground them in robot action spaces for zero-shot deployment.
  • Supports multiple robot morphologies, including DROID, Bridge, bimanual YAM robots, and G1 humanoids.
  • Few-shot adaptation further improves single-embodiment performance; the authors open-source the code and models.
item →