CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators
Ranking
Overall
85
Content
95
Popularity
61
Observed public metrics from 1 member.
Merged summary
TL;DR - CLAP is a cross-embodiment, action-conditioned video world model that learns shared physical dynamics from heterogeneous human and robot videos. It enables zero-shot simulation across robot platforms while matching or exceeding leading single-embodiment models in environments such as DROID.
- Unifies disparate action spaces through end-effector poses, language instructions, and learned latent actions.
- Uses curriculum learning to first acquire physical priors from unlabeled videos, then ground them in robot action spaces for zero-shot deployment.
- Supports multiple robot morphologies, including DROID, Bridge, bimanual YAM robots, and G1 humanoids.
- Few-shot adaptation further improves single-embodiment performance; the authors open-source the code and models.
Sources (1)
CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators
Public signals
Hugging Face upvotes 0 · Semantic Scholar citations 1 · Semantic Scholar influential citations 0
TL;DR - CLAP is a cross-embodiment, action-conditioned video world model that learns shared physical dynamics from heterogeneous human and robot videos. It enables zero-shot simulation across robot platforms while matching or exceeding leading single-embodiment models in environments such as DROID.
- Unifies disparate action spaces through end-effector poses, language instructions, and learned latent actions.
- Uses curriculum learning to first acquire physical priors from unlabeled videos, then ground them in robot action spaces for zero-shot deployment.
- Supports multiple robot morphologies, including DROID, Bridge, bimanual YAM robots, and G1 humanoids.
- Few-shot adaptation further improves single-embodiment performance; the authors open-source the code and models.