ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow
TL;DR - ShadowDancer enables frame-level control of video world models by learning appearance-invariant dynamics from paired videos that replay the same action with different visuals. Demonstrated actions can then transfer to new environments without labels, motion estimators, or fine-tuning.
- “Shadow pairs” preserve dynamics while independently varying appearance.
- Cross-shadow prediction isolates a unified action representation by discarding visual differences.
- The representation controls a block-causal world model across diverse dynamics families.
- It achieved an average 86% blinded win rate against latent-action and interactive-world-model baselines in rollout comparisons.