🛰️ Daily AI Frontier
‹ back to 2026-07-31

ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow

arXiv cs.CV Multimodal & Generative Jin Cao, Zian Meng, Kaipeng Zhang 2026-07-30

TL;DR - ShadowDancer enables frame-level control of video world models by learning appearance-invariant dynamics from paired videos that replay the same action with different visuals. Demonstrated actions can then transfer to new environments without labels, motion estimators, or fine-tuning.

  • “Shadow pairs” preserve dynamics while independently varying appearance.
  • Cross-shadow prediction isolates a unified action representation by discarding visual differences.
  • The representation controls a block-causal world model across diverse dynamics families.
  • It achieved an average 86% blinded win rate against latent-action and interactive-world-model baselines in rollout comparisons.

view merged work →