佐治亚理工学院徐丹飞:别被「视觉生成」骗了,视频预测≠机器人规划|RSS 2026
TL;DR - Georgia Tech’s Danfei Xu argues that realistic video prediction does not constitute robotic planning. His compositional world-model approach uses modular skills, factor graphs, and adaptive temporal attention to bridge predicted futures and executable actions.
- Factor graphs compose independently learned skills at test time while enforcing spatial, temporal, and multi-arm constraints.
- A “video-action gap” arises because abundant video data supports stronger generalization than scarce robot-action data.
- The Temporal Ratio measures whether control attends more to predicted future frames or the current observation.
- Policy guidance emphasizes future trajectories during reaching, then current visual and tactile feedback during contact and grasping.