AI 开始学习「想象未来」:ECCV 2026背后,中国学者如何卡位世界模型
TL;DR - ECCV 2026 reflects a convergence of world models, 3D representations, and embodied AI, shifting the field from visually plausible future generation toward physically grounded prediction that guides action. Chinese researchers are contributing through new model definitions, evaluation benchmarks, and integrations with robotic control.
- ECCV 2026 features 13 workshops focused on world models or embodied AI, highlighting growing attention to construction, evaluation, reliability, and action-loop integration.
- World Action Models distinguish passive video prediction from systems that forecast action consequences and use them for decision-making.
- 4DWorldBench evaluates spatial structure, temporal continuity, physical consistency, and downstream usefulness rather than visual quality alone.
- VLA-JEPA and eWAM integrate latent world prediction with vision-language-action models, while Gaussian Splatting offers a differentiable 3D representation for dynamic environments.