李飞飞:语言之后,AI要学会用"身体"理解世界
TL;DR - Fei-Fei Li argues that AI’s next frontier is spatial and physical intelligence: systems that can model, navigate, and act in the real world rather than merely process language. This could unlock robotics for care, manufacturing, agriculture, disaster response, and scientific discovery, but progress is constrained primarily by scarce physical-interaction data.
- Li defines world models through a “3P” framework: rendering how a world looks, simulating its physics and dynamics, and planning actions within it.
- Robotics is substantially harder than language modeling because it operates in a high-dimensional 3D world with limited sensor and interaction data.
- Promising data and model approaches include egocentric video, teleoperation, and vision-language-action architectures, though consumer robotics lacks the mature usage loop needed for crowdsourced data collection.
- World Labs is targeting high-fidelity simulation rather than simple 3D rendering, aiming to support robotics, games, visual effects, architecture, and industrial applications.