不堆 Transformer,斯坦福吴佳俊如何用物理重新定义多模态融合?|ECCV 2026
Ranking
Overall
82
Content
95
Popularity
N/A
No observed public metrics; popularity remains neutral/archived.
Merged summary
TL;DR - Stanford’s Jiajun Wu outlined a physics-grounded approach to multimodal perception that aligns vision, sound, and touch through shared properties such as geometry, material, and stiffness. This could improve generalization when audio and tactile data are scarce, rather than relying only on end-to-end Transformer fusion.
- PhysDreamer distills physical properties such as spatially varying Young’s modulus from video diffusion models by matching differentiable simulations to generated reference videos.
- WonderPlay uses coarse 3D physics simulations—represented through depth or optical flow—to condition video generation for controllable, action-driven physical interactions.
- DiffImpact and RealImpact apply differentiable acoustic rendering and measured sound fields to infer impact and material-related properties from audio.
- DexSkin extends the framework to low-cost tactile sensing for robotic manipulation, treating touch as another observation of the same underlying physics.
Sources (1)
不堆 Transformer,斯坦福吴佳俊如何用物理重新定义多模态融合?|ECCV 2026
Public signals
N/A
TL;DR - Stanford’s Jiajun Wu outlined a physics-grounded approach to multimodal perception that aligns vision, sound, and touch through shared properties such as geometry, material, and stiffness. This could improve generalization when audio and tactile data are scarce, rather than relying only on end-to-end Transformer fusion.
- PhysDreamer distills physical properties such as spatially varying Young’s modulus from video diffusion models by matching differentiable simulations to generated reference videos.
- WonderPlay uses coarse 3D physics simulations—represented through depth or optical flow—to condition video generation for controllable, action-driven physical interactions.
- DiffImpact and RealImpact apply differentiable acoustic rendering and measured sound fields to infer impact and material-related properties from audio.
- DexSkin extends the framework to low-cost tactile sensing for robotic manipulation, treating touch as another observation of the same underlying physics.