几何的「反攻」!港科大谭平:从局部先验到全局一致,3D 几何如何增强视觉大模型 | ECCV 2026
TL;DR - HKUST professor Tan Ping presents a hybrid approach to scalable 3D reconstruction and scene generation: learned models supply strong local visual priors, while explicit geometry and global optimization enforce consistency across views and long camera trajectories.
- For large-scale structure-from-motion, sampled keyframes and token compression form a lightweight neural scene representation, reducing Transformer memory demands while localizing additional frames.
- Global 3R combines feed-forward predictions with pose averaging and bundle adjustment; on reported tests, geometric optimization raised the share of relative rotations within 1° from roughly 70% to over 90%.
- ControlRoom and SpatialGen use explicit 3D layouts and multimodal, multi-view diffusion to improve object placement and spatial consistency; the SpatialGen dataset includes 57,000 rooms and about five million multimodal renderings.
- SpatialCraft infers a coarse 3D proxy from one image, converts it into a 3D Gaussian representation, and applies video diffusion refinement to support more stable long-range scene navigation.