无监督多模态融合:视觉/IMU/激光雷达互监督,免标注实现厘米级定位
TL;DR - This overview describes label-free fusion of cameras, IMUs, and LiDAR for robust centimeter-level localization in robotics and autonomous driving. Cross-sensor supervision could reduce annotation costs while improving resilience when individual sensors degrade.
- LiDAR geometry can supervise visual depth, while IMU measurements constrain long-term visual/LiDAR odometry drift.
- Self-supervised losses and cross-modal contrastive learning align 2D images with 3D point-cloud features without ground truth.
- Tightly coupled factor graphs or extended Kalman filters combine complementary sensor observations for higher accuracy and robustness.
- Remaining challenges include dynamic or degenerate environments, real-time computation, sensor failures, and cross-scene generalization.