🛰️ Daily AI Frontier
‹ back to 2026-09-03

打破黑盒猜想:大模型通往真正「空间智能」的破局之路

雷峰网 (AI科技评论) Multimodal & Generative 2026-09-03
Representative image for 打破黑盒猜想:大模型通往真正「空间智能」的破局之路

TL;DR - SpatialSV trains multimodal large language models to internalize explicit 3D geometry, improving spatial reasoning without adding inference-time overhead. Its interpretable depth and point-cloud reconstructions also reveal whether failures stem from missing objects, lost spatial anchors, or poor viewpoint alignment.

  • SpatialSV lifts intermediate MLLM features into depth maps, camera-ray maps, and point clouds, supervised with explicit geometric losses during training.
  • The auxiliary 2D-to-3D projection and prediction modules are removed for inference, leaving no additional latency or memory cost.
  • Across eight MLLMs, lower 3D reconstruction error strongly correlates with higher spatial-question-answering accuracy, enabling visual diagnosis of internal representation failures.
  • With only 50% of text annotations, automatically generated 3D supervision raised Qwen2.5-VL-3B accuracy from 47.2% to 53.9%, approaching the 55.3% full-annotation result.

view merged work →