行动即推理:南洋理工大学等提出S-Agent,空间智能迈入Agent时代
Merged summary
TL;DR - S-Agent turns spatial reasoning into an agentic workflow where a VLM plans actions and specialized tools supply 2D, 3D, and geometric evidence. It improves zero-shot benchmarks and enables action-trace distillation into smaller models.
- S-Agent maintains scene memory while iteratively selecting frames, locating entities, estimating depth and pose, aligning views, and deriving spatial relations.
- It achieves 46.4% on MMSI-Bench and 60.0% on ViewSpatial-Bench in zero-shot settings.
- About 292,000 filtered S-300K trajectories were used to fine-tune Qwen3-VL-8B, raising scores to 41.6% on MMSI-Bench and 46.8% on ViewSpatial-Bench.
- Ablations indicate that spatial experts converting noisy geometric outputs into task-relevant evidence are more useful than exposing raw depth, pose, or point-cloud data directly.
Sources (1)
行动即推理:南洋理工大学等提出S-Agent,空间智能迈入Agent时代
TL;DR - S-Agent turns spatial reasoning into an agentic workflow where a VLM plans actions and specialized tools supply 2D, 3D, and geometric evidence. It improves zero-shot benchmarks and enables action-trace distillation into smaller models.
- S-Agent maintains scene memory while iteratively selecting frames, locating entities, estimating depth and pose, aligning views, and deriving spatial relations.
- It achieves 46.4% on MMSI-Bench and 60.0% on ViewSpatial-Bench in zero-shot settings.
- About 292,000 filtered S-300K trajectories were used to fine-tune Qwen3-VL-8B, raising scores to 41.6% on MMSI-Bench and 46.8% on ViewSpatial-Bench.
- Ablations indicate that spatial experts converting noisy geometric outputs into task-relevant evidence are more useful than exposing raw depth, pose, or point-cloud data directly.