ACM MM 2026 | SIS-Bench:无人机具身空间智能的自我认知能力评估基准与增强方法
Merged summary
TL;DR - SIS-Bench evaluates multimodal models’ spatial cognition and self-awareness in UAV video, revealing major weaknesses in understanding self-motion. Its optical-flow-based SIS-Motion method improves both benchmark performance and zero-shot navigation transfer.
- SIS-Bench contains 1,646 real UAV videos, 4,856 multiple-choice questions, and 13 perception, memory, and reasoning tasks.
- The best evaluated model scored 71.6% overall versus 91.7% for humans; Gemini-3-Flash showed a 25.1-point gap between spatial and self cognition.
- Adding explicit motion features raised Qwen2.5-VL-7B accuracy from 73.1% to 76.9%.
- Gains were limited on higher-level reasoning and planning, so the work does not yet demonstrate closed-loop autonomous flight.
Sources (1)
ACM MM 2026 | SIS-Bench:无人机具身空间智能的自我认知能力评估基准与增强方法
TL;DR - SIS-Bench evaluates multimodal models’ spatial cognition and self-awareness in UAV video, revealing major weaknesses in understanding self-motion. Its optical-flow-based SIS-Motion method improves both benchmark performance and zero-shot navigation transfer.
- SIS-Bench contains 1,646 real UAV videos, 4,856 multiple-choice questions, and 13 perception, memory, and reasoning tasks.
- The best evaluated model scored 71.6% overall versus 91.7% for humans; Gemini-3-Flash showed a 25.1-point gap between spatial and self cognition.
- Adding explicit motion features raised Qwen2.5-VL-7B accuracy from 73.1% to 76.9%.
- Gains were limited on higher-level reasoning and planning, so the work does not yet demonstrate closed-loop autonomous flight.