Humans count the boxes in this image, hidden ones included, with 82.1% accuracy. The best off-the-shelf multimodal model manages 17.7%. Spatial-IQ is a diagnostic benchmark from NVIDIA Research that breaks 3D object counting into nine perceptual and cognitive sub-tasks, from counting columns to inferring the blocks that must be underneath to hold the structure up, and scores each one separately. Training on those sub-tasks lifted Qwen2.5-VL-32B object-counting accuracy from 2.9% to 62.6%. For researchers developing multimodal reasoning systems, this creates a practical loop: identify where spatial reasoning fails, target the missing capability, and verify that improvement reflects composition, not just a better final score.
Ranking
Overall
82
Content
95
Popularity
N/A
No observed public metrics; popularity remains neutral/archived.
Merged summary
TL;DR - NVIDIA Research’s Spatial-IQ benchmark diagnoses 3D object-counting failures across nine perceptual and reasoning sub-tasks. Targeted training raised Qwen2.5-VL-32B accuracy from 2.9% to 62.6%, showing the value of compositional supervision.
- Humans achieved 82.1% accuracy, versus 17.7% for the best off-the-shelf multimodal model.
- Tasks include counting visible structures and inferring hidden supporting blocks.
- Separate sub-task scores reveal specific spatial-reasoning weaknesses.
- The benchmark helps verify whether gains reflect improved composition rather than only higher final scores.
Sources (1)
Humans count the boxes in this image, hidden ones included, with 82.1% accuracy. The best off-the-shelf multimodal model manages 17.7%. Spatial-IQ is a diagnostic benchmark from NVIDIA Research that breaks 3D object counting into nine perceptual and cognitive sub-tasks, from counting columns to inferring the blocks that must be underneath to hold the structure up, and scores each one separately. Training on those sub-tasks lifted Qwen2.5-VL-32B object-counting accuracy from 2.9% to 62.6%. For researchers developing multimodal reasoning systems, this creates a practical loop: identify where spatial reasoning fails, target the missing capability, and verify that improvement reflects composition, not just a better final score.
Public signals
N/A
TL;DR - NVIDIA Research’s Spatial-IQ benchmark diagnoses 3D object-counting failures across nine perceptual and reasoning sub-tasks. Targeted training raised Qwen2.5-VL-32B accuracy from 2.9% to 62.6%, showing the value of compositional supervision.
- Humans achieved 82.1% accuracy, versus 17.7% for the best off-the-shelf multimodal model.
- Tasks include counting visible structures and inferring hidden supporting blocks.
- Separate sub-task scores reveal specific spatial-reasoning weaknesses.
- The benchmark helps verify whether gains reflect improved composition rather than only higher final scores.