🛰️ Daily AI Frontier
‹ back to 2026-08-02

Humans count the boxes in this image, hidden ones included, with 82.1% accuracy. The best off-the-shelf multimodal model manages 17.7%. Spatial-IQ is a diagnostic benchmark from NVIDIA Research that breaks 3D object counting into nine perceptual and cognitive sub-tasks, from counting columns to inferring the blocks that must be underneath to hold the structure up, and scores each one separately. Training on those sub-tasks lifted Qwen2.5-VL-32B object-counting accuracy from 2.9% to 62.6%. For researchers developing multimodal reasoning systems, this creates a practical loop: identify where spatial reasoning fails, target the missing capability, and verify that improvement reflects composition, not just a better final score.

Multimodal & Generative @NVIDIAAI 2026-07-31

TL;DR - NVIDIA Research’s Spatial-IQ benchmark diagnoses 3D object-counting failures across nine perceptual and reasoning sub-tasks. Targeted training raised Qwen2.5-VL-32B accuracy from 2.9% to 62.6%, showing the value of compositional supervision.

  • Humans achieved 82.1% accuracy, versus 17.7% for the best off-the-shelf multimodal model.
  • Tasks include counting visible structures and inferring hidden supporting blocks.
  • Separate sub-task scores reveal specific spatial-reasoning weaknesses.
  • The benchmark helps verify whether gains reflect improved composition rather than only higher final scores.

view merged work →