An Exam for Active Observers
Merged summary
TL;DR - ActiveVision is a 17-task benchmark testing whether multimodal models can iteratively redirect visual attention. Frontier models perform dramatically below humans, suggesting they lack a robust perception-reasoning feedback loop.
- GPT-5.5 achieved the best model score at just 10.6%, versus a 96.1% human average.
- It scored zero on 11 of 17 tasks; Claude Fable 5 solved only 3.5%.
- Allowing models to write and run vision code did not close much of the gap because the code was unreliable on realistic images.
- The findings motivate architectures and training objectives designed explicitly for active visual observation.
Sources (1)
An Exam for Active Observers
TL;DR - ActiveVision is a 17-task benchmark testing whether multimodal models can iteratively redirect visual attention. Frontier models perform dramatically below humans, suggesting they lack a robust perception-reasoning feedback loop.
- GPT-5.5 achieved the best model score at just 10.6%, versus a 96.1% human average.
- It scored zero on 11 of 17 tasks; Claude Fable 5 solved only 3.5%.
- Allowing models to write and run vision code did not close much of the gap because the code was unreliable on realistic images.
- The findings motivate architectures and training objectives designed explicitly for active visual observation.