🛰️ Daily AI Frontier
‹ back to 2026-07-20

An Exam for Active Observers

Research Multimodal & Generative

Merged summary

TL;DR - ActiveVision is a 17-task benchmark testing whether multimodal models can iteratively redirect visual attention. Frontier models perform dramatically below humans, suggesting they lack a robust perception-reasoning feedback loop.

  • GPT-5.5 achieved the best model score at just 10.6%, versus a 96.1% human average.
  • It scored zero on 11 of 17 tasks; Claude Fable 5 solved only 3.5%.
  • Allowing models to write and run vision code did not close much of the gap because the code was unreliable on realistic images.
  • The findings motivate architectures and training objectives designed explicitly for active visual observation.

Sources (1)

An Exam for Active Observers

arXiv cs.CV Jiarui Zhang, Muzi Tao, Shangshang Wang, Ollie Liu, Xuezhe Ma, Willie Neiswanger 2026-07-17 arXiv:2607.16165

TL;DR - ActiveVision is a 17-task benchmark testing whether multimodal models can iteratively redirect visual attention. Frontier models perform dramatically below humans, suggesting they lack a robust perception-reasoning feedback loop.

  • GPT-5.5 achieved the best model score at just 10.6%, versus a 96.1% human average.
  • It scored zero on 11 of 17 tasks; Claude Fable 5 solved only 3.5%.
  • Allowing models to write and run vision code did not close much of the gap because the code was unreliable on realistic images.
  • The findings motivate architectures and training objectives designed explicitly for active visual observation.
item →