🛰️ Daily AI Frontier
‹ back to 2026-07-20

An Exam for Active Observers

arXiv cs.CV Multimodal & Generative Jiarui Zhang, Muzi Tao, Shangshang Wang, Ollie Liu, Xuezhe Ma, Willie Neiswanger 2026-07-17

TL;DR - ActiveVision is a 17-task benchmark testing whether multimodal models can iteratively redirect visual attention. Frontier models perform dramatically below humans, suggesting they lack a robust perception-reasoning feedback loop.

  • GPT-5.5 achieved the best model score at just 10.6%, versus a 96.1% human average.
  • It scored zero on 11 of 17 tasks; Claude Fable 5 solved only 3.5%.
  • Allowing models to write and run vision code did not close much of the gap because the code was unreliable on realistic images.
  • The findings motivate architectures and training objectives designed explicitly for active visual observation.

view merged work →