🛰️ Daily AI Frontier
‹ back to 2026-08-22

Question-Guided Evidence Acquisition for Multimodal Visual Question Answering

arXiv cs.CV LLM Agents Alin-Ionut Popa 2026-08-20
Representative image for Question-Guided Evidence Acquisition for Multimodal Visual Question Answering

TL;DR - Q-Guide is a lightweight agent that improves document visual question answering by iteratively identifying missing evidence and invoking targeted tools for text reading, zooming, or region grounding. It substantially outperforms direct prompting and recent multi-agent systems, showing that focused inference-time perception matters more than complex orchestration.

  • Q-Guide replaces single-pass page encoding with question-guided evidence acquisition over multiple deliberate rounds.
  • It achieves 65.0% on DocVQA2026 versus 40.0% for baselines, and 32.4% on Manga109 versus 24.4%.
  • Improvements hold across Claude Opus 4.6, Sonnet 4.6, and Opus 4.5 backbones.
  • Most gains emerge within two to three rounds; planners, routers, and collaborating agents provide no additional benefit.

view merged work →