Question-Guided Evidence Acquisition for Multimodal Visual Question Answering
Ranking
Observed public metrics from 1 member.
Merged summary
TL;DR - Q-Guide is a lightweight agent that improves document visual question answering by iteratively identifying missing evidence and invoking targeted tools for text reading, zooming, or region grounding. It substantially outperforms direct prompting and recent multi-agent systems, showing that focused inference-time perception matters more than complex orchestration.
- Q-Guide replaces single-pass page encoding with question-guided evidence acquisition over multiple deliberate rounds.
- It achieves 65.0% on DocVQA2026 versus 40.0% for baselines, and 32.4% on Manga109 versus 24.4%.
- Improvements hold across Claude Opus 4.6, Sonnet 4.6, and Opus 4.5 backbones.
- Most gains emerge within two to three rounds; planners, routers, and collaborating agents provide no additional benefit.
Sources (1)
Question-Guided Evidence Acquisition for Multimodal Visual Question Answering
TL;DR - Q-Guide is a lightweight agent that improves document visual question answering by iteratively identifying missing evidence and invoking targeted tools for text reading, zooming, or region grounding. It substantially outperforms direct prompting and recent multi-agent systems, showing that focused inference-time perception matters more than complex orchestration.
- Q-Guide replaces single-pass page encoding with question-guided evidence acquisition over multiple deliberate rounds.
- It achieves 65.0% on DocVQA2026 versus 40.0% for baselines, and 32.4% on Manga109 versus 24.4%.
- Improvements hold across Claude Opus 4.6, Sonnet 4.6, and Opus 4.5 backbones.
- Most gains emerge within two to three rounds; planners, routers, and collaborating agents provide no additional benefit.