🛰️ Daily AI Frontier
‹ back to 2026-07-30

Towards Grounded GI Endoscopy VQA via Multi-Task Learning on Small VLMs

arXiv cs.CV Medical/Healthcare AI Itbaan Safwan, Ramail Khan, Muhammad Annas Shaikh, Muhammad Atif Tahir 2026-07-29
Representative image for Towards Grounded GI Endoscopy VQA via Multi-Task Learning on Small VLMs

TL;DR - A multi-task fine-tuning method improves both answer accuracy and visual grounding for small vision-language models on GI endoscopy VQA, supporting more evidence-aligned clinical responses.

  • Reuses expert polyp masks and Grad-CAM-derived weak supervision to create grounding tasks with minimal extra annotation.
  • Fine-tunes three small VLM backbones using low-rank adaptation under matched VQA-only and multi-task settings.
  • Consistently improves accuracy and alignment between answer tokens and relevant image regions.
  • Evaluated on Kvasir-VQA-x1 with both in-distribution and out-of-distribution data.

view merged work →