Towards Grounded GI Endoscopy VQA via Multi-Task Learning on Small VLMs
TL;DR - A multi-task fine-tuning method improves both answer accuracy and visual grounding for small vision-language models on GI endoscopy VQA, supporting more evidence-aligned clinical responses.
- Reuses expert polyp masks and Grad-CAM-derived weak supervision to create grounding tasks with minimal extra annotation.
- Fine-tunes three small VLM backbones using low-rank adaptation under matched VQA-only and multi-task settings.
- Consistently improves accuracy and alignment between answer tokens and relevant image regions.
- Evaluated on Kvasir-VQA-x1 with both in-distribution and out-of-distribution data.