Towards Grounded GI Endoscopy VQA via Multi-Task Learning on Small VLMs
Ranking
Overall
73
Content
85
Popularity
44
Observed public metrics from 1 member.
Merged summary
TL;DR - A multi-task fine-tuning method improves both answer accuracy and visual grounding for small vision-language models on GI endoscopy VQA, supporting more evidence-aligned clinical responses.
- Reuses expert polyp masks and Grad-CAM-derived weak supervision to create grounding tasks with minimal extra annotation.
- Fine-tunes three small VLM backbones using low-rank adaptation under matched VQA-only and multi-task settings.
- Consistently improves accuracy and alignment between answer tokens and relevant image regions.
- Evaluated on Kvasir-VQA-x1 with both in-distribution and out-of-distribution data.
Sources (1)
Towards Grounded GI Endoscopy VQA via Multi-Task Learning on Small VLMs
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - A multi-task fine-tuning method improves both answer accuracy and visual grounding for small vision-language models on GI endoscopy VQA, supporting more evidence-aligned clinical responses.
- Reuses expert polyp masks and Grad-CAM-derived weak supervision to create grounding tasks with minimal extra annotation.
- Fine-tunes three small VLM backbones using low-rank adaptation under matched VQA-only and multi-task settings.
- Consistently improves accuracy and alignment between answer tokens and relevant image regions.
- Evaluated on Kvasir-VQA-x1 with both in-distribution and out-of-distribution data.