Do Pathology Vision-Language Models Truly See Pathology?
Merged summary
TL;DR - PathBind is a 2,600-sample benchmark testing whether pathology vision-language models genuinely connect answers to visual evidence. Results reveal a substantial gap between VQA accuracy and visual-semantic grounding.
- Gemini-3-Pro averaged 53.5% across five pathology VQA benchmarks without visual input, exposing textual shortcuts.
- Domain training improved accuracy without proportional visual binding gains; Patho-R1-7B underperformed Qwen2.5-VL-7B on multimodal gain and attention IoU.
- Entity-level attention was diffuse and weakly query-specific on PathVG.
- Evaluations covered 18 VLMs on VQA and 10 on region-level grounding.
Sources (1)
Do Pathology Vision-Language Models Truly See Pathology?
TL;DR - PathBind is a 2,600-sample benchmark testing whether pathology vision-language models genuinely connect answers to visual evidence. Results reveal a substantial gap between VQA accuracy and visual-semantic grounding.
- Gemini-3-Pro averaged 53.5% across five pathology VQA benchmarks without visual input, exposing textual shortcuts.
- Domain training improved accuracy without proportional visual binding gains; Patho-R1-7B underperformed Qwen2.5-VL-7B on multimodal gain and attention IoU.
- Entity-level attention was diffuse and weakly query-specific on PathVG.
- Evaluations covered 18 VLMs on VQA and 10 on region-level grounding.