🛰️ Daily AI Frontier
‹ back to 2026-07-25

Do Pathology Vision-Language Models Truly See Pathology?

Research Medical/Healthcare AI

Merged summary

TL;DR - PathBind is a 2,600-sample benchmark testing whether pathology vision-language models genuinely connect answers to visual evidence. Results reveal a substantial gap between VQA accuracy and visual-semantic grounding.

  • Gemini-3-Pro averaged 53.5% across five pathology VQA benchmarks without visual input, exposing textual shortcuts.
  • Domain training improved accuracy without proportional visual binding gains; Patho-R1-7B underperformed Qwen2.5-VL-7B on multimodal gain and attention IoU.
  • Entity-level attention was diffuse and weakly query-specific on PathVG.
  • Evaluations covered 18 VLMs on VQA and 10 on region-level grounding.

Sources (1)

Do Pathology Vision-Language Models Truly See Pathology?

arXiv cs.CV Chengyang Zhang, Wenchuan Zhang, Bo Li, Xinyu Liu, Jiaming Yang, Mengran Li, Chenxun Deng, Jie Chen, Yang Zhang, Wei Ju, Yuhao Yi, Hong Bu, Jiancheng Lv 2026-07-23 arXiv:2607.21065

TL;DR - PathBind is a 2,600-sample benchmark testing whether pathology vision-language models genuinely connect answers to visual evidence. Results reveal a substantial gap between VQA accuracy and visual-semantic grounding.

  • Gemini-3-Pro averaged 53.5% across five pathology VQA benchmarks without visual input, exposing textual shortcuts.
  • Domain training improved accuracy without proportional visual binding gains; Patho-R1-7B underperformed Qwen2.5-VL-7B on multimodal gain and attention IoU.
  • Entity-level attention was diffuse and weakly query-specific on PathVG.
  • Evaluations covered 18 VLMs on VQA and 10 on region-level grounding.
item →