🛰️ Daily AI Frontier
‹ back to 2026-07-22

PathAgentBench: Benchmarking Evidence-Seeking Vision-Language Models on Whole-Slide Pathology Image

Research Medical/Healthcare AI

Ranking

Overall 78
Content 95
Popularity 40

Observed public metrics from 1 member.

Merged summary

TL;DR - PathAgentBench evaluates whether vision-language models can actively find and integrate diagnostic evidence in gigapixel pathology slides. Results expose a major gap between reasoning over curated regions and autonomously locating relevant tissue.

  • Covers 1,822 TCGA slides and 17,135 diagnostic paths annotated by ten board-certified pathologists.
  • Tests cross-modal matching, diagnostic-region localization, and multi-scale reasoning across 20 models.
  • Leading open-weight models exceed 93% multi-scale reasoning accuracy, but the best localization mean IoU remains below 0.09.
  • Autonomous evidence discovery deteriorates sharply at higher magnifications, reaching a 0.020 unconditional hit rate.

Sources (1)

PathAgentBench: Benchmarking Evidence-Seeking Vision-Language Models on Whole-Slide Pathology Image

arXiv cs.CV Dankai Liao, Tianyi Zhang, Yufeng Wu, Xinyue Zhang, Qiaochu Xue, Zeyu Liu, Dachun Zhao, Linghan Cai, Yueming Jin 2026-07-21 arXiv:2607.19261
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-07-31 14:12:04.426022 UTC

TL;DR - PathAgentBench evaluates whether vision-language models can actively find and integrate diagnostic evidence in gigapixel pathology slides. Results expose a major gap between reasoning over curated regions and autonomously locating relevant tissue.

  • Covers 1,822 TCGA slides and 17,135 diagnostic paths annotated by ten board-certified pathologists.
  • Tests cross-modal matching, diagnostic-region localization, and multi-scale reasoning across 20 models.
  • Leading open-weight models exceed 93% multi-scale reasoning accuracy, but the best localization mean IoU remains below 0.09.
  • Autonomous evidence discovery deteriorates sharply at higher magnifications, reaching a 0.020 unconditional hit rate.
item →