🛰️ Daily AI Frontier
‹ back to 2026-07-22

PathAgentBench: Benchmarking Evidence-Seeking Vision-Language Models on Whole-Slide Pathology Image

arXiv cs.CV Medical/Healthcare AI Dankai Liao, Tianyi Zhang, Yufeng Wu, Xinyue Zhang, Qiaochu Xue, Zeyu Liu, Dachun Zhao, Linghan Cai, Yueming Jin 2026-07-21

TL;DR - PathAgentBench evaluates whether vision-language models can actively find and integrate diagnostic evidence in gigapixel pathology slides. Results expose a major gap between reasoning over curated regions and autonomously locating relevant tissue.

  • Covers 1,822 TCGA slides and 17,135 diagnostic paths annotated by ten board-certified pathologists.
  • Tests cross-modal matching, diagnostic-region localization, and multi-scale reasoning across 20 models.
  • Leading open-weight models exceed 93% multi-scale reasoning accuracy, but the best localization mean IoU remains below 0.09.
  • Autonomous evidence discovery deteriorates sharply at higher magnifications, reaching a 0.020 unconditional hit rate.

view merged work →