PathView-Bench: Can Multimodal Large Language Models Achieve Fine-grained Multiscale Understanding of Pathology Images?
TL;DR - PathVU is a large, vision-anchored benchmark testing whether multimodal models understand pathology images across local regions and whole-slide views. Results across 18 models reveal substantial limitations in fine-grained, multiscale visual reasoning.
- Includes 14 VQA tasks, 61,673 images, and 308,070 samples from 23 public datasets.
- Covers 28 organs using over 7.25 million human-supervised labels and spatial annotations.
- Tests localization, recognition, quantity estimation, spatial reasoning, and insufficient-context judgment.
- Uses deterministic targets for reproducible, programmatic scoring rather than evaluating only final diagnoses or reports.