PathoArgus: Advancing Evidence-Grounded Long-Context Visual Reasoning across Gigapixel Whole-Slide and Multi-Slide Case Contexts
TL;DR - PathoArgus-Bench evaluates whether pathology AI systems genuinely ground answers in gigapixel whole-slide and multi-slide evidence rather than exploiting textual priors. Results show that conventional question-level accuracy substantially overstates reliable evidence-based reasoning.
- The benchmark contains 22,078 four-choice questions from 4,913 patients across 15 TCGA projects, spanning six pathology capabilities and three evidence-demand levels.
- Evidence State Quartets test grounding by holding question text fixed while moving, replacing, or removing the target slide set.
- GPT-5.6 achieved 57.09% overall accuracy but answered all four states correctly in only 19 of 483 quartets (3.93% QExact).
- The proposed fixed-budget PathoArgus reader reached 50.39% overall accuracy but only 1.86% QExact, indicating that better context access alone does not ensure consistent evidence use.