How Good are Foundation Models in Longitudinal MRI Disease Progression Reasoning?
TL;DR - This paper introduces a benchmark for evaluating vision-language models on longitudinal, multi-view MRI disease progression. Tests show persistent weaknesses in identifying change direction and quantifying volume, highlighting barriers to clinical deployment.
- Includes 3,920 expert-verified question-answer pairs from 890 patients and over 3,200 MRI timepoints.
- Covers seven cohorts spanning glioblastoma, neurodegeneration, vestibular schwannoma, and brain metastases.
- Evaluation of 16 models found moderate temporal alignment but systematic progression-reasoning failures.
- Multi-view input improved spatial localization but degraded temporal reasoning in compact models.