MVC-Bench: Benchmarking Calibration of Medical Vision-Language Models
TL;DR - MVC-Bench evaluates confidence calibration in medical vision-language models across clinical imaging modalities, distribution shifts, and prompt variations. It also introduces Multi-Class Margin regularization, which substantially improves calibration across the tested settings.
- Covers eight model backbones and three modalities: fundus imaging, histopathology, and chest X-rays.
- Compares post-hoc, train-time, zero-shot, and six prompt-tuning approaches across more than 1,638 controlled experiments.
- Measures accuracy alongside Expected, Maximum, and Adaptive Calibration Error, including under domain shifts and prompt or seed variations.
- Multi-Class Margin regularization achieves the lowest Expected Calibration Error in 10 of 12 in-domain settings and remains competitive under domain shift.