Beyond Benchmark Scores: Auditing Medical Vision-Language Models for Chest X-Ray Tuberculosis Screening
Ranking
Overall
82
Content
95
Popularity
N/A
No observed public metrics; popularity remains neutral/archived.
Merged summary
TL;DR - An audit of four vision-language models across 12,200 chest X-rays finds that tuberculosis-screening performance is highly sensitive to cohorts, prompts, control populations, prevalence, and thresholds. Strong benchmark results therefore do not establish reliable clinical portability.
- No model led across every cohort and reliability criterion; prompt changes significantly altered AUROC in 21 of 48 controlled comparisons.
- Replacing healthy controls with non-tuberculosis disease controls reduced AUROC by 0.075–0.306, with medical models performing especially poorly against pneumonia and lung tumors on VinDr-CXR.
- Thresholds calibrated for 95% sensitivity on TBX11K retained that target in only 4 of 16 external evaluations.
- A supervised model fell from 0.999 AUROC on TBX11K validation to 0.629 on two external cohorts, highlighting substantial distribution-shift risk.
Sources (1)
Beyond Benchmark Scores: Auditing Medical Vision-Language Models for Chest X-Ray Tuberculosis Screening
Public signals
N/A
TL;DR - An audit of four vision-language models across 12,200 chest X-rays finds that tuberculosis-screening performance is highly sensitive to cohorts, prompts, control populations, prevalence, and thresholds. Strong benchmark results therefore do not establish reliable clinical portability.
- No model led across every cohort and reliability criterion; prompt changes significantly altered AUROC in 21 of 48 controlled comparisons.
- Replacing healthy controls with non-tuberculosis disease controls reduced AUROC by 0.075–0.306, with medical models performing especially poorly against pneumonia and lung tumors on VinDr-CXR.
- Thresholds calibrated for 95% sensitivity on TBX11K retained that target in only 4 of 16 external evaluations.
- A supervised model fell from 0.999 AUROC on TBX11K validation to 0.629 on two external cohorts, highlighting substantial distribution-shift risk.