🛰️ Daily AI Frontier
‹ back to 2026-09-21

Beyond Benchmark Scores: Auditing Medical Vision-Language Models for Chest X-Ray Tuberculosis Screening

arXiv cs.CV Medical/Healthcare AI Mushir Akhtar, M. Tanveer, Mohd. Arshad 2026-09-18

TL;DR - An audit of four vision-language models across 12,200 chest X-rays finds that tuberculosis-screening performance is highly sensitive to cohorts, prompts, control populations, prevalence, and thresholds. Strong benchmark results therefore do not establish reliable clinical portability.

  • No model led across every cohort and reliability criterion; prompt changes significantly altered AUROC in 21 of 48 controlled comparisons.
  • Replacing healthy controls with non-tuberculosis disease controls reduced AUROC by 0.075–0.306, with medical models performing especially poorly against pneumonia and lung tumors on VinDr-CXR.
  • Thresholds calibrated for 95% sensitivity on TBX11K retained that target in only 4 of 16 external evaluations.
  • A supervised model fell from 0.999 AUROC on TBX11K validation to 0.629 on two external cohorts, highlighting substantial distribution-shift risk.

view merged work →