🛰️ Daily AI Frontier
‹ back to 2026-09-21

Beyond Benchmark Scores: Auditing Medical Vision-Language Models for Chest X-Ray Tuberculosis Screening

Research Medical/Healthcare AI

Ranking

Overall 82
Content 95
Popularity N/A

No observed public metrics; popularity remains neutral/archived.

Merged summary

TL;DR - An audit of four vision-language models across 12,200 chest X-rays finds that tuberculosis-screening performance is highly sensitive to cohorts, prompts, control populations, prevalence, and thresholds. Strong benchmark results therefore do not establish reliable clinical portability.

  • No model led across every cohort and reliability criterion; prompt changes significantly altered AUROC in 21 of 48 controlled comparisons.
  • Replacing healthy controls with non-tuberculosis disease controls reduced AUROC by 0.075–0.306, with medical models performing especially poorly against pneumonia and lung tumors on VinDr-CXR.
  • Thresholds calibrated for 95% sensitivity on TBX11K retained that target in only 4 of 16 external evaluations.
  • A supervised model fell from 0.999 AUROC on TBX11K validation to 0.629 on two external cohorts, highlighting substantial distribution-shift risk.

Sources (1)

Beyond Benchmark Scores: Auditing Medical Vision-Language Models for Chest X-Ray Tuberculosis Screening

arXiv cs.CV Mushir Akhtar, M. Tanveer, Mohd. Arshad 2026-09-18 arXiv:2609.21763
Public signals N/A
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-26 14:14:54.395451 UTC

TL;DR - An audit of four vision-language models across 12,200 chest X-rays finds that tuberculosis-screening performance is highly sensitive to cohorts, prompts, control populations, prevalence, and thresholds. Strong benchmark results therefore do not establish reliable clinical portability.

  • No model led across every cohort and reliability criterion; prompt changes significantly altered AUROC in 21 of 48 controlled comparisons.
  • Replacing healthy controls with non-tuberculosis disease controls reduced AUROC by 0.075–0.306, with medical models performing especially poorly against pneumonia and lung tumors on VinDr-CXR.
  • Thresholds calibrated for 95% sensitivity on TBX11K retained that target in only 4 of 16 external evaluations.
  • A supervised model fell from 0.999 AUROC on TBX11K validation to 0.629 on two external cohorts, highlighting substantial distribution-shift risk.
item →