🛰️ Daily AI Frontier
‹ back to 2026-09-11

Target leakage, not model class, explains reported accuracy in survey-based cardiovascular screening: a leakage-tiered audit of glass-box and tabular foundation models

Research Medical/Healthcare AI

Ranking

Overall 78
Content 95
Popularity 37

Observed public metrics from 1 member.

Merged summary

TL;DR - An audit of cardiovascular screening models finds that commonly reported high accuracy primarily reflects target leakage rather than model class. After removing post-diagnostic features, transparent glass-box models matched more complex alternatives while enabling faster inference and auditable fairness and uncertainty fixes.

  • Removing two post-diagnostic features reduced every model’s AUROC by 0.049–0.051 and compressed results into a 0.0045-wide band.
  • An explainable boosting machine was non-inferior within a 0.005 margin and scored patients about 104Ă— faster than the strongest tabular foundation model.
  • Editing the glass-box model’s shape functions reduced the sex disparity in detection rates to 0.010.
  • Mondrian calibration repaired subgroup conformal-coverage gaps, while frozen models transferred from 2022 to 2023 within 0.002 AUROC.

Sources (1)

Target leakage, not model class, explains reported accuracy in survey-based cardiovascular screening: a leakage-tiered audit of glass-box and tabular foundation models

arXiv cs.CL Raad Bin Tareaf, Murad Al-Rajab, Samia Loucif, Samer Ellaham, Cedric Schmitz 2026-09-10 arXiv:2609.11838
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-19 14:14:56.664125 UTC

TL;DR - An audit of cardiovascular screening models finds that commonly reported high accuracy primarily reflects target leakage rather than model class. After removing post-diagnostic features, transparent glass-box models matched more complex alternatives while enabling faster inference and auditable fairness and uncertainty fixes.

  • Removing two post-diagnostic features reduced every model’s AUROC by 0.049–0.051 and compressed results into a 0.0045-wide band.
  • An explainable boosting machine was non-inferior within a 0.005 margin and scored patients about 104Ă— faster than the strongest tabular foundation model.
  • Editing the glass-box model’s shape functions reduced the sex disparity in detection rates to 0.010.
  • Mondrian calibration repaired subgroup conformal-coverage gaps, while frozen models transferred from 2022 to 2023 within 0.002 AUROC.
item →