Target leakage, not model class, explains reported accuracy in survey-based cardiovascular screening: a leakage-tiered audit of glass-box and tabular foundation models
TL;DR - An audit of cardiovascular screening models finds that commonly reported high accuracy primarily reflects target leakage rather than model class. After removing post-diagnostic features, transparent glass-box models matched more complex alternatives while enabling faster inference and auditable fairness and uncertainty fixes.
- Removing two post-diagnostic features reduced every model’s AUROC by 0.049–0.051 and compressed results into a 0.0045-wide band.
- An explainable boosting machine was non-inferior within a 0.005 margin and scored patients about 104Ă— faster than the strongest tabular foundation model.
- Editing the glass-box model’s shape functions reduced the sex disparity in detection rates to 0.010.
- Mondrian calibration repaired subgroup conformal-coverage gaps, while frozen models transferred from 2022 to 2023 within 0.002 AUROC.