Target leakage, not model class, explains reported accuracy in survey-based cardiovascular screening: a leakage-tiered audit of glass-box and tabular foundation models
Ranking
Overall
78
Content
95
Popularity
37
Observed public metrics from 1 member.
Merged summary
TL;DR - An audit of cardiovascular screening models finds that commonly reported high accuracy primarily reflects target leakage rather than model class. After removing post-diagnostic features, transparent glass-box models matched more complex alternatives while enabling faster inference and auditable fairness and uncertainty fixes.
- Removing two post-diagnostic features reduced every model’s AUROC by 0.049–0.051 and compressed results into a 0.0045-wide band.
- An explainable boosting machine was non-inferior within a 0.005 margin and scored patients about 104Ă— faster than the strongest tabular foundation model.
- Editing the glass-box model’s shape functions reduced the sex disparity in detection rates to 0.010.
- Mondrian calibration repaired subgroup conformal-coverage gaps, while frozen models transferred from 2022 to 2023 within 0.002 AUROC.
Sources (1)
Target leakage, not model class, explains reported accuracy in survey-based cardiovascular screening: a leakage-tiered audit of glass-box and tabular foundation models
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - An audit of cardiovascular screening models finds that commonly reported high accuracy primarily reflects target leakage rather than model class. After removing post-diagnostic features, transparent glass-box models matched more complex alternatives while enabling faster inference and auditable fairness and uncertainty fixes.
- Removing two post-diagnostic features reduced every model’s AUROC by 0.049–0.051 and compressed results into a 0.0045-wide band.
- An explainable boosting machine was non-inferior within a 0.005 margin and scored patients about 104Ă— faster than the strongest tabular foundation model.
- Editing the glass-box model’s shape functions reduced the sex disparity in detection rates to 0.010.
- Mondrian calibration repaired subgroup conformal-coverage gaps, while frozen models transferred from 2022 to 2023 within 0.002 AUROC.