ECG Mirage: Revealing and Mitigating the Underutilisation of ECGs in Vision-Language Models for Clinical Prediction
Ranking
Overall
78
Content
95
Popularity
39
Observed public metrics from 1 member.
Merged summary
TL;DR - This paper identifies “ECG Mirage,” where vision-language models appear effective at clinical prediction but fail to meaningfully use the correct patient’s ECG. Visual prompt tuning improves both predictive performance and reliance on patient-specific ECG information without modifying the VLM backbone.
- Tests compare matched ECGs, outcome-discordant mismatched ECGs, and text-only inputs while keeping clinical context and targets fixed.
- Across four VLMs on MDS-ED, matched ECGs provide no consistent advantage for predicting ICU admission or clinical deterioration.
- Supervised learning followed by conditional direct preference optimization of restricted visual prompts yields 70.6% balanced accuracy for ICU admission and 67.5% for deterioration.
- The tuned models increase matched-versus-mismatched performance gaps to roughly 16.5 and 5.5 percentage points for the two tasks, respectively.
Sources (1)
ECG Mirage: Revealing and Mitigating the Underutilisation of ECGs in Vision-Language Models for Clinical Prediction
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - This paper identifies “ECG Mirage,” where vision-language models appear effective at clinical prediction but fail to meaningfully use the correct patient’s ECG. Visual prompt tuning improves both predictive performance and reliance on patient-specific ECG information without modifying the VLM backbone.
- Tests compare matched ECGs, outcome-discordant mismatched ECGs, and text-only inputs while keeping clinical context and targets fixed.
- Across four VLMs on MDS-ED, matched ECGs provide no consistent advantage for predicting ICU admission or clinical deterioration.
- Supervised learning followed by conditional direct preference optimization of restricted visual prompts yields 70.6% balanced accuracy for ICU admission and 67.5% for deterioration.
- The tuned models increase matched-versus-mismatched performance gaps to roughly 16.5 and 5.5 percentage points for the two tasks, respectively.