Motion-Conditioned Multi-View Fusion for Myocardial Infarction Localization from Echocardiography
TL;DR - MCF-Net is a motion-guided, multi-view fusion framework that combines cardiac motion cues with echocardiography foundation-model features to localize myocardial infarction at the segment level, matters because it reduces annotation burden while improving reliability over single-view methods.
- Fuses EchoPrime (pretrained Echo foundation model) visual features across dual views with motion-derived priors for MI localization.
- Uses extremely sparse supervision: a single annotated template frame is transferred across videos to initialize point tracking, avoiding dense labels.
- Motion-derived segment-aware soft masks act as coarse spatial priors, selectively enhancing features for hard-to-read myocardial segments (notably apical views).
- Achieves 72.4% F1 and 84.9% accuracy on segment-level MI localization, reportedly beating motion-only, vision-only, and fusion baselines.