Multimodal Pretraining for Generalizable EEG Representation Learning
Merged summary
TL;DR - A multimodal EEG foundation model jointly learns from raw signals, time-frequency views, and text to improve generalizable seizure detection. It achieves state-of-the-art CHB-MIT performance while exposing substantial challenges in patient-independent evaluation.
- Combines Mamba, ViT-style, and lightweight text encoders in a shared embedding space.
- Uses masked modeling, cross-view contrastive alignment, and temporal consistency losses without labeled pretraining data.
- Achieves 0.874 AUROC as a single model and 0.878 as an ensemble on the standard CHB-MIT split.
- Records 0.558 mean balanced accuracy across 19 subjects under leave-one-subject-out evaluation and supports interpretable seizure localization.
Sources (1)
Multimodal Pretraining for Generalizable EEG Representation Learning
TL;DR - A multimodal EEG foundation model jointly learns from raw signals, time-frequency views, and text to improve generalizable seizure detection. It achieves state-of-the-art CHB-MIT performance while exposing substantial challenges in patient-independent evaluation.
- Combines Mamba, ViT-style, and lightweight text encoders in a shared embedding space.
- Uses masked modeling, cross-view contrastive alignment, and temporal consistency losses without labeled pretraining data.
- Achieves 0.874 AUROC as a single model and 0.878 as an ensemble on the standard CHB-MIT split.
- Records 0.558 mean balanced accuracy across 19 subjects under leave-one-subject-out evaluation and supports interpretable seizure localization.