Stress-Testing EEG Foundation Models for Clinical Decoding: Dataset Identity and Targeted Negative Controls
Merged summary
TL;DR - A benchmark of six EEG foundation models finds that clinical decoding performance is highly sensitive to dataset identity, evaluation splits, baselines, and negative controls. Pretraining showed a clear benefit mainly for cross-subject seizure detection.
- Classical EEG features substantially outperformed frozen REVE embeddings on Korean dementia classification.
- Frozen embeddings identified datasets almost perfectly but weakly decoded Korean diagnoses, indicating strong dataset-specific signals.
- Random initialization, random projections, and PCA sometimes matched or exceeded pretrained representations.
- On CHB-MIT ictal detection, REVE achieved 0.793 AUROC, beating a randomly initialized encoder by 9.2 percentage points.
Sources (1)
Stress-Testing EEG Foundation Models for Clinical Decoding: Dataset Identity and Targeted Negative Controls
TL;DR - A benchmark of six EEG foundation models finds that clinical decoding performance is highly sensitive to dataset identity, evaluation splits, baselines, and negative controls. Pretraining showed a clear benefit mainly for cross-subject seizure detection.
- Classical EEG features substantially outperformed frozen REVE embeddings on Korean dementia classification.
- Frozen embeddings identified datasets almost perfectly but weakly decoded Korean diagnoses, indicating strong dataset-specific signals.
- Random initialization, random projections, and PCA sometimes matched or exceeded pretrained representations.
- On CHB-MIT ictal detection, REVE achieved 0.793 AUROC, beating a randomly initialized encoder by 9.2 percentage points.