Self-supervision drives representational convergence in medical foundation models more than clinical supervision
TL;DR - A controlled study of medical foundation models finds that representational convergence is driven more by self-supervised pretraining objectives than by clinical supervision, model scale, or capability. The limited shared geometry still enables useful cross-encoder and cross-hospital classifier transfer.
- Compared 18 image and 7 text encoders across five imaging modalities, including 650,982 chest radiographs.
- Matched self-supervised encoders aligned most (40.4%), versus label-supervised (21.1%) and image-text models (3.3%).
- Convergence did not significantly increase with model size and neither extended to clinical language nor matched radiologists’ similarity judgments.
- Linear classifiers transferred across encoders and five held-out hospitals, retaining roughly 85% of within-encoder performance.