SCALPEL: Semantic Cross-modal Alignment via LLM-Powered Encoder Learning for Medical Vision-Language Representation
TL;DR - SCALPEL is a medical vision-language pretraining framework that adapts a generative medical LLM into a stronger text encoder while reducing training costs and clinically significant alignment errors. It reports state-of-the-art results across retrieval, zero-shot classification, and visual question answering benchmarks.
- Contrastive clinical-report fine-tuning produces more isotropic LLM text representations.
- Offline feature caching enables memory-efficient asymmetric image-text alignment.
- An anatomy-negation-aware objective penalizes laterality confusion and false negation mismatches.
- Evaluations span MIMIC-CXR, CheXpert, and IU X-Ray.