A report-grounded vision-language foundation model for colonoscopy from 280000 routine reports
TL;DR - EndoCLIP is a colonoscopy vision-language foundation model trained on 125,756 lesion-level image-text pairs recovered from 280,476 routine reports. It shows that weakly linked clinical documentation can provide scalable supervision for retrieval, report generation, and diagnosis.
- Outperforms general-purpose and biomedical vision-language encoders across retrieval, structured report generation, and six multicentre classification tasks.
- Delivers stronger zero-shot and linear-probe performance.
- Its linear probe approaches expert-reader performance for benign-versus-malignant classification in a blinded study with 12 endoscopists.
- Enables clinical targets to be expressed through language rather than task-specific annotations.