A report-grounded vision-language foundation model for colonoscopy from 280000 routine reports
Ranking
Overall
81
Content
100
Popularity
37
Observed public metrics from 1 member.
Merged summary
TL;DR - EndoCLIP is a colonoscopy vision-language foundation model trained on 125,756 lesion-level image-text pairs recovered from 280,476 routine reports. It shows that weakly linked clinical documentation can provide scalable supervision for retrieval, report generation, and diagnosis.
- Outperforms general-purpose and biomedical vision-language encoders across retrieval, structured report generation, and six multicentre classification tasks.
- Delivers stronger zero-shot and linear-probe performance.
- Its linear probe approaches expert-reader performance for benign-versus-malignant classification in a blinded study with 12 endoscopists.
- Enables clinical targets to be expressed through language rather than task-specific annotations.
Sources (1)
A report-grounded vision-language foundation model for colonoscopy from 280000 routine reports
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - EndoCLIP is a colonoscopy vision-language foundation model trained on 125,756 lesion-level image-text pairs recovered from 280,476 routine reports. It shows that weakly linked clinical documentation can provide scalable supervision for retrieval, report generation, and diagnosis.
- Outperforms general-purpose and biomedical vision-language encoders across retrieval, structured report generation, and six multicentre classification tasks.
- Delivers stronger zero-shot and linear-probe performance.
- Its linear probe approaches expert-reader performance for benign-versus-malignant classification in a blinded study with 12 endoscopists.
- Enables clinical targets to be expressed through language rather than task-specific annotations.