🛰️ Daily AI Frontier
‹ back to 2026-07-30

A report-grounded vision-language foundation model for colonoscopy from 280000 routine reports

arXiv cs.AI Medical/Healthcare AI Jia Yu, Yan Zhu, Yili He, Zilong Wang, Xinyang Jiang, Peiyao Fu, Ruijie Yang, Tianyi Chen, Siyuan Li, Zhihua Wang, Fei Wu, Quanlin Li, Xian Yang, Pinghong Zhou, Shuo Wang 2026-07-30
Representative image for A report-grounded vision-language foundation model for colonoscopy from 280000 routine reports

TL;DR - EndoCLIP is a colonoscopy vision-language foundation model trained on 125,756 lesion-level image-text pairs recovered from 280,476 routine reports. It shows that weakly linked clinical documentation can provide scalable supervision for retrieval, report generation, and diagnosis.

  • Outperforms general-purpose and biomedical vision-language encoders across retrieval, structured report generation, and six multicentre classification tasks.
  • Delivers stronger zero-shot and linear-probe performance.
  • Its linear probe approaches expert-reader performance for benign-versus-malignant classification in a blinded study with 12 endoscopists.
  • Enables clinical targets to be expressed through language rather than task-specific annotations.

view merged work →