🛰️ Daily AI Frontier
‹ back to 2026-07-30

A report-grounded vision-language foundation model for colonoscopy from 280000 routine reports

Research Medical/Healthcare AI

Ranking

Overall 81
Content 100
Popularity 37

Observed public metrics from 1 member.

Representative image for A report-grounded vision-language foundation model for colonoscopy from 280000 routine reports

Merged summary

TL;DR - EndoCLIP is a colonoscopy vision-language foundation model trained on 125,756 lesion-level image-text pairs recovered from 280,476 routine reports. It shows that weakly linked clinical documentation can provide scalable supervision for retrieval, report generation, and diagnosis.

  • Outperforms general-purpose and biomedical vision-language encoders across retrieval, structured report generation, and six multicentre classification tasks.
  • Delivers stronger zero-shot and linear-probe performance.
  • Its linear probe approaches expert-reader performance for benign-versus-malignant classification in a blinded study with 12 endoscopists.
  • Enables clinical targets to be expressed through language rather than task-specific annotations.

Sources (1)

A report-grounded vision-language foundation model for colonoscopy from 280000 routine reports

arXiv cs.AI Jia Yu, Yan Zhu, Yili He, Zilong Wang, Xinyang Jiang, Peiyao Fu, Ruijie Yang, Tianyi Chen, Siyuan Li, Zhihua Wang, Fei Wu, Quanlin Li, Xian Yang, Pinghong Zhou, Shuo Wang 2026-07-30 arXiv:2607.28466
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-08-28 14:31:42.411718 UTC

TL;DR - EndoCLIP is a colonoscopy vision-language foundation model trained on 125,756 lesion-level image-text pairs recovered from 280,476 routine reports. It shows that weakly linked clinical documentation can provide scalable supervision for retrieval, report generation, and diagnosis.

  • Outperforms general-purpose and biomedical vision-language encoders across retrieval, structured report generation, and six multicentre classification tasks.
  • Delivers stronger zero-shot and linear-probe performance.
  • Its linear probe approaches expert-reader performance for benign-versus-malignant classification in a blinded study with 12 endoscopists.
  • Enables clinical targets to be expressed through language rather than task-specific annotations.
item →