RadSight: Towards Perceptually Reliable Multimodal Radiology Image Understanding
TL;DR - RadSight is a radiology-focused multimodal model designed to improve diagnostic reliability through stronger low-level visual perception. It outperforms existing medical MLLMs across a new large-scale benchmark covering 2D and 3D imaging tasks.
- Perception-Bench contains 1.13 million samples across six perception and clinical evaluation dimensions.
- Existing MLLMs struggle with basic lesion properties such as location, size, and density.
- RadSight uses dual 2D/3D encoders and a four-stage curriculum from vision-language alignment to diagnostic interpretation.
- Training uses an 8.37 million-sample perception-oriented corpus, producing especially strong gains in spatial grounding and clinical diagnosis.