UltraSAM3: A Concept-Driven Foundation Model for Universal Ultrasound Image Segmentation
TL;DR - UltraSAM3 adapts SAM3 into a concept-driven foundation model for universal ultrasound segmentation, letting clinicians specify targets by text instead of expert-drawn visual prompts. It matters because ultrasound's speckle noise, low contrast and ambiguous boundaries have kept segmentation stuck in task-specific or prompt-dependent models.
- Trained on image–mask–concept triplets from a large-scale corpus spanning 37 public ultrasound datasets and 13 anatomical categories, aligning ultrasound visual patterns with clinically meaningful concepts across organs and lesions.
- Replaces visual-prompt dependence with text-based target specification, addressing the usability gap in existing foundation models like SAM-style segmenters.
- Adds an instruction-guided agent that parses complex natural-language queries into concise ultrasound concept prompts, reported to improve robustness on complex user instructions.
- Reported to outperform representative concept- and text-driven biomedical segmentation baselines on multi-organ benchmarks, external datasets, and visual-prompt-enhanced settings (no numeric metrics given in the abstract).