Big, Bright, or Invisible: A Frozen-Feature Benchmark of 3D CT Foundation Models
Ranking
Observed public metrics from 1 member.
Merged summary
TL;DR - A benchmark of ten frozen 3D CT foundation-model encoders across three thoracic CT cohorts (including an unseen internal clinical set) using k-NN, zero-shot prompting, and linear probing. It matters because it shows no universal state-of-the-art exists, and that physical detectability — not architecture — is the dominant limit on diagnostic breadth.
- Rankings fluctuate substantially with evaluation context; models pairing fine-grained image tokenization with vision-language alignment generally lead, but a lightweight supervised encoder stays competitive, indicating explicit labels can substitute for scale.
- The primary performance determinant is a physical bottleneck: a finding's detectability scales with its contrast against surrounding tissue and its spatial extent.
- Controlled within-organ comparisons show widespread/high-contrast abnormalities (devices, effusions) are reliably recovered, while small, low-contrast focal lesions fail across all encoders.
- Authors attribute this to globally pooled embeddings and argue region- or lesion-level pretraining is needed to represent small, low-contrast structures.
Sources (1)
Big, Bright, or Invisible: A Frozen-Feature Benchmark of 3D CT Foundation Models
TL;DR - A benchmark of ten frozen 3D CT foundation-model encoders across three thoracic CT cohorts (including an unseen internal clinical set) using k-NN, zero-shot prompting, and linear probing. It matters because it shows no universal state-of-the-art exists, and that physical detectability — not architecture — is the dominant limit on diagnostic breadth.
- Rankings fluctuate substantially with evaluation context; models pairing fine-grained image tokenization with vision-language alignment generally lead, but a lightweight supervised encoder stays competitive, indicating explicit labels can substitute for scale.
- The primary performance determinant is a physical bottleneck: a finding's detectability scales with its contrast against surrounding tissue and its spatial extent.
- Controlled within-organ comparisons show widespread/high-contrast abnormalities (devices, effusions) are reliably recovered, while small, low-contrast focal lesions fail across all encoders.
- Authors attribute this to globally pooled embeddings and argue region- or lesion-level pretraining is needed to represent small, low-contrast structures.