🛰️ Daily AI Frontier
‹ back to 2026-09-04

MetaStructAtlas: A Grounded 3D Vision-Language Dataset and Benchmark for Functional and Structural Reasoning in Whole-Body PET/CT

arXiv cs.CV Medical/Healthcare AI Chenguang Zheng, Le Xue, Yichi Zhang, Wenbo Zhang, Zehui Ling, Gang Feng, Xin Gao, Yuan Qi, Yuan Cheng, Zixin Hu, Mei Tian 2026-09-03
Representative image for MetaStructAtlas: A Grounded 3D Vision-Language Dataset and Benchmark for Functional and Structural Reasoning in Whole-Body PET/CT

TL;DR - MetaStructAtlas is a large-scale, grounded 3D vision-language dataset and benchmark for jointly reasoning over metabolic and anatomical evidence in whole-body PET/CT. It addresses a gap in medical VLM research, which has largely focused on regional CT imaging.

  • Includes 490 co-registered 3D PET/CT volumes, 50,470 organ-level segmentation masks, and grounded radiology reports.
  • Introduces MetaStructVQA, a benchmark with 100,565 grounded visual question-answer pairs.
  • Covers anatomical, morphological, and metabolic reasoning while linking diagnostic questions to visual evidence across modalities.
  • Evaluates state-of-the-art 3D medical vision-language models to establish a baseline for whole-body nuclear-medicine reasoning.

view merged work →