🛰️ Daily AI Frontier
‹ back to 2026-08-04

HarMoE: Multi-Source Chest Radiograph Pretraining with Dataset-Disentangled Experts

Research Medical/Healthcare AI

Ranking

Overall 67
Content 80
Popularity 36

Observed public metrics from 1 member.

Representative image for HarMoE: Multi-Source Chest Radiograph Pretraining with Dataset-Disentangled Experts

Merged summary

TL;DR - HarMoE is a chest X-ray vision-language pretraining framework that harmonizes many heterogeneous multi-label classification datasets instead of relying mainly on MIMIC-CXR image-report pairs, using dataset-aware mixture-of-experts to keep clinical semantics separate from dataset identity. It matters because it shows scaling radiology VLMs can come from cleaner, broader labeled supervision rather than more free-text reports.

  • Core problem: differences in label ontologies, annotation protocols, acquisition pipelines, and report styles cause models to entangle clinical semantics with dataset identity, hurting transfer even as data scale grows.
  • Method: a shared backbone learns cross-dataset medical semantics while source-specific variation is confined to lightweight residual experts placed in deeper decoder layers.
  • Supervision: training uses a unified disease vocabulary with masked multi-dataset supervision, so complementary annotations across sources can be combined without creating false negatives.
  • Reported gains over strong baselines on zero-shot classification, out-of-distribution transfer, and grounding; code plus an 873k-image harmonized dataset are slated for release at github.com/Roypic/harmoe.

Sources (1)

HarMoE: Multi-Source Chest Radiograph Pretraining with Dataset-Disentangled Experts

arXiv cs.CV Haozhe Luo, Ziyu Zhou, Shelley Zixin Shu, Mauricio Reyes 2026-08-03 arXiv:2608.02252
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-08-17 09:48:44.832919 UTC

TL;DR - HarMoE is a chest X-ray vision-language pretraining framework that harmonizes many heterogeneous multi-label classification datasets instead of relying mainly on MIMIC-CXR image-report pairs, using dataset-aware mixture-of-experts to keep clinical semantics separate from dataset identity. It matters because it shows scaling radiology VLMs can come from cleaner, broader labeled supervision rather than more free-text reports.

  • Core problem: differences in label ontologies, annotation protocols, acquisition pipelines, and report styles cause models to entangle clinical semantics with dataset identity, hurting transfer even as data scale grows.
  • Method: a shared backbone learns cross-dataset medical semantics while source-specific variation is confined to lightweight residual experts placed in deeper decoder layers.
  • Supervision: training uses a unified disease vocabulary with masked multi-dataset supervision, so complementary annotations across sources can be combined without creating false negatives.
  • Reported gains over strong baselines on zero-shot classification, out-of-distribution transfer, and grounding; code plus an 873k-image harmonized dataset are slated for release at github.com/Roypic/harmoe.
item →