🛰️ Daily AI Frontier
‹ back to 2026-09-25

Synthetic Hospital: An Open, Verifiable, Physician-Validated Longitudinal EHR Benchmark

Research Medical/Healthcare AI

Ranking

Overall 85
Content 100
Popularity N/A

No observed public metrics; popularity remains neutral/archived.

Representative image for Synthetic Hospital: An Open, Verifiable, Physician-Validated Longitudinal EHR Benchmark

Merged summary

TL;DR - Synthetic Hospital is an open, fully synthetic benchmark for evaluating language models on realistic longitudinal electronic health records without exposing protected health information. Its verifiable ground truth and simulated EHR infrastructure enable reproducible testing of clinical AI systems.

  • Includes 1,268 synthetic patients and 5,602 encounters grounded in public medical-education sources and standard clinical ontologies.
  • Provides complete provenance plus interoperability APIs, role-based access, and a function-calling interface modeled on real hospital systems.
  • Physicians distinguished synthetic from real charts at only 53% accuracy in a blinded review.
  • The best of 10 evaluated models achieved 0.73 severity-weighted F1 on longitudinal problem-list reconstruction and missed roughly half of clinically relevant findings in chart summaries.

Sources (1)

Synthetic Hospital: An Open, Verifiable, Physician-Validated Longitudinal EHR Benchmark

arXiv cs.AI Christine Park, Valerie Chen, Tim Dettmers 2026-09-24 arXiv:2609.30027
Public signals N/A
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-26 14:14:08.856518 UTC

TL;DR - Synthetic Hospital is an open, fully synthetic benchmark for evaluating language models on realistic longitudinal electronic health records without exposing protected health information. Its verifiable ground truth and simulated EHR infrastructure enable reproducible testing of clinical AI systems.

  • Includes 1,268 synthetic patients and 5,602 encounters grounded in public medical-education sources and standard clinical ontologies.
  • Provides complete provenance plus interoperability APIs, role-based access, and a function-calling interface modeled on real hospital systems.
  • Physicians distinguished synthetic from real charts at only 53% accuracy in a blinded review.
  • The best of 10 evaluated models achieved 0.73 severity-weighted F1 on longitudinal problem-list reconstruction and missed roughly half of clinically relevant findings in chart summaries.
item →