🛰️ Daily AI Frontier
‹ back to 2026-09-09

Performance of Clinical AI System and Physicians and Frontier Language Models in primary care diagnostics

arXiv cs.CL Medical/Healthcare AI Andy Nkansah, Hanna Plotnitskaya, Stanislau Salavei, Anna Kozlova, Piotr Gibas, Julian Milek, Viktar Harbachou, Aleksey Ropan, Pavel Satalkin 2026-09-08

TL;DR - In 150 synthetic Polish-language primary-care consultations, the clinical AI system Doctorina outperformed eight physicians on diagnosis, workup, and initial treatment. The results suggest that adaptive information gathering may enable clinical AI to deliver stronger end-to-end diagnostic performance than standalone evaluation alone captures.

  • Doctorina achieved 82.0% Top-1 diagnostic concordance versus 57.0% for physicians, a 25-point difference (95% CI: 17.7–32.7).
  • Primary-or-reference-differential concordance reached 97.3% for Doctorina versus 85.0% for physicians.
  • Doctorina also led physicians in normalized workup scores (89.4 vs. 66.9) and treatment scores (83.7 vs. 61.2).
  • A second Doctorina run reproduced its advantages; Kimi K3 ranked next diagnostically, while Claude Opus 5 had the highest management estimate among closely matched leading models.

view merged work →