🛰️ Daily AI Frontier
‹ back to 2026-07-29

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents

arXiv cs.AI Medical/Healthcare AI Korosh Vatanparvar, Ashutosh Joshi, Maria Xenochristou, Mohammad Abuzar Hashemi, Prasad Kasu, Deepak Bansal, Daniel Lopez-Martinez, Anchal Nema, Ramya Ganesan, Will Kimbrough, Alex Woody, Yadunandana Rao, Dilek Hakkani-Tur, Wilko Schulz-Mahlendorf 2026-07-28
Representative image for PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents

TL;DR - PatientAgentBench evaluates patient-facing healthcare agents through sustained, tool-using conversations grounded in realistic health records. Testing 10 models across 1,200 scenarios reveals that even frontier systems retain clinically important safety and workflow gaps.

  • Uses an LLM jury with over 100 clinician-grounded criteria across six dimensions.
  • Jury ratings achieved 79–93% adjacent agreement with licensed clinicians.
  • Triage was most discriminating, with pass rates ranging from 32% to 88%.
  • Frontier models failed 1–3% of safety and workflow cases, including unverified tool outputs and omitted crisis resources.

view merged work →