🛰️ Daily AI Frontier
‹ back to 2026-07-29

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents

Research Medical/Healthcare AI

Ranking

Overall 80
Content 95
Popularity 45

Observed public metrics from 1 member.

Representative image for PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents

Merged summary

TL;DR - PatientAgentBench evaluates patient-facing healthcare agents through sustained, tool-using conversations grounded in realistic health records. Testing 10 models across 1,200 scenarios reveals that even frontier systems retain clinically important safety and workflow gaps.

  • Uses an LLM jury with over 100 clinician-grounded criteria across six dimensions.
  • Jury ratings achieved 79–93% adjacent agreement with licensed clinicians.
  • Triage was most discriminating, with pass rates ranging from 32% to 88%.
  • Frontier models failed 1–3% of safety and workflow cases, including unverified tool outputs and omitted crisis resources.

Sources (1)

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents

arXiv cs.AI Korosh Vatanparvar, Ashutosh Joshi, Maria Xenochristou, Mohammad Abuzar Hashemi, Prasad Kasu, Deepak Bansal, Daniel Lopez-Martinez, Anchal Nema, Ramya Ganesan, Will Kimbrough, Alex Woody, Yadunandana Rao, Dilek Hakkani-Tur, Wilko Schulz-Mahlendorf 2026-07-28 arXiv:2607.25485
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-08-17 09:52:52.947594 UTC

TL;DR - PatientAgentBench evaluates patient-facing healthcare agents through sustained, tool-using conversations grounded in realistic health records. Testing 10 models across 1,200 scenarios reveals that even frontier systems retain clinically important safety and workflow gaps.

  • Uses an LLM jury with over 100 clinician-grounded criteria across six dimensions.
  • Jury ratings achieved 79–93% adjacent agreement with licensed clinicians.
  • Triage was most discriminating, with pass rates ranging from 32% to 88%.
  • Frontier models failed 1–3% of safety and workflow cases, including unverified tool outputs and omitted crisis resources.
item →