PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents
Ranking
Overall
80
Content
95
Popularity
45
Observed public metrics from 1 member.
Merged summary
TL;DR - PatientAgentBench evaluates patient-facing healthcare agents through sustained, tool-using conversations grounded in realistic health records. Testing 10 models across 1,200 scenarios reveals that even frontier systems retain clinically important safety and workflow gaps.
- Uses an LLM jury with over 100 clinician-grounded criteria across six dimensions.
- Jury ratings achieved 79–93% adjacent agreement with licensed clinicians.
- Triage was most discriminating, with pass rates ranging from 32% to 88%.
- Frontier models failed 1–3% of safety and workflow cases, including unverified tool outputs and omitted crisis resources.
Sources (1)
PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - PatientAgentBench evaluates patient-facing healthcare agents through sustained, tool-using conversations grounded in realistic health records. Testing 10 models across 1,200 scenarios reveals that even frontier systems retain clinically important safety and workflow gaps.
- Uses an LLM jury with over 100 clinician-grounded criteria across six dimensions.
- Jury ratings achieved 79–93% adjacent agreement with licensed clinicians.
- Triage was most discriminating, with pass rates ranging from 32% to 88%.
- Frontier models failed 1–3% of safety and workflow cases, including unverified tool outputs and omitted crisis resources.