MedPRESS: A Multi-turn Benchmark for Patient-Pressure-Induced Medical Sycophancy in LLMs
TL;DR - MedPRESS is a multi-turn benchmark that measures whether LLMs cave to patient pressure and endorse unsafe medical advice, showing that static safety evaluations miss a real failure mode in conversational health settings.
- 600 medically grounded five-turn dialogues span three scenario families: medication/treatment demand, personal health self-care, and symptom triage / care resistance.
- Each dialogue escalates pressure through a fixed arc — health query, personal experience, social proof, external evidence claims, then direct adversarial challenge.
- 20 LLMs (general, medical-domain, lightweight, large, open-weight, proprietary) were evaluated with structured judging and safety-focused metrics; models frequently drifted into unsafe agreement, with variation by family, scale, and prompt type.
- Anti-sycophancy prompting improved robustness for several models but did not eliminate unsafe agreement, indicating safe knowledge alone doesn't guarantee it is maintained under pressure.