MedAgent-R1: Faithfulness-Aware Reinforcement Learning for Evidence-Grounded Medical Reasoning
TL;DR - MedAgent-R1 uses faithfulness-aware reinforcement learning to make medical retrieval agents ground their reasoning in cited evidence rather than fabricate plausible justifications. It sharply reduces citation fabrication while preserving accuracy and improving safety performance.
- Outcome-only RL raised accuracy by 5 points but increased citation fabrication from 16.5% to 31.8%, a failure mode termed “confident hallucination.”
- Its reward design gates accuracy credit on evidence grounding and adds retrieval-validity and conciseness signals to prevent reward exploitation.
- MedAgent-R1 reduced citation fabrication to 4.7%, increased evidence completeness from 58.7 to 82.6, and maintained 75.1% accuracy.
- It gained 13.2 points on HealthBench Safety and surpassed GPT-4o on reported faithfulness measures, though not on overall accuracy.