MedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation Agents
Merged summary
TL;DR - MedDDC-Eval evaluates multi-turn medical consultation agents by separating history-taking quality from diagnosis generation. This enables more controlled comparisons and supports training agents to gather diagnostically useful evidence efficiently.
- A shared frozen diagnostic reader scores policy-elicited histories, reducing confounding from policy-specific diagnosis generation.
- Changing only the reader shifted diagnosis F1 by 2.2–19.0 points and reversed 18%–36% of policy rankings.
- The framework measures diagnostic usefulness, information acquisition, and trajectory efficiency across Record and Dialogue splits.
- GRPO post-training of Qwen3-32B improved total scores by 9.7 and 4.6 points on the two held-out splits.
Sources (1)
MedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation Agents
TL;DR - MedDDC-Eval evaluates multi-turn medical consultation agents by separating history-taking quality from diagnosis generation. This enables more controlled comparisons and supports training agents to gather diagnostically useful evidence efficiently.
- A shared frozen diagnostic reader scores policy-elicited histories, reducing confounding from policy-specific diagnosis generation.
- Changing only the reader shifted diagnosis F1 by 2.2–19.0 points and reversed 18%–36% of policy rankings.
- The framework measures diagnostic usefulness, information acquisition, and trajectory efficiency across Record and Dialogue splits.
- GRPO post-training of Qwen3-32B improved total scores by 9.7 and 4.6 points on the two held-out splits.