🛰️ Daily AI Frontier
‹ back to 2026-07-22

MedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation Agents

Research Medical/Healthcare AI

Merged summary

TL;DR - MedDDC-Eval evaluates multi-turn medical consultation agents by separating history-taking quality from diagnosis generation. This enables more controlled comparisons and supports training agents to gather diagnostically useful evidence efficiently.

  • A shared frozen diagnostic reader scores policy-elicited histories, reducing confounding from policy-specific diagnosis generation.
  • Changing only the reader shifted diagnosis F1 by 2.2–19.0 points and reversed 18%–36% of policy rankings.
  • The framework measures diagnostic usefulness, information acquisition, and trajectory efficiency across Record and Dialogue splits.
  • GRPO post-training of Qwen3-32B improved total scores by 9.7 and 4.6 points on the two held-out splits.

Sources (1)

MedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation Agents

arXiv cs.CL Guofeng Zhang, Yizeng Quan, Huaiyi Fang, Jianwei Lv, Jinyao Liu, Xunxu Duan, Lening An, Yu Ouyang, Junfeng Wang 2026-07-21 arXiv:2607.18999

TL;DR - MedDDC-Eval evaluates multi-turn medical consultation agents by separating history-taking quality from diagnosis generation. This enables more controlled comparisons and supports training agents to gather diagnostically useful evidence efficiently.

  • A shared frozen diagnostic reader scores policy-elicited histories, reducing confounding from policy-specific diagnosis generation.
  • Changing only the reader shifted diagnosis F1 by 2.2–19.0 points and reversed 18%–36% of policy rankings.
  • The framework measures diagnostic usefulness, information acquisition, and trajectory efficiency across Record and Dialogue splits.
  • GRPO post-training of Qwen3-32B improved total scores by 9.7 and 4.6 points on the two held-out splits.
item →