🛰️ Daily AI Frontier
‹ back to 2026-08-12

ConRub-Med: Reinforcement Learning with Consensus Rubrics for Open-Ended Medical Question Answering

Research Medical/Healthcare AI

Ranking

Overall 66
Content 80
Popularity 33

Observed public metrics from 1 member.

Representative image for ConRub-Med: Reinforcement Learning with Consensus Rubrics for Open-Ended Medical Question Answering

Merged summary

TL;DR - ConRub-Med is an RL recipe that replaces expensive physician-written rubrics with consensus-filtered, model-generated rubrics to supervise open-ended medical QA, where cheap outcome verifiers don't exist. It shows scalable rubric supervision can beat larger-sample baselines on hard clinical benchmarks.

  • Rubric construction: three heterogeneous LLMs independently propose atomic criteria, and a separate reviewer model keeps only criteria with semantic support from all three generators.
  • Three-State scoring separates correct coverage, missing information, and incorrect claims, with errors given negative rather than zero credit.
  • GRPO variant: when all responses in a group get identical rewards, a pairwise judge supplies sequence-level advantages only if both candidate orderings agree; untied groups use vanilla GRPO.
  • Results: ranks first on 6 of 9 benchmarks with the best medical and generalization averages; 38.98 ± 1.04 on HealthBench-Hard from 5,166 prompts vs. InfiMed-ORBIT's 33.60 (8K) and 37.30 (28K). Blinded ratings by two medical experts favored the full pipeline's rubric panels over single-generator panels.

Sources (1)

ConRub-Med: Reinforcement Learning with Consensus Rubrics for Open-Ended Medical Question Answering

arXiv cs.CL Taojie Zhu, Yuan Xia, Tao Sun, Yizhi Wang, Yan Chen, Qunshan He, Tian Guan, Jian Wang, Jinjie Gu, Junwei Liu, Yonghong He 2026-08-11 arXiv:2608.10996
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-08-13 10:09:30.951836 UTC

TL;DR - ConRub-Med is an RL recipe that replaces expensive physician-written rubrics with consensus-filtered, model-generated rubrics to supervise open-ended medical QA, where cheap outcome verifiers don't exist. It shows scalable rubric supervision can beat larger-sample baselines on hard clinical benchmarks.

  • Rubric construction: three heterogeneous LLMs independently propose atomic criteria, and a separate reviewer model keeps only criteria with semantic support from all three generators.
  • Three-State scoring separates correct coverage, missing information, and incorrect claims, with errors given negative rather than zero credit.
  • GRPO variant: when all responses in a group get identical rewards, a pairwise judge supplies sequence-level advantages only if both candidate orderings agree; untied groups use vanilla GRPO.
  • Results: ranks first on 6 of 9 benchmarks with the best medical and generalization averages; 38.98 ± 1.04 on HealthBench-Hard from 5,166 prompts vs. InfiMed-ORBIT's 33.60 (8K) and 37.30 (28K). Blinded ratings by two medical experts favored the full pipeline's rubric panels over single-generator panels.
item →