ConRub-Med: Reinforcement Learning with Consensus Rubrics for Open-Ended Medical Question Answering
Taojie Zhu, Yuan Xia, Tao Sun, Yizhi Wang, Yan Chen, Qunshan He, Tian Guan, Jian Wang, Jinjie Gu, Junwei Liu, Yonghong He
Tsinghua University · Ant Group · Zhejiang University
cs.CL
Submitted: 2026-08-11
Updated: 2026-08-12
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: ConRub-Med: Reinforcement Learning with Consensus Rubrics for Open-Ended Medical Question Answering Summary This paper introduces ConRub-Med, a reinforcement learning (RL) approach for open-ended
Terminology
Summary
ConRub-Med: Reinforcement Learning with Consensus Rubrics for Open-Ended Medical Question Answering
Summary
This paper introduces ConRub-Med, a reinforcement learning (RL) approach for open-ended medical question answering that uses consensus-based rubrics to provide feedback during policy optimization. The method addresses the challenge that open-ended medical questions lack cheap, automatic outcome verifiers, unlike mathematics or coding where answers can be checked automatically. The paper states: "Many open-ended medical questions lack a comparably cheap, single outcome verifier. Their answers are not simply right or wrong: they may express the same content in different forms, be partly correct yet incomplete, or contain one unsafe claim amid otherwise useful content."
The method has three main components:
-
Consensus Rubric Construction: Three heterogeneous language models (GPT-5 Mini, Gemini 2.5 Pro, and Claude Sonnet 4.5) independently propose 20–30 atomic criteria for each prompt. A separate reviewer model (DeepSeek-V4-Pro) filters these criteria, retaining only those with semantic support from all three generators. The paper explains:
Three models generate criteria independently. A separate reviewer filters them and retains only criteria supported by all three.
This process produces a training dataset of 5,166 prompts with 49,046 content criteria and 10,332 global controls (59,378 total criterion instances). -
Three-State Scoring: The criterion judge distinguishes between correct coverage, missing information, and incorrect claims, assigning scores of +1, 0, and −1 respectively. The paper notes:
Three-State scoring asks the criterion judge to distinguish correct, missing, and wrong content, assigning negative credit to errors.
This differs from binary scoring, which cannot distinguish omission from incorrect assertions. The reward is computed as the average of these criterion scores plus a soft overlength penalty. -
Pairwise Sequence Advantages for Exact Reward Ties: When all responses in a complete GRPO group receive identical final rewards (which would result in zero vanilla GRPO advantages), a separate pairwise judge compares response pairs in both candidate orders. The paper states: "When every response in a complete Group Relative Policy Optimization (GRPO) group receives the same final reward, a pairwise judge provides sequence advantages only if both candidate orders agree, without changing the scalar rewards." Only order-consistent preferences become nonzero sequence advantages (+1 for winner, −1 for loser), while inconsistent or tied comparisons yield zero. Non-tied groups use vanilla GRPO.
The training objective combines group-relative advantages with asymmetric clipping, using the formula: L GRPO(θ) = −(1/k)Σ(1/T i)Σ min(ρ i,t(θ)A train i,t, clip(ρ i,t(θ), 1−ε low, 1+ε high)A train i,t), with ε low = 0.20 and ε high = 0.28.
Key Results:
-
Main performance: Among evaluated open-weight models, ConRub-Med ranks first on six of nine benchmarks: HealthBench-Hard, MedXpertQA-Text, DiagnosisArena-MCQ, WritingBench, GPQA-Diamond, and MMLU-Medical. It achieves the highest medical average (51.76) and generalization average (74.27).
-
HealthBench-Hard: ConRub-Med scores 38.98 ± 1.04 (mean ± SD), compared with InfiMed-ORBIT's 33.60 with 8,000 samples and 37.30 with 28,000 samples. The entire observed range (37.94–40.02) exceeds the strongest listed InfiMed-ORBIT setting.
-
Expert evaluation: In a blinded study matched by question, two medical experts rated criteria from the full consensus pipeline as more clinically relevant than criteria produced by a single generator. Both experts rated 99.3% of consensus criteria as clinically relevant versus 86.8% and 86.0% for single-generator panels, with gains of 12.5 and 13.2 percentage points (95% CIs [6.6, 19.1] and [6.6, 20.6]).
-
Ablation results: Consensus Rubrics outperform single-generator rubrics (medical average 49.26 vs. 47.51–48.93). Three-State scoring raises the medical average by 2.25 points (from 49.26 to 51.51) while changing generalization scores little. Pairwise Advantages raise the generalization average from 62.96 to 74.27 (including +30.32 on IFEval-Loose) while changing the medical average only slightly (from 51.51 to 51.76).
Analysis findings:
-
Binary scoring reaches higher training rewards (0.780–0.802) than Three-State scoring (0.751–0.753), yet yields lower medical performance, consistent with reward hacking where longer answers accumulate more positive credits even with unsupported content.
-
Exact-tie rates rise from 2.80% in epoch 1 to 7.73% in epoch 3, with reward 1.0 accounting for 64.6% and 85.2% of those ties respectively.
-
Of 847 exact-tie groups, the bidirectional judge accepts 120 preferences and provides nonzero sequence advantages to 90 groups (10.63%). The pairwise judge selects the longer response in 60 of 120 accepted comparisons, indicating no systematic length bias.
-
A case study shows the pairwise judge identifying a unit-conversion error (treating mrem/hr as mrem/s and multiplying by 3,600) that produces identical scalar rewards and criterion profiles.
Limitations include: the three generators may share errors; unanimous support can exclude useful criteria identified by only one or two models; the expert study covers only 17 matched RaR-Medicine questions; pairwise advantages are restricted to exact ties; and controlled policy comparisons use only one training seed due to compute costs.
Improvements for AI systems
Improvements to AI Systems:
-
Implement consensus-based rubric generation for open-ended tasks. Instead of relying on a single LLM to define evaluation criteria, use multiple heterogeneous models to propose criteria and a separate reviewer to retain only those with unanimous semantic support. This reduces single-model bias and increases clinical relevance (99.3% vs. 86% expert-rated relevance). The improved system can generate more robust, unbiased evaluation rubrics for any open-ended domain (e.g., legal advice, counseling, technical support) where correctness is multi-faceted.
-
Adopt three-state scoring (correct/missing/incorrect) instead of binary scoring. The system should assign +1, 0, and −1 for correct, missing, and wrong content respectively. This prevents reward hacking where models pad answers with unsupported but non-contradictory content to accumulate positive credits. The improved system can train models to prioritize factual completeness and penalize hallucinations explicitly, leading to higher-quality outputs in medical and other safety-critical domains.
-
Integrate pairwise sequence advantages for exact reward ties. When all responses in a group receive identical scalar rewards, use a bidirectional pairwise judge to break ties only when both orderings agree. This resolves cases where scalar rewards fail to distinguish subtle errors (e.g., unit-conversion mistakes) that produce identical criterion profiles. The improved system can detect and correct for hidden errors in otherwise uniformly-scored outputs, improving generalization (e.g., +30.32 on IFEval-Loose) without altering scalar reward semantics.
-
Use asymmetric clipping with group-relative advantages. Apply ε low=0.20 and ε high=0.28 in the GRPO objective to allow larger positive updates than negative ones. This encourages exploration of better responses while preventing destructive updates from poor samples. The improved system can train more stably and achieve higher performance ceilings on benchmarks like HealthBench-Hard (38.98 vs. 33.60–37.30 for baseline).
-
Dynamically monitor exact-tie rates during training. Track the percentage of GRPO groups with identical rewards (rising from 2.80% to 7.73% across epochs). When tie rates exceed a threshold, automatically activate the pairwise judge to provide additional learning signal. The improved system can self-regulate its training signal diversity, preventing stagnation in later epochs.
-
Combine consensus rubrics with three-state scoring and pairwise advantages in a unified pipeline. The full system (ConRub-Med) achieves the highest medical average (51.76) and generalization average (74.27) among open-weight models. The improved system can be applied to any open-ended QA task—including clinical decision support, patient education, and diagnostic reasoning—to produce answers that are more complete, safer, and less prone to subtle factual errors.
Abstract
Reinforcement learning with verifiable rewards has been especially effective in mathematics and coding, where answers can be checked automatically. Many open-ended medical questions lack comparably cheap outcome verifiers: responses may be partly correct, incomplete, or contain clinically consequential errors. Rubrics written or validated by physicians offer strong clinical grounding, but involving experts in every instance is costly. Model-generated rubrics make this supervision scalable. We introduce ConRub-Med to preserve useful distinctions as rubric feedback moves from construction to policy optimization. For each prompt, three heterogeneous language models propose atomic criteria independently; a separate model reviews them, retaining only criteria with semantic support from all three generators. Three-State scoring distinguishes correct coverage, missing information, and incorrect claims. Errors receive negative rather than zero credit. When every response in a complete Group Relative Policy Optimization (GRPO) group receives the same final reward, a pairwise judge provides sequence advantages only if both candidate orders agree, without changing the scalar rewards. Groups without ties use vanilla GRPO. In a blinded study matched by question, two medical experts rate panels from the full pipeline as more clinically relevant than panels produced by one generator. Across the evaluated open models, ConRub-Med ranks first on six of nine benchmarks and achieves the highest medical and generalization averages. Using the resulting rubric dataset of 5,166 prompts, it scores 38.98 plus or minus 1.04 (mean plus or minus SD) on HealthBench-Hard, compared with InfiMed-ORBIT's 33.60 with 8,000 samples and 37.30 with 28,000.
Sources
- HealthBench: Evaluating Large Language Models Towards Improved Human Health
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
- Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy
- Bridging Offline and Online Reinforcement Learning for LLMs
- MedGemma 1.5 Technical Report
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Qwen3 Technical Report
- AMIR-GRPO: Inducing Implicit Preference Signals into GRPO
- Instruction-Following Evaluation for Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering