Fathom-Vaidya: Advancing Medical Reasoning with Rubric-Based Rewards
cs.AI, cs.CL, cs.LG
Submitted: 2026-09-21
Updated: 2026-09-21
Comments: 18 pages, 5 Figures, Correspondence to kunal.singh@fractal.ai
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Deploying Large Language Models (LLMs) in healthcare requires robust performance across two complementary dimensions - diagnostic reasoning: the convergent, evidence-driven task of inferring a
Terminology
Abstract
Deploying Large Language Models (LLMs) in healthcare requires robust performance across two complementary dimensions - diagnostic reasoning: the convergent, evidence-driven task of inferring a patient's condition from clinical data to produce a diagnosis, and clinical healthcare reasoning: the broader, navigational judgment required to communicate, plan, and adapt across multi-turn clinical interactions where a single correct answer may not exist. Recent benchmarks such as HealthBench and MedXpertQA reveal persistent weaknesses in both areas, exposing failures in complex diagnostic scenarios and limitations in contextual, patient-centered dialogue. We introduce a sequential training framework that targets these facets using synthetic data and rubric-based reinforcement learning. First, we improve diagnostic reasoning using MedBullets-derived questions with rule- and rubric-guided Reinforcement Learning (RL). We then shift to clinical reasoning by generating 5.3k synthetic multi-turn scenarios, each paired with multi-dimensional rubrics to comprehensively assess the response. This approach yields over 10% improvement on MedXpertQA, and our 30B model achieves 50.1% accuracy on HealthBench-Hard, surpassing proprietary baselines including GPT-5 (thinking). Our results show that targeted synthetic datasets and rubric-based training can systematically improve both diagnostic and interactive clinical reasoning in medical LLMs.
Sources
- HealthBench: Evaluating Large Language Models Towards Improved Human Health
- HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs
- HuatuoGPT-II, One-stage Training for Medical Adaption of LLMs
- Evaluating Large Language Models Trained on Code
- Training Verifiers to Solve Math Word Problems
- HealthBench Professional: Evaluating Large Language Models on Real Clinician Chats
- m1: Unleash the Potential of Test-Time Scaling for Medical Reasoning with Large Language Models
- What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams
- Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- Baichuan-M2: Scaling Medical Capability with Large Verifier System
- MedReason: Eliciting Factual Medical Reasoning Steps in LLMs via Knowledge Graphs
- MedCaseReasoning: Evaluating and learning diagnostic reasoning from clinical case reports
- Qwen3 Technical Report
- Self-Rewarding Rubric-Based Reinforcement Learning for Open-Ended Reasoning
- Group Sequence Policy Optimization
- DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models
- MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection