A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench

arXiv:2608.12138 · cs.CL, cs.AI, cs.HC, cs.IR, cs.LG · Submitted 2026-08-12 · Read on arXiv

Praveen Reddy, Charuta Mandke, Suvrankar Datta, Sarah Khan, Siddharth Reddy Anthireddy, Shitij Arora, Vishal Singh

cs.CL, cs.AI, cs.HC, cs.IR, cs.LG

Submitted: 2026-08-12

Updated: 2026-08-13

Comments: 2 tables

Code: https://github.com/openai/simple-evals.4

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 75/100

The gist: This paper evaluates VITA, a retrieval-augmented generation (RAG) system purpose-built for context-specific knowledge retrieval in India, against several frontier general-purpose large language

Terminology

Summary

This paper evaluates VITA, a retrieval-augmented generation (RAG) system purpose-built for context-specific knowledge retrieval in India, against several frontier general-purpose large language models (LLMs) on the HealthBench clinical reasoning benchmark. The authors report that VITA ranked first overall, earning 51.9% of possible rubric points, compared with 46.1% for GPT-5.4, 44.3% for o4-mini, 42.6% for Gemini 3.1 Pro, and 37.3% for Claude Sonnet 4.6 across 4,023 English-language HealthBench questions (80.5% of the full 5,000-question benchmark and 94.7% of its English subset). In head-to-head question-level comparisons, VITA achieved the highest score on 45.4% of questions (1,827 of 4,023) — 2.6 times more than the next best system (GPT-5.4, 716 wins). Per-axis analysis revealed "VITA's advantage was concentrated in clinical accuracy (55.9% vs. 49.5% for GPT-5.4), completeness (51.8% vs. 42.6%), and context awareness (50.3% vs. 45.1%), while general-purpose LLMs scored higher on communication quality and instruction following."

To address objections that comparators had been superseded and that a GPT-family judge may favor GPT-family models, co-authors with no equity or financial interest in VITA re-ran the evaluation on a random 500-question subset against newer models (GPT-5.5, Claude Opus 4.8, Gemini 3.5 Pro, Grok 4.3), graded by DeepSeek-V4-Pro, an open-weight frontier judge sharing no lineage with any system tested. Under these conditions, "VITA and GPT-5.5 were statistically indistinguishable on mean per-question score, but VITA ranked first on points-weighted score and produced the highest-scoring response on more questions (questions won) than any other system." Specifically, VITA scored 51.0% mean per-question (95% CI 48.6–53.4) versus 52.0% for GPT-5.5 (95% CI 49.4–54.5), but VITA won 109 questions versus 80 for GPT-5.5, and VITA's points-weighted score was 49.1% versus 48.3% for GPT-5.5. "VITA's advantages in clinical accuracy and completeness persisted under the neutral judge; its context-awareness advantage did not, and the communication gap, which is the most subjective metric in the Healthbench, widened."

The authors offer two observations. First, VITA's performance demonstrates that a purpose-built clinical AI system can outperform the most capable frontier models on the very benchmark used to claim general LLM superiority. They hypothesize that corpus specificity is a meaningful design variable in RAG-based clinical AI, citing prior work showing that large, unfiltered corpora introduce retrieval noise and that curated corpora combining clinical guidelines with high-quality systematic reviews outperform broad literature databases on clinical question answering tasks. They note VITA's particular strength on LMIC-specific scenarios, giving the example that VITA scored 51 of 67 possible rubric points on a question about Nipah virus exposure from raw date palm sap in Bangladesh, compared with 33 of 67 for GPT-5.4. They also note that general-purpose LLMs scored substantially higher on communication quality, consistent with evidence that HealthBench rubrics encode Western communication norms that may undervalue responses calibrated to other contexts, and cite a prior multi-site study where 37 physician evaluators in India and Bangladesh rated VITA significantly higher than ChatGPT Plus across six clinical dimensions, with physicians' largest rated advantage for VITA lying in evidence quality. Second, the pace of innovation in clinical AI demands evaluation frameworks that are both context-sensitive and continuous, as static, point-in-time benchmarks developed in high-income country contexts cannot fully capture performance across the diversity of settings in which medicine is practiced, nor keep pace with rapid model iteration.

The authors acknowledge important limitations: "The GPT-4.1 judge and physician-written rubrics were developed primarily in Western clinical contexts and may systematically undervalue responses calibrated to LMIC settings. The English-only evaluation does not capture HealthBench's non-English scenarios, where VITA's multilingual corpus may confer additional advantages or face different challenges. Under a neutral judge with current-generation comparators, VITA's aggregate lead was no longer statistically distinguishable from the best frontier model; the top of the ranking should be read as parity rather than a clear first place. Despite these limitations, they conclude that the finding that a purpose-built clinical AI system matches or outperforms frontier general-purpose LLMs on an independent, openly reproducible benchmark — with results verifiable from supplementary data — demonstrates that the question of which class of system performs better is far from settled. The answer depends critically on which specialized systems are evaluated, in which clinical contexts, and against which standards of care."

Improvements for AI systems

Improvements to AI systems:

  1. Add domain-specific retrieval with curated, context-aware corpora. Implement a RAG pipeline that filters and prioritizes clinical guidelines, systematic reviews, and region-specific outbreak protocols (e.g., Nipah virus from date palm sap) over broad, unfiltered literature. This directly boosts clinical accuracy and completeness in low- and middle-income country (LMIC) scenarios.

  2. Introduce a dual-scoring mechanism for subjective outputs. Separate objective clinical reasoning (accuracy, completeness, context awareness) from communication quality. Train a lightweight classifier to detect when a response is clinically correct but stylistically misaligned with Western rubrics, then adjust scoring or generate alternative phrasings without altering medical content.

  3. Implement continuous benchmark re-evaluation with neutral judges. Build an automated evaluation loop that periodically re-tests the system against the latest frontier models using an open-weight judge (e.g., DeepSeek-V4-Pro) on random subsets. This prevents overfitting to a single judge family and detects when performance parity shifts.

  4. Add multilingual and non-English clinical reasoning support. Extend the retrieval corpus and fine-tuning to include non-English clinical scenarios from HealthBench, enabling the system to handle context-specific questions in languages beyond English, where general-purpose LLMs currently underperform.

  5. Develop a context-awareness calibration module. For each query, detect the geographic and cultural setting (e.g., urban India vs. rural Bangladesh) and dynamically adjust retrieval weights toward local guidelines, epidemiological data, and treatment protocols, improving performance on region-specific questions.

  6. Create a points-weighted winner optimization target. Instead of only maximizing mean per-question score, train the system to also maximize the number of questions where it produces the highest-scoring response (as VITA did), which better reflects real-world clinical decision support where a single correct answer matters.

What the improved AI system can do:

  • Achieve parity or superiority over frontier general-purpose LLMs on clinical benchmarks in LMIC contexts, with higher clinical accuracy and completeness.

  • Provide clinically correct answers that are also communicatively appropriate for Western rubrics, or flag when a rubric is culturally biased.

  • Stay current with rapid model iteration by automatically re-benchmarking against new models and adjusting its retrieval and generation strategies.

  • Answer clinical questions in multiple languages with context-specific retrieval, improving safety and relevance in non-English healthcare settings.

  • Win more individual clinical questions (not just average scores), making it more reliable for point-of-care decision support where each answer matters.

Abstract

General-purpose large language models (LLMs) have recently been reported to match or exceed specialized clinical AI tools on medical benchmarks, but such comparisons draw on a narrow set of systems and on benchmarks developed largely in high-income settings. We evaluate VITA, a retrieval-augmented generation (RAG) system purpose-built for contextual knowledge retrieval in India and other low- and middle-income (LMIC) settings. VITA retrieves from a curated corpus of disease-specific guidelines, India-specific antimicrobial resistance data, national formulary constraints, and resource-limited care protocols; its architecture and corpus are proprietary, but the benchmark, the physician-written rubrics, and our full response and scoring outputs are public for independent verification. On 4,023 English-language HealthBench questions (80.5% of the benchmark), scored with a GPT-4.1 judge, VITA ranked first with 51.9% of possible rubric points, ahead of GPT-5.4 (46.1%), o4-mini (44.3%), Gemini 3.1 Pro (42.6%), and Claude Sonnet 4.6 (37.3%), and scored highest on 45.4% of questions. To test robustness to newer models and judge lineage, a 500-question subset was re-run against current-generation models (GPT-5.5, Claude Opus 4.8, Gemini 3.5 Pro, Grok 4.3) and graded by a neutral open-weight judge (DeepSeek-V4-Pro) sharing no lineage with any system tested. Here the gap narrowed to parity: VITA and GPT-5.5 were statistically indistinguishable on mean per-question score, while VITA led on points-weighted score and won the most questions. VITA's advantages in accuracy and completeness persisted under the neutral judge; its communication scores were lower. These results indicate that a purpose-built clinical RAG system remains competitive with frontier LLMs on an open benchmark, consistent with corpus specificity as a design variable that improves grounding at some cost to communication polish.

Sources

Related papers