OpenMedLM: Prompt engineering can out-perform fine-tuning in medical question-answering with open-source large language models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "OpenMedLM: Prompt engineering can out-perform fine-tuning in medical question-answering with open-source large language models".
Jane: The paper was written by Jenish Maharjana, Anurag Garikipatia, Navan Preet Singha, Leo Cyrusa, Mayank Sharmaa et al. from Montera, Inc dba Forta.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: We are looking at this paper titled "OpenMedLM: Prompt engineering can out-perform fine-tuning in medical question-answering with open-source large language models," and it sets a huge topic for discussion right from the start.
Jane: It’s great to see the names of one.AI and all those researchers listed, but it really tells us that we’re looking at a collective effort to challenge current boundaries in healthcare AI.
Lu: The fact that this title points to open-source models is a massive deal, suggesting we aren're not limited to proprietary systems like GPT-four or PaLM.
Meng: I’m interested in the practical implications of this approach; if the model is open source, it means implementation costs can drop significantly for developing custom medical tools.
Lalam: It implies a future where the most advanced medical knowledge isn't gated by massive corporate budgets, which changes the entire landscape of who can access quality AI assistance.
Tom: The initial results in this paper are quite impressive, particularly when you look at how they tested Yi 34B against those four key benchmarks: MedQA, MedMCQA, PubMedQA, and the MMLU medical subset.
Jane: These benchmarks represent incredibly complex tests for how well the AI truly understands clinical knowledge across different domains.
Lu: The overall findings show a clear trend toward superior performance in these open-source models compared to previous records that required heavy fine-tuning.
Meng: They aren't just making marginal gains; they are maximizing the performance of existing, accessible tools by pushing the limits of what they can do right out of the box.
Lalam: It’s about unlocking latent knowledge that was already present in these foundation models and making it useful through a smart prompting strategy.
Tom: The results are quite striking, especially seeing that seventy-two point six percent accuracy on MedQA is a big deal, but how did they achieve this without the heavy fine-tuning?
Jane: They found that OpenMedLM was achieving state-of-the-art performance for open-source models, which is a huge leap forward for the accessibility of these tools.
Lu: It proves that even generalist models have deep, complex knowledge that we can just "prompt" out when we know how to ask the right questions.
Meng: It’s a practical demonstration that resource efficiency and high performance aren't competing concepts in this domain.
Lalam: This is about making sure the power of medical knowledge isn't locked behind proprietary walls or massive computational budgets, democratizing access for everyone who needs it.
Summary: Tom: We’ve seen the impressive results from the open-source models in "OpenMedLM: Prompt engineering can out-perform fine-tuning in medical question-answering with open-source large language models," but how did they manage to achieve such high scores?
Jane: The paper dives deep into the specific components of OpenMedLM, which are really what make this experiment unique compared to just using a simple prompt.
Lu: I love how they're combining techniques like zero-shot and CoT (Chain-of-Thought) to build up the reasoning step by step, mimicking human expertise.
Meng: It shows how you can inject specific, high-quality context into a massive model without having to retrain millions of parameters which would take weeks.
Lalam: This creates a culture where the AI isn't just guessing; it's actively thinking through the possibilities based on the structure we provide it.
Tom: And then, as they build up those prompts, they show us how each component adds value in an ablation study to prove which is most effective.
Jane: The way they combine CoT with kNN selection is particularly clever; it's finding relevant examples from training data to ground the reasoning process.
Lu: I think that’s what makes the whole system feel robust, mimicking a structured thought process rather than just relying on random chance.
Meng: It provides a clear pathway for optimization, letting us see exactly which components contribute most to performance gains in a predictable way.
Lalam: This is about creating an active collaborator with the AI, using its knowledge base in a highly structured way that promotes deeper understanding of the subject matter.
Improvements: Tom: We've seen the mechanics of OpenMedLM and its impressive results, but what does this all mean for the real world? The paper suggests that prompt engineering can achieve SOTA performance on these benchmarks without needing to become a specialized model itself.
Jane: This is a huge departure from the idea that specialized knowledge always requires expensive, dedicated models and massive computational power.
Lu: It truly shows that the inherent potential of generalist AI models, when guided correctly by one.AI's team, is extremely powerful enough to surpass specialized tuning efforts.
Meng: We have to be careful though; while this is fantastic for academic and multiple-choice tasks like MedQA, real clinical scenarios are much messier than a standardized test question.
Lalam: The implication is that we can build more accessible tools for doctors and patients who rely on open-source infrastructure, making AI less of an exclusive luxury.
Tom: It’s clear the future isn't solely about building bigger models; it's about better instructing and guiding the ones we already have available.
Jane: And while the paper acknowledges limitations—like needing to handle complex, open-ended clinical scenarios—it still offers a very promising path forward for continuous improvement.
Lu: We're seeing that the future isn't just about scale, but about how much better instructing and guiding existing models is in a new way of thinking.
Meng: I think we need to keep refining the prompt engineering techniques so that the practical deployment moves beyond these academic benchmarks into real-world application.
Lalam: We should be excited to start using OpenMedLM as a foundational starting point for developing AI tools that support global health initiatives and patient care worldwide.
The Wrap-up: Tom: It's clear we have a lot of exciting avenues for future research here that are worth watching closely, especially since the prompt engineering approach is so effective.
Jane: We’ve learned that OpenMedLM provides a way for open-source AI to achieve strong performance on medical tasks through smart prompting, which is a huge win for accessibility.
Lu: It truly shows that emergent abilities are present in these models when we guide them properly, proving the power of prompt engineering is immense in this field.
Meng: We need to keep researching how to make this practical, ensuring that the implementation is robust and can handle real-world complexities outside these benchmarks.
Lalam: I hope this opens the door for more accessible AI tools that will truly benefit patients worldwide by empowering them with knowledge through smart design.
Tom: The entire team agrees that we’ve seen a massive leap forward in how we approach these medical tasks today.
Lu: I think the authors have really delivered on their promise to showcase how prompt engineering can outperform fine-tuning with open-source large language models.
Meng: It definitely sets a high bar for practical application of resource-efficient AI that I want to see achieved in the coming years.
Lalam: For Lalam, this is about using accessible technology to empower healthcare workers globally and improving health equity for everyone.
Montera, Inc dba Forta
cs.CL, cs.AI, cs.IR
Submitted: 2024-02-29
Updated: 2024-02-29
Journal ref: Sci Rep 14, 14156 (2024)
DOI: 10.1038/s41598-024-64827-6
Code: https://github.com/01-ai/Yi
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 85/100
The gist: OpenMedLM presents a novel prompting platform designed to achieve state-of-the-art (SOTA) performance for open-source (OS) large language models in medical question answering, demonstrating that
Key concepts
- Prompt Engineering
- This technique involves structuring inputs and providing context to a large language model (LLM) to guide its output. The system uses methods like Chain-of-Thought (CoT) to build reasoning step by step, mimicking structured human expertise without needing model retraining.
- Fine-Tuning
- This is the process of heavily customizing a general AI model using massive amounts of specialized data. The paper suggests that prompt engineering can achieve high performance in medical tasks without requiring this resource-intensive and time-consuming method.
- Open-Source Large Language Models (LLMs)
- These are foundational models whose weights are publicly available, meaning they are not limited to proprietary systems. This accessibility allows developers to build custom medical tools without being restricted by massive corporate budgets.
Terminology
Summary
OpenMedLM presents a novel prompting platform designed to achieve state-of-the-art (SOTA) performance for open-source (OS) large language models in medical question answering, demonstrating that robust prompt engineering can outperform computationally intensive fine-tuning. This research is critically important because it addresses the challenge of expanding equitable access to medical knowledge
by providing a viable alternative to proprietary, expensive models. By showcasing emergent properties within accessible OS foundation models, OpenMedLM provides a path for researchers and developers to achieve high performance without the need for specialized fine-tuning or massive computational resources.
Methodology and Benchmarks
The study evaluated several open-source foundation LLMs ranging in size from 7B to 70B, utilizing Yi 34B as the primary model due to its superior zero-shot performance. The models were tested across four rigorous medical benchmarks: MedQA (USMLE questions), MedMCQA (Indian post-graduate medical exams), PubMedQA (questions based on PubMed abstracts), and the MMLU medical-subset, which contains 1,785 unique questions from nine clinically relevant topics.
Prompt Engineering Techniques
The OpenMedLM platform leverages a series of advanced prompting strategies to optimize the model’s performance:
-
Zero-Shot Prompting: A simple instruction accompanied by the multiple choice question and options.
-
Few-Shot Prompting (FS): Multiple examples are provided as context, including five randomly selected questions or five k-nearest neighbors (kNN) of similar questions from the training set.
-
Chain-of-Thought (CoT) Prompting: The prompt includes intermediate steps of reasoning and deduction, often with a simulated thought process of how a physician might deduce the answer. This is applied using both random and kNN examples.
-
Ensemble/Self-Consistency: The model is run multiple times, and a majority voting strategy is implemented to increase confidence in the output.
The OpenMedLM Framework
The implementation of OpenMedLM follows a sequential ablation study, where each subsequent prompting approach builds upon previous techniques:
-
Zero-shot prompting (Instruction + Question/Options).
-
Few-shot prompting (Instruction + Random Examples).
-
CoT Prompting with random examples (Instruction + Explanation/Answer).
-
kNN CoT Prompting (using the 5 most similar questions identified via kNN algorithm, ensuring accuracy by resubmitting up to three times if GPT-4 fails to generate a correct explanation).
-
Ensemble/Self-Consistency (Running the prompt five times and shuffling options for each run).
This framework allows researchers to observe how each sub-component contributes to the overall performance, providing a detailed understanding of the mechanistic details behind achieving SOTA results on accessible models.
Results and Findings
The combined effect of all sub-components of OpenMedLM achieved OS SOTA performance across three out of the four benchmarks:
-
On MedQA, OpenMedLM delivered 72.6% accuracy, surpassing previous SOTA by 2.4%.
-
On the MMLU medical-subset, it achieved 81.6% accuracy, establishing itself as the first OS LLM to surpass this threshold.
The ablation study demonstrated that while specialized models like Meditron 70B outperformed Yi 34B in zero-shot scenarios (e.g., MedQA at 65.4% vs 58.4%), the the sequential addition of prompting techniques allowed Yi 34B to achieve a higher accuracy than Meditron on the MedQA dataset (72.6%). This confirms that prompt engineering can significantly optimize performance, achieving a total improvement of over 14% over zero-shot prompting in several benchmarks.
Improvements for AI systems
Operationalizing OpenMedLM for AI System Enhancement
Based on the findings in this paper, we are not merely improving a model; we are implementing a robust Prompt Engineering Pipeline that fundamentally shifts how open-source foundation models (specifically Yi 34B or similar large, open-access architectures) achieve high performance in specialized domains like medicine.
The following improvements detail the specific architectural and operational changes required to integrate the OpenMedLM methodology into any AI system:
A. Shift from Fine-Tuning to Prompt Engineering (Optimization Paradigm)
-
Implementation: The core optimization strategy moves away from computationally expensive, task-specific fine-tuning (which introduces risks like catastrophic forgetting) toward a sophisticated, multi-stage prompt construction process.
-
Goal: Achieving state-of-the-art (SOTA) performance on specialized benchmarks using generalist foundation models without the overhead of extensive retraining.
B. Contextual Example Selection via kNN Integration (Relevance)
-
Implementation: Before generating a prompt for any specific question, a k-Nearest Neighbors (kNN) algorithm must be executed against the model's training dataset. This ensures that the five provided few-shot examples are highly contextually relevant to the target question, rather than being randomly selected.
-
Goal: Elevating the quality of in-context learning (ICL) by ensuring prompt examples accurately reflect similar medical concepts, thereby reducing irrelevant information noise.
C. Forced Reasoning via Chain-of-Thought (CoT) Integration (Verifiability)
-
Implementation: The prompt must be structured to compel the model to generate a step-by-step reasoning path (Chain-of-Thought) for every few-shot example provided, along with the correct answer. This forces the the model to demonstrate its deductive process.
-
Goal: Transforming simple multiple-choice selection into a verifiable diagnostic process, allowing developers and end users to trace how the AI arrived at its conclusion.
D. Robustness through Ensemble Voting (Reliability)
-
Implementation: The final prompt must be executed multiple times (5 runs), with each run performed while randomly shuffling the presentation of multiple-choice options. The system then aggregates these results using a majority voting scheme to determine the definitive answer.
-
Goal: Mitigating stochastic variance and ensuring output reliability, even when the underlying foundation model exhibits probabilistic behavior.
E. Sequential Ablation Implementation (Optimization Flow) The components must be added sequentially, not in isolation:
-
Start with Zero-Shot Prompting (Baseline).
-
Add Random Few-Shot Prompting (Initial boost).
-
Add CoT Reasoning to the Few-Shot Examples (Significant performance surge).
-
Replace Random Examples with kNN-selected CoT Examples (Contextual refinement).
-
Apply Self-Consistency/Ensemble Voting to the final prompt structure.
By implementing this OpenMedLM pipeline, the improved AI system will possess the following specific capabilities:
1. High-Accuracy Medical Question Answering:
- The system can accurately answer complex, multi-choice medical questions sourced from high-stakes examinations (e.g., USMLE/MedQA and Indian PG exams/MedMCQA at levels significantly higher than standard open-source models).
2. Justification of Answers:
- Unlike traditional black-box LLMs, the system provides a Chain-of-Thought explanation for every answer, allowing it to act as a preliminary diagnostic aid where the reasoning is auditable.
3. Open and Accessible Deployment:
- The system can run on open-source infrastructure (e.g using Yi 34B), providing full transparency and compliance required in healthcare environments, without requiring the massive computational investment of proprietary fine-tuning.
4. Contextually Informed Decision Making:
- The system can leverage its vast training knowledge, but specifically guided by the kNN selection process, ensuring that it draws analogies and examples from its training data that are highly relevant to the specific medical query at hand.
Abstract
LLMs have become increasingly capable at accomplishing a range of specialized-tasks and can be utilized to expand equitable access to medical knowledge. Most medical LLMs have involved extensive fine-tuning, leveraging specialized medical data and significant, thus costly, amounts of computational power. Many of the top performing LLMs are proprietary and their access is limited to very few research groups. However, open-source (OS) models represent a key area of growth for medical LLMs due to significant improvements in performance and an inherent ability to provide the transparency and compliance required in healthcare. We present OpenMedLM, a prompting platform which delivers state-of-the-art (SOTA) performance for OS LLMs on medical benchmarks. We evaluated a range of OS foundation LLMs (7B-70B) on four medical benchmarks (MedQA, MedMCQA, PubMedQA, MMLU medical-subset). We employed a series of prompting strategies, including zero-shot, few-shot, chain-of-thought (random selection and kNN selection), and ensemble/self-consistency voting. We found that OpenMedLM delivers OS SOTA results on three common medical LLM benchmarks, surpassing the previous best performing OS models that leveraged computationally costly extensive fine-tuning. The model delivers a 72.6% accuracy on the MedQA benchmark, outperforming the previous SOTA by 2.4%, and achieves 81.7% accuracy on the MMLU medical-subset, establishing itself as the first OS LLM to surpass 80% accuracy on this benchmark. Our results highlight medical-specific emergent properties in OS LLMs which have not yet been documented to date elsewhere, and showcase the benefits of further leveraging prompt engineering to improve the performance of accessible LLMs for medical applications.
Sources
- MusicLM: Generating Music From Text
- What learning algorithm is in-context learning? Investigations with linear models
- Gemini: A Family of Highly Capable Multimodal Models
- MEDITRON-70B: Scaling Medical Pretraining for Large Language Models
- Mistral 7B
- Mixtral of Experts
- Large Language Models are Zero-Shot Reasoners
- Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine
- Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
- Emergent Abilities of Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering