LLMs-Healthcare: Current Applications and Challenges of Large Language Models in various Medical Specialties
Ummara Mumtaz, Awais Ahmed, Summaya Mumtaz
cs.CL, cs.AI
Submitted: 2024-02-26
Updated: 2026-08-11
Comments: 26 pages and one figure
Journal ref: Artificial Intelligence in Health ,Published online: 2 April 2024
DOI: 10.36922/aih.2558
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 37/100
The gist: This paper provides a "comprehensive overview of the latest advancements in utilizing Large Language Models (LLMs) within the healthcare sector, emphasizing their transformative impact across various
Terminology
Summary
This paper provides a comprehensive overview of the latest advancements in utilizing Large Language Models (LLMs) within the healthcare sector, emphasizing their transformative impact across various medical domains.
The authors note that LLMs represent a paradigm shift in AI's capability to understand, generate, and interact using human language
and are increasingly utilized to tackle critical challenges in patient care,
including assisting healthcare professionals in making complex diagnostic decisions, and easing the administrative burdens often associated with healthcare provision.
Applications Across Medical Specialties
-
Cancer Care (Oncology): LLMs are being explored to
enhance diagnostic accuracy, personalize therapy options, and streamline patient care.
In studies regarding breast tumor boards,ChatGPT's recommendations aligned with the tumor board's decisions in seven out of the ten cases, marking a 70% concordance.
In emergency room triage for metastatic prostate cancer,ChatGPT demonstrated a high sensitivity of 95.7% for determining patient admissions.
Regarding radiologic decision-making,GPT-4 outperformed ChatGPT-3.5
in accuracy for breast cancer screening and breast pain imaging. For advanced solid tumors, ChatGPT showed areasonably high agreement between ChatGPT's suggestions and the NCCN guidelines,
with an average valid therapy quotient (VTQ) of 0.77. However, limitations exist, as the modelgenerally offered non-specific recommendations
andmade errors in patient-specific therapy suggestions
in some primary breast cancer studies. In CNS tumor management, while ChatGPT showedcompetence in recommending adjuvant treatments,
itsdiagnostic accuracy was limited, with a notable discrepancy in glioma classifications.
-
Dermatology: To address the
shortage of specialized medical professionals
and the complexity of interpreting skin imagery, researchers introducedSkinGPT-4, an innovative interactive dermatology diagnostic system underpinned by an advanced visual Large Language Model.
Trained on 52,929 images, SkinGPT-4enabled [it] to articulate medical features in skin disease images using natural language and make precise diagnoses.
In evaluations,78.76% of the diagnoses rendered by SkinGPT-4 were validated as either accurate or relevant by the dermatologists.
-
Neurodegenerative Disorders (Dementia & Alzheimer's): LLMs are used for
predicting neurodegenerative disorders.
In studies using Mayo Clinic cases,ChatGPT-4 led the pack with an impressive accuracy rate of 84%,
compared to 76% for both ChatGPT-3.5 and Google Bard. Additionally,GPT-3's text embeddings offer a promising approach for early dementia diagnosis
by outperforming traditional acoustic methods when analyzing spontaneous speech. Furthermore, theAD-BERT
model, which utilizes a BERT framework,outperformed seven baseline models
in predicting the progression from Mild Cognitive Impairment (MCI) to Alzheimer's Disease (AD) using unstructured EHR notes. -
Dentistry: While research is
notably scarce,
emerging frameworks utilizeMulti-Modal LLMs
to incorporatevisual, auditory, and textual data.
This allows forcomprehensive analysis,
such as using vision-language models to evaluate dental X-rays and CT scans todetect any decay on the tooth
and subsequentlypropose a comprehensive treatment plan.
-
Mental Health (Psychiatry and Psychology): LLMs can
refine diagnostic precision, optimize treatment outcomes, and enable more tailored patient care.
Specifically,Med-PaLM 2 demonstrated its prowess in evaluating psychiatric states... showcasing an impressive accuracy rate ranging between 80% and 84%
when predicting psychiatric risk from narratives. Research also shows thatinstruction fine-tuning notably enhances LLMs' effectiveness
in predicting mental states from online text. However, the use of conversational agents like Replika has shownmixed outcomes,
as they may struggle withcontent moderation, consistent interactions, memory retention, and preventing user dependency.
-
Other Specialties:
-
Nephrology: LLMs assist in
diagnosing kidney diseases, providing treatment guidance, and monitoring renal function.
In multiple-choice testing,GPT-4 demonstrated superior performance, garnering a score of 73.3%,
outperforming Claude 2 (54.4%). -
Gastroenterology: ChatGPT has shown
substantial potential in providing valuable insights
when answering queries regarding symptoms, diagnostic tests, and treatments. -
Allergy and Immunology: Models like GPT-4 and Google Med-PaLM2
significantly enhance the diagnostic process
andcan tailor treatment plans to suit individual patient needs.
Handling Diverse Medical Data Types
The paper outlines several methods for processing medical data for LLM input:
-
Clinical Notes: These
serve as rich patient information repositories
and arepreprocessed to ensure they are in a format that's easily digestible,
such asanonymizing patient data to maintain privacy.
-
X-rays/Images: Images are often processed by
computer-aided detection (CAD) model[s]
toderive the outputs in tensor form,
which are thentranslated into natural language
for the LLM. -
Radiological Reports: These are
processed as texts within the report to be input for LLMs in medicine.
-
Speech Data: Audio is
converted into a textual format through automatic speech recognition (ASR) systems,
with models likeWav2vec 2.0 emerging as a leading contender.
-
Tabular Data: This requires
serialization [of] the feature columns into coherent sequences of natural language tokens that the LLM can interpret.
Challenges and Conclusion
The integration of LLMs faces significant hurdles, including accuracy and precision,
the capacity of LLMs to consider the comprehensive clinical picture,
and integration of LLMs into existing medical workflows.
Ethical and practical concerns include data privacy, interoperability,
ensuring content sensitivity and safety,
and the need for human oversight in verifying the information provided by LLMs.
The authors conclude that while the horizon of LLMs in healthcare is expansive and promising,
these models should be viewed as supplementary tools that augment, rather than replace, the expertise of medical professionals.
Improvements for AI systems
1. Retrieval-Augmented Generation (RAG) for Oncology Decision Support
-
Improvement: Integrate real-time clinical decision support (CDS) databases and updated NCCN guidelines directly into the LLM via a RAG framework.
-
Capability: The system will generate highly specific, patient-tailored therapy suggestions by cross-referencing a patient's unique biomarkers and tumor classifications against the most current medical literature, eliminating the
non-specific recommendations
and guideline discrepancies currently observed in oncology tasks.
2. Stateful, Guardrailed Conversational Architectures for Mental Health
-
Improvement: Implement long-term memory (LTM) modules and real-time, high-sensitivity safety/content-moderation layers.
-
Capability: The system will provide consistent, empathetic therapeutic interactions that remember a patient's long-term history to avoid repetitive or contradictory responses, while simultaneously detecting high-risk psychological markers to trigger immediate human intervention, preventing user dependency and ensuring safety.
3. Unified Multimodal Latent Fusion Models
-
Improvement: Move away from
text-translation
of images (CAD-to-text) toward end-to-end multimodal transformers that process visual, tabular, and textual data in a single latent space. -
Capability: The system will synthesize a
comprehensive clinical picture
by directly correlating visual anomalies in X-rays or CT scans with biochemical fluctuations in tabular lab results and nuances in clinical notes, providing a holistic diagnostic reasoning path rather than analyzing data types in isolation.
4. Longitudinal Temporal Transformers for Neurodegenerative Prediction
-
Improvement: Develop transformer architectures specifically designed for longitudinal temporal analysis of speech embeddings and unstructured EHR notes.
-
Capability: The system will track subtle, minute linguistic and acoustic shifts in a patient's speech over months or years, allowing for the high-accuracy prediction of the progression from Mild Cognitive Impairment (MCI) to Alzheimer’s Disease (AD) by identifying patterns that traditional models miss.
5. Automated Cross-Modal Dental Treatment Planning
-
Improvement: Build specialized Vision-Language Models (VLM) trained on paired dental imaging (X-ray/CT) and corresponding clinical treatment records.
-
Capability: The system will automatically detect dental decay or structural abnormalities from visual scans and immediately generate a comprehensive, multi-modal treatment plan that integrates the visual findings with the patient's textual medical history.
Abstract
We aim to present a comprehensive overview of the latest advancements in utilizing Large Language Models (LLMs) within the healthcare sector, emphasizing their transformative impact across various medical domains. LLMs have become pivotal in supporting healthcare, including physicians, healthcare providers, and patients. Our review provides insight into the applications of Large Language Models (LLMs) in healthcare, specifically focusing on diagnostic and treatment-related functionalities. We shed light on how LLMs are applied in cancer care, dermatology, dental care, neurodegenerative disorders, and mental health, highlighting their innovative contributions to medical diagnostics and patient care. Throughout our analysis, we explore the challenges and opportunities associated with integrating LLMs in healthcare, recognizing their potential across various medical specialties despite existing limitations. Additionally, we offer an overview of handling diverse data types within the medical field.
Sources
- Emergent Abilities of Large Language Models
- SkinGPT-4: An Interactive Dermatology Diagnostic System with Visual Large Language Model
- Exploring Multimodal Approaches for Alzheimer's Disease Detection Using Patient Speech Transcript and Audio Data
- Large language models improve Alzheimer's disease diagnosis using multi-modality data
- The Capability of Large Language Models to Measure Psychiatric Functioning
- Mental-LLM: Leveraging Large Language Models for Mental Health Prediction via Online Text Data
- A Comparative Study of Open-Source Large Language Models, GPT-4 and Claude 2: Multiple-Choice Test Taking in Nephrology
- ChatCAD: Interactive Computer-Aided Diagnosis on Medical Image using Large Language Models
- Language Models are Few-shot Learners for Prognostic Prediction
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering