LLMs-Healthcare: Current Applications and Challenges of Large Language Models in various Medical Specialties

arXiv:2311.12882 · cs.CL, cs.AI · Submitted 2024-02-26 · Read on arXiv

Ummara Mumtaz, Awais Ahmed, Summaya Mumtaz

cs.CL, cs.AI

Submitted: 2024-02-26

Updated: 2026-08-11

Comments: 26 pages and one figure

Journal ref: Artificial Intelligence in Health ,Published online: 2 April 2024

DOI: 10.36922/aih.2558

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 37/100

The gist: This paper provides a "comprehensive overview of the latest advancements in utilizing Large Language Models (LLMs) within the healthcare sector, emphasizing their transformative impact across various

Terminology

Summary

This paper provides a comprehensive overview of the latest advancements in utilizing Large Language Models (LLMs) within the healthcare sector, emphasizing their transformative impact across various medical domains. The authors note that LLMs represent a paradigm shift in AI's capability to understand, generate, and interact using human language and are increasingly utilized to tackle critical challenges in patient care, including assisting healthcare professionals in making complex diagnostic decisions, and easing the administrative burdens often associated with healthcare provision.

Applications Across Medical Specialties

  • Cancer Care (Oncology): LLMs are being explored to enhance diagnostic accuracy, personalize therapy options, and streamline patient care. In studies regarding breast tumor boards, ChatGPT's recommendations aligned with the tumor board's decisions in seven out of the ten cases, marking a 70% concordance. In emergency room triage for metastatic prostate cancer, ChatGPT demonstrated a high sensitivity of 95.7% for determining patient admissions. Regarding radiologic decision-making, GPT-4 outperformed ChatGPT-3.5 in accuracy for breast cancer screening and breast pain imaging. For advanced solid tumors, ChatGPT showed a reasonably high agreement between ChatGPT's suggestions and the NCCN guidelines, with an average valid therapy quotient (VTQ) of 0.77. However, limitations exist, as the model generally offered non-specific recommendations and made errors in patient-specific therapy suggestions in some primary breast cancer studies. In CNS tumor management, while ChatGPT showed competence in recommending adjuvant treatments, its diagnostic accuracy was limited, with a notable discrepancy in glioma classifications.

  • Dermatology: To address the shortage of specialized medical professionals and the complexity of interpreting skin imagery, researchers introduced SkinGPT-4, an innovative interactive dermatology diagnostic system underpinned by an advanced visual Large Language Model. Trained on 52,929 images, SkinGPT-4 enabled [it] to articulate medical features in skin disease images using natural language and make precise diagnoses. In evaluations, 78.76% of the diagnoses rendered by SkinGPT-4 were validated as either accurate or relevant by the dermatologists.

  • Neurodegenerative Disorders (Dementia & Alzheimer's): LLMs are used for predicting neurodegenerative disorders. In studies using Mayo Clinic cases, ChatGPT-4 led the pack with an impressive accuracy rate of 84%, compared to 76% for both ChatGPT-3.5 and Google Bard. Additionally, GPT-3's text embeddings offer a promising approach for early dementia diagnosis by outperforming traditional acoustic methods when analyzing spontaneous speech. Furthermore, the AD-BERT model, which utilizes a BERT framework, outperformed seven baseline models in predicting the progression from Mild Cognitive Impairment (MCI) to Alzheimer's Disease (AD) using unstructured EHR notes.

  • Dentistry: While research is notably scarce, emerging frameworks utilize Multi-Modal LLMs to incorporate visual, auditory, and textual data. This allows for comprehensive analysis, such as using vision-language models to evaluate dental X-rays and CT scans to detect any decay on the tooth and subsequently propose a comprehensive treatment plan.

  • Mental Health (Psychiatry and Psychology): LLMs can refine diagnostic precision, optimize treatment outcomes, and enable more tailored patient care. Specifically, Med-PaLM 2 demonstrated its prowess in evaluating psychiatric states... showcasing an impressive accuracy rate ranging between 80% and 84% when predicting psychiatric risk from narratives. Research also shows that instruction fine-tuning notably enhances LLMs' effectiveness in predicting mental states from online text. However, the use of conversational agents like Replika has shown mixed outcomes, as they may struggle with content moderation, consistent interactions, memory retention, and preventing user dependency.

  • Other Specialties:

  • Nephrology: LLMs assist in diagnosing kidney diseases, providing treatment guidance, and monitoring renal function. In multiple-choice testing, GPT-4 demonstrated superior performance, garnering a score of 73.3%, outperforming Claude 2 (54.4%).

  • Gastroenterology: ChatGPT has shown substantial potential in providing valuable insights when answering queries regarding symptoms, diagnostic tests, and treatments.

  • Allergy and Immunology: Models like GPT-4 and Google Med-PaLM2 significantly enhance the diagnostic process and can tailor treatment plans to suit individual patient needs.

Handling Diverse Medical Data Types

The paper outlines several methods for processing medical data for LLM input:

  • Clinical Notes: These serve as rich patient information repositories and are preprocessed to ensure they are in a format that's easily digestible, such as anonymizing patient data to maintain privacy.

  • X-rays/Images: Images are often processed by computer-aided detection (CAD) model[s] to derive the outputs in tensor form, which are then translated into natural language for the LLM.

  • Radiological Reports: These are processed as texts within the report to be input for LLMs in medicine.

  • Speech Data: Audio is converted into a textual format through automatic speech recognition (ASR) systems, with models like Wav2vec 2.0 emerging as a leading contender.

  • Tabular Data: This requires serialization [of] the feature columns into coherent sequences of natural language tokens that the LLM can interpret.

Challenges and Conclusion

The integration of LLMs faces significant hurdles, including accuracy and precision, the capacity of LLMs to consider the comprehensive clinical picture, and integration of LLMs into existing medical workflows. Ethical and practical concerns include data privacy, interoperability, ensuring content sensitivity and safety, and the need for human oversight in verifying the information provided by LLMs. The authors conclude that while the horizon of LLMs in healthcare is expansive and promising, these models should be viewed as supplementary tools that augment, rather than replace, the expertise of medical professionals.

Improvements for AI systems

1. Retrieval-Augmented Generation (RAG) for Oncology Decision Support

  • Improvement: Integrate real-time clinical decision support (CDS) databases and updated NCCN guidelines directly into the LLM via a RAG framework.

  • Capability: The system will generate highly specific, patient-tailored therapy suggestions by cross-referencing a patient's unique biomarkers and tumor classifications against the most current medical literature, eliminating the non-specific recommendations and guideline discrepancies currently observed in oncology tasks.

2. Stateful, Guardrailed Conversational Architectures for Mental Health

  • Improvement: Implement long-term memory (LTM) modules and real-time, high-sensitivity safety/content-moderation layers.

  • Capability: The system will provide consistent, empathetic therapeutic interactions that remember a patient's long-term history to avoid repetitive or contradictory responses, while simultaneously detecting high-risk psychological markers to trigger immediate human intervention, preventing user dependency and ensuring safety.

3. Unified Multimodal Latent Fusion Models

  • Improvement: Move away from text-translation of images (CAD-to-text) toward end-to-end multimodal transformers that process visual, tabular, and textual data in a single latent space.

  • Capability: The system will synthesize a comprehensive clinical picture by directly correlating visual anomalies in X-rays or CT scans with biochemical fluctuations in tabular lab results and nuances in clinical notes, providing a holistic diagnostic reasoning path rather than analyzing data types in isolation.

4. Longitudinal Temporal Transformers for Neurodegenerative Prediction

  • Improvement: Develop transformer architectures specifically designed for longitudinal temporal analysis of speech embeddings and unstructured EHR notes.

  • Capability: The system will track subtle, minute linguistic and acoustic shifts in a patient's speech over months or years, allowing for the high-accuracy prediction of the progression from Mild Cognitive Impairment (MCI) to Alzheimer’s Disease (AD) by identifying patterns that traditional models miss.

5. Automated Cross-Modal Dental Treatment Planning

  • Improvement: Build specialized Vision-Language Models (VLM) trained on paired dental imaging (X-ray/CT) and corresponding clinical treatment records.

  • Capability: The system will automatically detect dental decay or structural abnormalities from visual scans and immediately generate a comprehensive, multi-modal treatment plan that integrates the visual findings with the patient's textual medical history.

Abstract

We aim to present a comprehensive overview of the latest advancements in utilizing Large Language Models (LLMs) within the healthcare sector, emphasizing their transformative impact across various medical domains. LLMs have become pivotal in supporting healthcare, including physicians, healthcare providers, and patients. Our review provides insight into the applications of Large Language Models (LLMs) in healthcare, specifically focusing on diagnostic and treatment-related functionalities. We shed light on how LLMs are applied in cancer care, dermatology, dental care, neurodegenerative disorders, and mental health, highlighting their innovative contributions to medical diagnostics and patient care. Throughout our analysis, we explore the challenges and opportunities associated with integrating LLMs in healthcare, recognizing their potential across various medical specialties despite existing limitations. Additionally, we offer an overview of handling diverse data types within the medical field.

Sources

Related papers