HealthcareNLP: where are we and what is next?
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "HealthcareNLP: where are we and what is next?".
Tom: This tutorial provides an overview of current achievements and future challenges in Healthcare Natural Language Processing (HealthcareNLP) by structuring the field into three hierarchical layers: data/resource, NLP-Eval, and patients.
Jane: First, who's behind it and why it matters.
Title and authors: Tom: So we've got the title of "HealthcareNLP: where are we and what is next?" from Lifeng Han and his team, and it immediately tells us this paper isn't just a random collection of ideas; it’s an overview meant to guide the entire field forward. They are essentially pointing out what’s missing in existing reviews, which is a big deal.
Jane: Exactly, Tom; they explicitly state that current reviews often overlook things like synthetic data generation for privacy or explainable clinical NLP for better integration, which suggests there's a real need to focus on those specific areas. It sets the stage for what we need to tackle next in the field.
Lu: It’s about identifying these overlooked sub-areas—like how we can use AI to generate synthetic data while keeping patient privacy absolutely tight—and then proposing new methodologies, such as retrieval augmented generation or neural symbolic integration of LLMs and KGs. That level of foresight is quite creative in its scope.
Meng: I’m curious about the practical reality of integrating those advanced methods; for example, how do you actually operationalize a neural symbolic integration when you're dealing with sensitive clinical notes? We need to know if these ideas translate into something that can run reliably in a hospital setting.
Lalam: If we look at the whole picture presented in "HealthcareNLP: where are we and what is next?", I see the potential for AI to create systems that are not only technically sound but also ethically robust, which is crucial for building patient trust. It’s about making sure the technology serves the human element of healthcare effectively.
The paper's summary: Tom: Moving into the summary of "HealthcareNLP: where we are and what is next?", it lays out this clear hierarchical view, breaking the field down into three distinct layers: data/resource, NLP-Eval, and patients. This structure helps us see the whole pipeline from data acquisition right up to the point where a patient receives simplified information.
Jane: That's exactly right; it shows that Healthcare NLP isn't just one big monolithic task but a series of interconnected challenges across these three domains. The summary clearly outlines what we have achieved, like standard NER and relation extraction, while simultaneously highlighting the gaps where new approaches are needed.
Lu: What really stands out in the summary is how they map specific tasks onto those layers; for instance, linking clinical coding to knowledge graphs is positioned as a key area where automation can significantly reduce human workload by connecting entities into structured data. That’s a powerful concept for efficiency.
Meng: So, when they talk about the NLP-Eval layer, are we talking about just standard tasks like NER and RE, or are they pushing toward more complex evaluations that involve things like sentiment analysis on patient interviews? I need to understand the scope of what's being tested at that stage.
Lalam: The summary makes it clear that the evaluation isn't stopping at simple extraction; it includes assessing how models handle nuances like sentiment in patient-centered text data, which speaks directly to understanding the patient experience. That level of depth is really important for building empathetic AI tools.
The paper's improvements: Tom: Now we look at the specific improvements the paper suggests for moving forward; they are essentially advocating for incorporating synthetic data generation, explainable clinical NLP, and advanced techniques like retrieval augmented generation or neural symbolic integration of LLMs and KGs. These are the big technical pushes they propose.
Jane: Those improvements suggest a clear path: we need better ways to handle data scarcity through synthetic methods while simultaneously making our models more transparent so clinicians can trust the results. That combination of generative data and transparency is really compelling for real-world adoption.
Lu: The suggestion for neural symbolic integration with LLMs and KGs is fascinating; it’s about trying to merge the power of large language models with the structured knowledge that KGs provide, which I think opens up incredible possibilities for reasoning over complex medical scenarios. It's a way to give the AI more clinical grounding.
Meng: From an engineering perspective, if we implement explainable AI directly into the pipeline as suggested, how do we ensure that the explanations are accurate and don't just add noise to the process? We need mechanisms that guarantee those justifications are reliable for clinical decision support.
Lalam: I think when you combine synthetic data generation with improved transparency, it gives us a much stronger foundation for building tools that can support shared decision-making between doctors and patients, making that communication layer much more effective.
Conclusion: Tom: So to wrap up the "HealthcareNLP: where we are we and what is next?" discussion, the main point is that the field needs a deliberate focus on these three layers—data resources, evaluation methods, and patient-centric applications—and actively pursue things like synthetic data and explainable clinical NLP.
Jane: That’s right; they emphasize that just having great models isn't enough; we need to solve the underlying problems of data scarcity, privacy concerns, and ensuring these systems are understandable for clinicians. It gives us a very actionable roadmap for where research should be heading next in Healthcare NLP.
Lu: I feel the biggest implication is that by focusing on these interconnected components, we move from just building isolated tools to creating holistic systems that actually serve the entire healthcare workflow, which is where the real potential lies for innovation.
Meng: From a practical standpoint, this paper gives us a clearer set of goals for our teams—we know exactly what kind of data we need to generate and what level of interpretability we should aim for in our deployments. It helps us prioritize development efforts effectively.
Lalam: Ultimately, the vision presented in "HealthcareNLP: where are we and what is next?" points toward an AI that is deeply integrated, ethically sound, and genuinely supportive of both the medical staff and the patients themselves. That’s a very hopeful place to be for this technology.
4D PICTURE consortium
cs.CL
Submitted: 2025-12-09
Updated: 2026-10-07
Comments: Presented Tutorial at LREC 2026 https://lrec2026.info/
Code: https://github.com/4dpicture/HealthNLP
Project page: https://ucrel.github.io/pymusas
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 83/100
The gist: This tutorial provides an overview of current achievements and future challenges in Healthcare Natural Language Processing (HealthcareNLP) by structuring the field into three hierarchical layers:
Key concepts
- Data/Resource Layer
- This layer deals with the foundational aspects of handling clinical data. It includes creating annotation guidelines for privacy, establishing ethical governance for hosting research data, and using synthetic data to solve problems caused by limited or scarce patient information.
- NLP-Eval Layer
- This layer focuses on specific NLP tasks used to evaluate models. These include standard Named Entity Recognition (NER) for finding diseases, relation extraction to identify links between symptoms and treatments, and clinical coding to automate the linking of entities into knowledge graphs.
- Patients Layer
- This layer centers on applications directly involving patients. It covers public involvement, improving health literacy, and using NLP to support shared decision-making by simplifying complex clinical information for better communication between doctors and patients.
Terminology
Summary
This tutorial provides an overview of current achievements and future challenges in Healthcare Natural Language Processing (HealthcareNLP) by structuring the field into three hierarchical layers: data/resource, NLP-Eval, and patients. It matters because it introduces essential sub-areas overlooked in existing reviews, such as synthetic data generation for privacy, explainable clinical NLP for better integration, and methodologies like retrieval augmented generation (RAG) and neural symbolic integration of LLMs and KGs.
The gist
This proposed tutorial focuses on Healthcare Domain Applications of NLP, what we have achieved around HealthcareNLP and the challenges that lie ahead for the future.
Data/Resource Layer
This layer addresses foundational aspects of data handling, encompassing:
-
Annotation guidelines: detailing how clinical documents shall be annotated and anonymised to protect privacy and address ethical concerns.
-
Ethical approval and Data Governance: outlining how data should be hosted and how access is granted for research purposes.
-
Synthetic data: discussing its utility in addressing
data sharing and scarcity issues.
NLP-Eval Layer
This layer covers specific NLP tasks that evaluate models, including:
2.1 Standard named entity recognition (NER) in biomedical and clinical domains (e.g., disease, symptoms, diagnoses, treatments).
2.2 Relation extraction (RE) such as drugs and side-effects
or symptoms and treatments.
2.3 Sentiment analysis of NLP models on patient-centered text data like patient interview transcripts.
2.4 Clinical Coding (normalisation/linking), which automates human work by linking entities into clinical knowledge graphs and dictionaries.
2.5 Machine Translation (MT) for translating clinical text to address low-resource eHealth model training obstacles
or multilingual consultation environments.
2.6 Text Simplification and Summarisation to help patients better understand documents, with methods ranging from historical rule-based, to statistical, and nowadays neural network models with attention mechanisms.
Patients Layer
This layer focuses on patient-centric applications:
3.1 Public and Patient Involvement and Engagement (PPIE) practice examples.
3.2 Patient health literacy.
3.3 Support for shared decision making (SDM) communication between clinicians and patients, including translation, simplification, and summarisation
tasks to support SDM communication support.
Models Applied
The tutorial covers the evolution of NLP approaches used across these layers:
= rule/ML
= LLM+KG
= Retrieval-augmented generation (RAG)
The covered models include traditional rule/ML, LLMs with KG, retrieval-augmented generation, explainable AI (XAI).
Real-world Applications and Future Perspectives
The tutorial includes:
• Real-world applications of projects like the 4D PICTURE consortium.
• Hands-on experience of using healthcareNLP platforms.
Future perspectives map NLP tasks to applications such as Biomedical Literature Mining,
Clinical text understanding and information extraction,
Decision support,
and various other domains including "Pharmacovigilance & Drug Discovery and
Public Health Surveillance."
Target Audience
The tutorial is designed for:
• NLP researchers interested in applying models to healthcare.
• Healthcare practitioners and data scientists learning how models are developed for healthcare.
• Undergraduate and postgraduate students interested in AI and NLP applications, with no prior knowledge required.
Outline Summary
The 4-hour tutorial covers: Data resources (annotation guidelines, anonymisation, governance, synthetic data usage), HealthcareNLP Tasks and Evaluations (NER, RE, Linking, Sentiment Classification), Models applied (rule/statistical/neural models up to LLM+KG and RAG), and Patients layer topics including PPIE, Health Literacy, and SDM. It also includes real-world applications from the UK / NL and EU projects.
Technical Requirements
Participants are expected to have access to a good internet connection for online code and data experiments.
Diversity Considerations
The tutorial specifically includes an MT sub-topic on translating healthcare domain data into low-resource languages for building systems for low-resource communities.
Reading List Highlights (Topic-aware Literature)
Key literature mentioned includes work on: De-identification, Healthcare Relation Extraction, Clinical NER and Linking, Clinical Coding with explainability, Biomedical Text Simplification (like the PLABA task), Synthetic Data for HealthcareNLP, Annotation Guidelines, Machine Translation in the biomedical domain, Sentiment Modelling of patient-centered text data (e.g., Dumbach et al., 2024), LLM prompting (in-context learning), and AI for patient-clinician communication.
Presenters
The tutorial is presented by Lifeng Han, Paul Rayson, Suzan Verberne, Andrew Moore, and Goran Nenadic.
Improvements for AI systems
Here are specific improvements to AI systems derived from the concepts in this paper, categorized by the layer they address:
) Data/Resource Layer Improvements:
-
Enhance data privacy and ethical compliance by integrating automated de-identification pipelines using methods like those described in Shaji et al. (2025) and comprehensive risk assessments (Dhivin Shaji et al., 2025).
-
Implement robust data governance frameworks that mandate clear protocols for hosting, access control, and usage permissions for clinical data to ensure ethical research practices.
-
Develop synthetic data generation modules leveraging LLMs and Knowledge Graphs (LLM+KG) techniques (e.g., Mlm4synmed) to create high-quality, privacy-preserving datasets that address data scarcity issues in specialized clinical domains.
) NLP-Eval Layer Improvements:
-
Deploy multi-modal evaluation pipelines that combine traditional rule/ML methods with state-of-the-art neural models (attention mechanisms) for comprehensive task assessment across NER, RE, and sentiment analysis.
-
Integrate Explainable AI (XAI) techniques directly into the clinical NLP pipeline (e.g., using hierarchical label-wise attention networks like Dong et al., 2021) to provide transparent justifications for automated clinical coding and entity linking decisions.
-
Implement Retrieval-Augmented Generation (RAG) systems tailored for clinical text, allowing models to ground their outputs in specific patient records or knowledge bases, thereby improving factual accuracy in relation extraction and summarization tasks.
-
Develop specialized Machine Translation (MT) components optimized for low-resource languages within the biomedical domain to facilitate multilingual consultations and global eHealth model training.
-
Create automated clinical coding systems that map extracted entities into structured Clinical Knowledge Graphs (KGs), significantly reducing manual effort while maintaining high precision through entity linking strategies (e.g., stacked and voted ensembles on LLMs).
) Patients Layer Improvements:
-
Develop Patient-Facing NLP interfaces capable of performing real-time text simplification and summarization of complex clinical documents, translating technical jargon into lay language for improved health literacy.
-
Build Shared Decision Making (SDM) support agents that analyze patient consultation transcripts or records to identify key points, translate them into understandable summaries, and proactively suggest information for clinicians to use in patient-clinician communication.
-
Implement Patient Public Involvement and Engagement (PPIE) tools that allow patients to interact with AI systems to better understand their diagnoses, treatment plans, and personalized health information.
) Overall System Capabilities:
The improved AI system would function as a comprehensive Healthcare NLP Co-pilot
capable of:
-
Accurately extracting structured clinical data (NER/RE/Linking) from unstructured notes.
-
Providing interpretable reasoning for its extracted results (XAI).
-
Generating simplified, accessible summaries of complex medical information for patients and clinicians alike (Simplification/Summarization).
-
Supporting multilingual interactions and translation in clinical settings.
-
Synthesizing new, high-quality training data that respects patient privacy constraints for the development of future healthcare models.
Abstract
This tutorial focused on Healthcare Domain Applications of NLP, what we have achieved around HealthcareNLP, and the challenges that lie ahead for the future. Existing reviews in this domain either overlook some important tasks, such as synthetic data generation for addressing privacy concerns, or explainable clinical NLP for improved integration and implementation, or fail to mention important methodologies, including retrieval augmented generation and the neural symbolic integration of LLMs and KGs. In light of this, the goal of this tutorial is to provide an introductory overview of the most important sub-areas of a patient- and resource-oriented HealthcareNLP, with three layers of hierarchy: data/resource layer: annotation guidelines, ethical approvals, governance, synthetic data; NLP-Eval layer: NLP tasks such as NER, RE, sentiment analysis, and linking/coding with categorised methods, leading to explainable HealthAI; patients layer: Patient Public Involvement and Engagement (PPIE), health literacy, translation, simplification, and summarisation (also NLP tasks), and shared decision-making support. A hands-on session will be included in the tutorial for the audience to use HealthcareNLP applications. The target audience includes NLP practitioners in the healthcare application domain, NLP researchers who are interested in domain applications, healthcare researchers, and students from NLP fields. The type of tutorial is "Introductory to CL/NLP topics (HealthcareNLP)" and the audience does not need prior knowledge to attend this. Tutorial materials: https://github.com/4dpicture/HealthNLP
Sources
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering