Domain-Adapted Small Language Models for Reliable Clinical Triage

arXiv:2604.26766 · cs.CL, cs.AI, cs.LG · Submitted 2026-04-29 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Domain-Adapted Small Language Models for Reliable Clinical Triage".

Jane: Accurate and consistent Emergency Severity Index (ESI) assignment remains a persistent challenge in emergency departments, where highly variable free-text triage documentation contributes to mistriage and workflow inefficiencies.

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So, we're talking about this paper today called "Domain-Adapted Small Language Models for Reliable Clinical Triage," and I'm really stoked because it tackles that persistent headache in emergency departments—getting the Emergency Severity Index assignment right when there's so much messy free-text documentation.

Jane: Exactly, Tom; it sounds like the authors are looking at how we can use smaller language models to make that triage process more consistent and less prone to those kinds of mistakes.

Lu: This research is fascinating because it zeroes in on making these tools reliable for real-world clinical settings where things move fast and documentation is often a bit chaotic.

Meng: I'm curious, does this mean we're looking at something that could actually be put into practice quickly, or are we still stuck with massive compute demands?

Lalam: From my perspective as an AI, the core idea here is about building models that deeply understand the specific language and context of pediatric triage documentation so they can make safer decisions.

Tom: Right, Lalam; it's not just about general language understanding; it’s about domain adaptation to a very specific medical environment.

Jane: That’s right, Lalam; the paper claims that by fine-tuning these models on data from places like Children’s National Hospital, they can substantially lower errors in how ESI levels are assigned compared to using standard methods.

Lu: And it seems they found that using specific formats for the input data made a huge difference in how accurate those predictions turned out.

Tom: It sounds like the main thrust of "Domain-Adapted Small Language Models for Reliable Clinical Triage" is showing that smaller, specialized models can perform very well when trained on highly relevant clinical examples, which is a really important concept for accessibility too.

Meng: But what does this mean practically? Can we actually deploy these kinds of models in a hospital setting without needing huge infrastructure?

Lalam: The work suggests that the Qwen2 point 5-7B model, after being adapted using Quantized Low-Rank Adaptation on their silver dataset, showed the best balance when it came to accuracy and efficiency for this task.

Jane: So we're seeing a model that’s not just accurate but also relatively light and stable during training.

Paper summary: Tom: And they did some serious testing across different ways of feeding the model information, comparing raw notes versus clinical vignettes, which tells us exactly what kind of input helps the AI understand the situation best.

Lu: That systematic comparison of prompting pipelines is a very rigorous way to show where these models actually shine in a complex environment.

Jane: It’s helpful to see how they systematically explored different ways of getting information into the model, moving from just raw text to structured data, and it points toward a clearer path for how we might build better triage tools.

Meng: From an engineering standpoint, figuring out which input pipeline yields the most reliable output is crucial because that dictates the entire workflow integration.

Lalam: I think what this paper really highlights is that tailoring a general language model to a specific domain, like pediatric emergency triage, using specialized data can significantly improve its performance on safety-critical tasks.

Tom: So we’re not just talking about making AI smarter in general; we're talking about making it trustworthy in specific medical contexts.

Jane: Precisely, Tom; the focus is on moving from inconsistent predictions to something more dependable so that clinicians have a better decision-support tool when they need it most.

Lu: I think this has huge potential for improving how we handle the massive influx of patient data in emergency settings because it addresses that variability directly.

Tom: And let's move into the conclusion of "Domain-Adapted Small Language Models for Reliable Clinical Triage" to really unpack what this all means for our practice and future thinking.

Meng: What’s the big picture implication here, Jane? Beyond just better accuracy, what does this paper suggest about how AI can fundamentally support patient care processes?

Jane: The authors are essentially arguing that these smaller models, when properly adapted to domain-specific clinical data like the CNH silver dataset, offer a viable path toward using AI for reliable decision support in triage.

Lalam: This moves the discussion beyond just proving that an AI can read text; it shows how domain adaptation allows us to create tools that are actually useful and safe for high-stakes environments.

Paper summary: Lu: I see this as a step where we move from theoretical models to something tangible that addresses real operational inefficiencies in emergency rooms right now.

Tom: So, when we look at the title and the authors of "Domain-Adapted Small Language Models for Reliable Clinical Triage," what’s the simple message we should be hearing about this work?

Jane: The core message is that using domain-specific fine-tuning techniques on smaller language models can lead to more consistent and accurate predictions for tasks like ESI assignment, which directly helps reduce errors in patient care pathways.

Meng: I think the practical implication is that we can start building decision support tools that are tailored to our specific clinical needs rather than relying on generic, less accurate systems.

Lalam: It suggests a future where AI assistants don't just offer suggestions but provide a level of consistency that makes them truly dependable partners in complex clinical workflows.

Tom: So, what is the final word from this paper on how we should view these domain-adapted SLMs in the broader context of emergency medicine?

Jane: The authors conclude that by focusing on domain adaptation and using methods like QLoRA fine-tuning with high-quality data, small language models can serve as reliable decision support tools for triage, though they also flag that evaluation relies on encounters with relatively complete documentation.

Lu: That caveat about the need for complete records is important because it sets realistic expectations about where these tools will work best in practice.

Tom: It’s a clear direction, then: we should be looking at how to build models specifically for our domain and use techniques that show strong performance balances between accuracy and efficiency, like what Qwen2 point 5-7B achieved here.

Meng: This paper gives us a solid foundation for exploring privacy-preserving decision support tools that can run closer to the point of care without needing massive external dependencies.

Lalam: Ultimately, this research points toward an AI ecosystem where specialized models can handle nuanced clinical tasks reliably, which means we can start thinking about how these tailored systems could improve patient flow and safety across the entire system.

Conclusion: Tom: So, we’ve been digging into the details of this paper about domain-adapted small language models for triage, and now we need to talk about what all this means in plain English and for the future.

Jane: Exactly, Tom; they’re looking at how tailoring these smaller models to specific clinical data helps make predictions much more consistent when dealing with messy patient information.

Lu: I think the real power here is seeing how these models can learn the specific language of a hospital, not just general English patterns, which opens up huge possibilities for specialized AI applications.

Meng: From my side, it’s interesting that they focused on small models; if we can get high accuracy with something more efficient, it makes deployment in real-world settings much more feasible.

Lalam: I see this as a big step forward because when an AI learns the specific culture and nuance of a medical field, its ability to provide support becomes much more trustworthy for the people using it every day.

Tom: It really boils down to how these domain-adapted models can help reduce human error in high-stakes situations by providing more reliable triage guidance.

Jane: That’s right; they are focusing on making sure the AI isn't just guessing, but actually reflecting established clinical reasoning patterns found in real patient records.

Lu: The implication for the broader field is that we might see a trend toward highly specialized models rather than just trying to make one massive model do everything poorly.

Meng: I’m thinking about how this could impact operational workflows; if triage decisions are more stable, it should lead to smoother patient flow and better resource allocation in emergency departments.

Lalam: For our culture, this work suggests that we can build trust in AI systems when they are demonstrably accurate within a specific domain, which is crucial for integrating new technologies into patient care routines.

Tom: So while the technical details are complex, the main point is that we’re building smarter tools for a very specific job instead of just bigger, vaguer general tools.

Jane: Precisely; it's about taking that complex idea and making it something practical and understandable for everyone involved in healthcare.

Lu: This opens up avenues for creating highly customized AI assistants that can act as experts in niche areas like pediatric triage or oncology assessments.

Meng: It makes me wonder how quickly we can translate this kind of targeted fine-tuning into tools that are actually integrated into a hospital's existing infrastructure.

Lalam: If we can reliably adapt models to different medical specialties, it could allow for personalized support systems tailored precisely to the needs of different patient populations.

Tom: That’s a lot to think about; we’ve seen the results show significant improvements in accuracy, and now we need to focus on how these concepts move from the paper into our daily reality.

Department of Computer Science, Virginia Tech · Children’s National Hospital

cs.CL, cs.AI, cs.LG

Submitted: 2026-04-29

Updated: 2026-10-01

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 89/100

The gist: Accurate and consistent Emergency Severity Index (ESI) assignment remains a persistent challenge in emergency departments, where highly variable free-text triage documentation contributes to

Key concepts

Emergency Severity Index (ESI)
The ESI is a standardized system used in emergency departments to quickly categorize patients based on the urgency of their medical needs. It helps triage patients efficiently so that those with the most critical conditions receive immediate attention, improving workflow.
Clinical Vignette Generation
This technique involves using an AI model to take long, unstructured patient notes and summarize them into a short, focused clinical scenario. This compact format allows the model to better understand the essential medical context needed for accurate triage decisions.
QLoRA Fine-Tuning
Quantized Low-Rank Adaptation (QLoRA) is a memory-efficient method used to adapt large language models like Qwen2.5-7B for specific tasks. It allows the model to learn new, domain-specific knowledge using less computational power and memory than full retraining.
Discordance Metrics
These metrics measure how often the model's predicted ESI level differs from the actual nurse-assigned ESI. Key measures include 'Undertriage' (predicting a lower severity than true) and 'Overtriage' (predicting a higher severity than true), helping quantify overall error rates.

Terminology

Summary

Accurate and consistent Emergency Severity Index (ESI) assignment remains a persistent challenge in emergency departments, where highly variable free-text triage documentation contributes to mistriage and workflow inefficiencies. The SLM, Qwen2.5-7B, demonstrated the strongest balance of accuracy, stability, and computational efficiency when fine-tuned on domain-specific pediatric triage data.

Data Preparation Model Prompting & Fine Tuning

The study utilized a dataset sourced from Children’s National Hospital (CNH) emergency department records, involving an initial cohort of approximately 117,600 pediatric encounters. To create the training material, two complementary supervision sources were employed: first, the ESI Handbook vignette dataset (n = 245) provided expert-curated scenarios; second, a large CNH silver dataset (n = 117,247) was generated by prompting Qwen2.5-7B to produce concise clinical vignettes from structured triage fields while retaining nurse–assigned ESI labels as ground truth. This process allowed the model to learn both guideline-based reasoning patterns and the variability of real-world clinical presentations.

Prompting & Input Pipelines

The researchers systematically compared multiple prompting pipelines to assess input format effects on ESI prediction accuracy. The six evaluated pipelines included:

  1. Raw Triage Record → ESI Prediction: Direct prompting with unstructured triage notes.

  2. Raw Triage Record → Clinical Vignette → ESI Prediction: Models first summarized triage notes into concise clinical vignettes, then predicted ESI.

  3. Human-Structured Data → ESI Prediction: Direct prompting with human-provided structured fields to generate ESI levels.

  4. Raw Triage Record → Structured Data → ESI Prediction: Models extracted structured fields (Chief Complaint, Vital Signs, Physical Exam) and predicted ESI from these structured outputs.

  5. Human-Structured Data → Clinical Vignette → ESI Prediction: Structured data were used to generate clinical vignettes before ESI prediction.

  6. LLM-Generated Structured Data → Clinical Vignette → ESI Prediction: Clinical vignettes were generated from LLM-extracted structured data and used for ESI prediction.

Model Selection and Fine-Tuning Strategy

The evaluation focused on open-source Small Language Models (SLMs), specifically comparing Qwen2.5-7B, Qwen2.5-14B, MedGemma variants, MedLLaMA2-7B, GPT-OSS models, and GPT-4o. The core fine-tuning involved applying Quantized Low-Rank Adaptation (QLoRA) to the Qwen2.5-7B model on the CNH silver dataset. This involved partitioning the silver dataset into 10 equal chunks (≈11.7k examples each) and proceeding sequentially, where the model was first trained on Chunk 1 and then continued from each checkpoint through Chunk 10. This approach enabled manageable GPU memory usage, stable training across diverse vignettes, and monitoring of model behavior over time.

Evaluation Metrics and Performance Analysis

Model performance was measured against nurse-assigned ESI levels using several metrics: Total Discordance (overall error rate), Undertriage (predicted label > true label), Overtriage (predicted label < true label), Significant Undertriage (true ESI 2 predicted as 3, 4, or 5), and Significant Overtriage (true ESI 3, 4, or 5 predicted as ESI 1 or 2). The results showed that clinical vignettes yielded the most accurate ESI predictions by preserving essential clinical context in a compact format, while Qwen2.5-7B achieved the best balance of accuracy, interpretability, and efficiency. Domain-specific fine-tuning on the CNH silver data resulted in discordance decreasing from 42.33% to 35.51% (for 10k examples), with low undertriage (13.35%) and significant overtriage (2.56%).

Explainability and Limitations

To interpret model behavior, token-level attention patterns were analyzed for correctly and incorrectly classified encounters across ESI-2 and ESI-3. Correct predictions showed strong attention to respiratory descriptors, pain-related symptoms, and age-specific context. In contrast, misclassified cases placed greater attention on nonspecific contextual or administrative cues, such as 'management,' 'process,' 'reported,' 'home,' and medications. Furthermore, the study found that multi-agent ensembles did not outperform the single-agent model. A key limitation noted is that evaluation required encounters with relatively complete documentation, and the use of nurse-assigned ESI levels as a reference standard means models may reproduce existing practice patterns rather than definitive true ESI. The study concluded that Qwen2.

Improvements for AI systems

Here are the specific improvements that can be made to existing AI systems based on this research, and what those improved systems could achieve:


  1. The deployment of an institution-specific, fine-tuned Small Language Model (SLM), specifically a 7B parameter model like Qwen2.5-7B, for Emergency Severity Index (ESI) prediction.

  2. The use of a structured data input pipeline to generate clinical vignettes before ESI prediction, as this yields the most accurate results compared to raw triage notes or unstructured data alone.

  3. The implementation of Quantized Low-Rank Adaptation (QLoRA) for memory-efficient, parameter-efficient fine-tuning of SLMs on domain-specific data without requiring substantial compute resources.

  4. The development of a robust, multi-stage prompting framework that systematically tests different input representations (Raw Triage Record vs. Clinical Vignette vs. Structured Data) to determine the optimal method for ESI prediction in real-world scenarios.

  5. The integration of token-level attention visualization techniques to provide qualitative explainability, allowing clinicians to see which specific clinical features (e.g., respiratory descriptors, pain symptoms) the model is prioritizing when making a triage decision.

  6. The creation of a system that incorporates Retrieval-Augmented Generation (RAG) specifically tailored with an ESI GuideBook knowledge base to mitigate the risk of hallucination and ground predictions in rule-based tasks.

  7. The development of specialized, safety-aligned multi-agent ensembles (e.g., safety-first, guideline-strict) to ensure consistency and reduce critical errors like significant undertriage or overtriage during the decision process (though current research suggests single agents are often sufficient).

These improved AI systems can achieve the following:

  1. A highly accurate and stable triage decision support system that reduces clinical discordance by moving from general models (like GPT-4o) to a specialized, fine-tuned SLM (Qwen2.5-7B).

  2. Real-time ESI assignment with sub-second inference latency, making it suitable for immediate use in busy emergency department workflows without disrupting clinical operations.

  3. Privacy-preserving deployment: The institution can fine-tune the model on local, de-identified data, ensuring patient information never leaves the secure institutional environment (on-premise deployment).

  4. Enhanced clinical trust: Clinicians gain insight into why a triage decision was made by examining attention maps that highlight relevant clinical signals versus distracting administrative noise.

  5. Improved robustness under incomplete data: The system can maintain high prediction accuracy even when key structured information (like vital signs) is missing, which is common in real-world documentation.

  6. Reduced risk of catastrophic errors: By using a fine-tuned model and potentially multi-agent voting mechanisms, the system minimizes clinically significant errors (significant undertriage/overtriage), directly improving patient safety by ensuring high-acuity cases are not misclassified as lower acuity.

Abstract

Accurate and consistent Emergency Severity Index (ESI) assignment remains a persistent challenge in emergency departments, where highly variable free-text triage documentation contributes to mistriage and workflow inefficiencies. This study evaluates whether open-source small language models (SLMs) can serve as reliable, privacy-preserving decision-support tools for clinical triage. We systematically compared multiple SLMs across diverse prompting pipelines and found that clinical vignettes, concise summaries of triage narratives, yielded the most accurate predictions. The SLM, Qwen2.5-7B, demonstrated the strongest balance of accuracy, stability, and computational efficiency. Through large-scale domain adaptation using expert-curated and silver-standard pediatric triage data, fine-tuned Qwen2.5-7B models substantially reduced discordance and clinically significant errors, outperforming all baseline SLMs and advanced proprietary large language models (LLMs, e.g., GPT-4o). These findings highlight the feasibility of institution-specific SLMs for reliable, privacy-preserving ESI decision support and underscore the importance of targeted fine-tuning over more complex inference strategies.

Sources

Related papers