ELICITED: EHR-grounded Longitudinal Interactive Conversations for Information-seeking Triage Evaluation and Decision-making
Haohao Zhu, Xiaolin Shi, Jiayu Zhou
University of Michigan · Ellipsis Health
cs.CL
Submitted: 2026-08-10
Updated: 2026-08-11
Code: https://github.com/illidanlab/elicited
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: ELICITED: EHR-GROUNDED LONGITUDINAL INTERACTIVE CONVERSATIONS FOR INFORMATION SEEKING TRIAGE EVALUATION AND DECISION MAKING introduces EHR2Dial-Triage, an agentic conversation-generation framework
Terminology
Summary
ELICITED: EHR-GROUNDED LONGITUDINAL INTERACTIVE CONVERSATIONS FOR INFORMATION SEEKING TRIAGE EVALUATION AND DECISION MAKING introduces EHR2Dial-Triage, an agentic conversation-generation framework and benchmark grounded in MIMIC-IV-ED for studying multi-turn emergency-department (ED) triage. The paper states: We introduce EHR2Dial-Triage, an EHR-to-dialogue framework and benchmark for studying multi-turn ED triage under explicit role-based and temporal information boundaries.
The framework constructs triage conversations under explicit role-based and temporal information boundaries. Each ED encounter is divided into information available at the beginning of triage, patient-reportable information that may be elicited during the conversation, and future or evaluation-only information that remains hidden.
A clinician model asks focused questions using currently available evidence and dialogue history, while a patient model responds using eligible case information. A verifier checks each proposed disclosure and records its supporting EHR event and first accepted dialogue turn.
The paper explains the motivation: "Emergency-department (ED) triage requires clinicians to rapidly identify patients who need immediate attention, determine who can safely wait, and prioritize limited clinical resources. At presentation, however, the available information is often incomplete and may be limited to a brief chief complaint and initial vital signs. It further notes that
Most existing ED benchmarks evaluate acuity prediction from a fixed clinical snapshot... it does not fully capture the interactive process through which triage-relevant evidence is elicited and interpreted."
The benchmark contains 4,041 dialogues with a median of eight clinician–patient rounds per dialogue. The paper reports: "The benchmark contains 4,041 dialogues, with a median of eight clinician–patient rounds per dialogue. Most dialogues contain between five and thirteen rounds, providing sufficient space for multi-turn information elicitation while maintaining a bounded interaction length. It also notes
EHR2Dial-Triage contains an average of 17.3 speaker turns per dialogue, with clinician and patient utterances averaging 17.1 and 12.8 words, respectively."
The framework supports controlled evaluation of information elicitation, evidence use, five-level Emergency Severity Index (ESI) prediction, and patient communication across models and patient personas. The paper states: "EHR2Dial-Triage supports controlled evaluation of information elicitation, evidence use, five-level Emergency Severity Index prediction, and patient communication across models, and patient personas. Together, these capabilities provide a structured setting for studying conversational triage as a dynamic process of clinical information acquisition, reasoning, and communication."
The paper describes two complementary experiments. Experiment I evaluates LLMs as triage clinicians: Experiment I evaluates LLMs as triage clinicians by examining what questions they ask, what source-supported patient information they elicit, and how the acquired evidence changes a fixed ESI reader's assessment.
Experiment II evaluates whether models can interpret a completed triage dialogue: Experiment II evaluates whether models can predict the recorded five-level Emergency Severity Index (ESI), and generate an appropriate patient-facing closing.
Key findings from Experiment I show that models behave similarly early but diverge later: At Round 2, FactCov ranges only from 18.9% to 19.8%, and coverage remains tightly clustered at Round 6 52.7%–53.9%... By Round 10, FactCov ranges from 58.1% to 68.0%, and HistoryMedCov from 43.1% to 56.5%.
The paper notes: GLM-4.7-Flash reaches the highest terminal coverage, followed closely by MedGemma-1.5-4B and gpt-oss-20B.
In Experiment II, the paper reports: Claude Fable 5 performs best overall on ESI prediction, with the highest Macro-AUC, Macro-F1, and QWK.
For patient-facing generation, GPT-5.6 performs best on the grounding-oriented generation metrics, with the highest Key Recall and Grounded Precision.
The paper also includes a safety analysis on under-triage: Under-triage is clinically important because assigning a patient to a lower urgency level may delay evaluation or treatment.
It finds: "For most models with enough severe cases for comparison, confidence on these errors is similar to or higher than confidence on correct predictions. GPT-5.6 and Claude Fable 5 show the opposite pattern, with lower confidence when severe under-triage occurs. The paper concludes:
These results raise an important safety concern. A model can make a large triage error without showing a corresponding drop in confidence, making such failures difficult to detect from reported scores alone."
The paper concludes: "Across 4,041 conversations, our experiments show that information acquisition, ESI prediction, evidence grounding, and patient-facing communication capture distinct dimensions of model performance. Models that perform strongly on final acuity prediction do not necessarily achieve the strongest information acquisition or grounded communication, highlighting the value of evaluating conversational triage as a multi-stage clinical process."
Improvements for AI systems
Improvements to AI Systems:
-
Dynamic Information-Elicitation Planning: Implement a clinician model that explicitly tracks a
coverage state
of known versus missing triage-relevant facts (e.g., HistoryMed, HistorySurg, Meds) and selects its next question to maximize marginal information gain, rather than relying on static question templates. This improves early-stage efficiency (targeting >20% FactCov by Round 2) and late-stage completeness (targeting >68% FactCov by Round 10). -
Temporal-Boundary-Aware Reasoning: Add a hard constraint layer to the model's generation process that prevents it from using or referencing future or evaluation-only EHR events (e.g., final diagnosis, disposition) during the triage dialogue. The system will maintain a separate
hidden information buffer
and a verifier module that blocks any disclosure or inference not yet supported by the current turn's eligible evidence, reducing hallucinated clinical reasoning. -
Confidence-Calibrated Triage with Under-Triage Detection: Enhance the ESI prediction head to output a calibrated confidence score that is explicitly penalized when high-confidence predictions coincide with severe under-triage (e.g., predicting ESI 4/5 when true ESI is 1/2). Train a secondary
risk-of-error
classifier on dialogue features (e.g., low coverage, unanswered critical questions) to flag cases where the model should defer or escalate, addressing the observed failure mode where confidence does not drop on large errors. -
Multi-Stage Evaluation and Reward Shaping: Modify the training objective to use a composite reward that weights four distinct metrics—information coverage (FactCov), evidence grounding (Grounded Precision), ESI prediction accuracy (Macro-F1), and patient-facing communication quality (Key Recall)—rather than optimizing only final acuity. This prevents models from
gaming
one stage (e.g., predicting well without asking good questions) and encourages balanced performance across the full triage process. -
Patient-Persona-Adaptive Communication: Implement a patient model that adjusts its response verbosity and disclosure style based on the persona (e.g., elderly, anxious, low-health-literacy) and the clinician's question phrasing, using a persona-conditional generation head. This improves realism in simulated conversations and allows the clinician model to practice tailoring questions and closing statements to diverse patient populations.
What the Improved AI System Can Do:
-
Conduct a full multi-turn ED triage conversation where it asks the most informative next question given only currently available EHR data, achieving higher terminal fact coverage (target >70%) while never leaking future clinical events.
-
Predict the 5-level ESI with calibrated confidence, and automatically flag cases where it is at high risk of under-triage, triggering a human-in-the-loop review or escalation protocol.
-
Generate a patient-facing closing statement that is both grounded in the elicited evidence (high Key Recall) and empathetic, adapting to the patient's persona and emotional state.
-
Be evaluated and improved simultaneously on four distinct clinical competencies—information acquisition, evidence use, acuity prediction, and communication—rather than a single aggregate score, enabling targeted fine-tuning of weak stages.
-
Serve as a safe, realistic training simulator for medical students and residents, where they can practice triage questioning and receive feedback on coverage gaps, evidence grounding, and communication quality without patient risk.
Abstract
Emergency-department (ED) triage requires clinicians to rapidly identify patients who need immediate attention, determine who can safely wait, and prioritize limited clinical resources. At presentation, however, information may be limited to a chief complaint and initial vital signs. Clinically important details, including symptom onset and progression, associated symptoms, medical history, and medication use, are often obtained through focused conversation. Effective triage therefore requires clinicians to identify information gaps, ask appropriate follow-up questions, and update their assessment as new evidence becomes available. Most existing ED benchmarks evaluate acuity prediction from a fixed clinical snapshot. Although this formulation measures predictive performance after patient information has been assembled, it does not capture the interactive process through which triage-relevant evidence is elicited and interpreted. Existing medical dialogue datasets support the study of clinical communication, but dialogue statements are not always linked to temporally ordered events in the electronic health record (EHR). We introduce EHR2Dial-Triage, an agentic conversation-generation framework and benchmark grounded in MIMIC-IV-ED. The framework constructs triage conversations under explicit role-based and temporal information boundaries. Each accepted patient disclosure is linked to its supporting EHR event and the first dialogue turn at which it becomes available. EHR2Dial-Triage enables controlled evaluation of information elicitation, evidence use, five-level Emergency Severity Index prediction, and patient-facing communication across models and patient personas. It provides a structured setting for studying conversational triage as a dynamic process of clinical information acquisition, reasoning, and communication.
Sources
- gpt-oss-120b & gpt-oss-20b Model Card
- Synthetic Patient-Physician Dialogue Generation from Clinical Notes Using LLM
- GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
- MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents
- MeDxAgent: Multi-Agent Consultation for Interactive Medical Diagnosis
- MedGemma 1.5 Technical Report
- TriageSim: A Conversational Emergency Triage Simulation Framework from Structured Electronic Health Records
- Qwen3 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering