Trajectory-Aware Clinical Risk Prediction via Severity-Grounded Knowledge Graphs and Retrieval-Augmented Generation

arXiv:2607.18270 · cs.AI, cs.CL · Submitted 2026-06-02 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Trajectory-Aware Clinical Risk Prediction via Severity-Grounded Knowledge Graphs and Retrieval-Augmented Generation".

Jane: Trajectory-Aware Clinical Risk Prediction via Severity-Grounded Knowledge Graphs and Retrieval-Augmented Generation proposes TRACER, a novel framework that enhances mortality and readmission prediction by integrating severity-grounded medical knowledge graphs,

Tom: First, who's behind it and why it matters.

Paper summary: Tom: We’ve touched on what TRACER is designed to do, focusing on how it uses severity scores and retrieval to improve clinical risk prediction. Now, let's get a clearer picture of the core thesis of this work by looking at the summary provided in the paper.

Jane: Right, Tom. The central claim of this paper is that current methods fail because they treat all medical concepts equally and don't account for disease severity or how a patient progresses across multiple visits. TRACER proposes solving this by building a knowledge graph where diagnoses get severity scores derived from literature, then using that graph to find patient-specific progression paths.

Lu: I see the explicit problem they are trying to solve: KDD ’twenty-six showed that ignoring the clinical importance of a diagnosis leads to missed critical red-flag events, and this paper is directly addressing that by measuring how likely an LLM is to predict one concept given another using patient-specific prompts.

Meng: So, the summary suggests they are building a structure where the knowledge itself has inherent clinical meaning—the severity scores—which then guides the retrieval of relevant patient journeys instead of just relying on raw visit data.

Lalam: That focus on clinical meaning embedded in the structure is really important for me; it suggests we can move towards an AI that understands medical context more deeply, not just statistical correlations between data points.

Tom: Exactly, Lalam! The mechanism involves constructing this SMKG enriched with severity information from literature and then using patient-specific prompts to navigate that graph for relevant paths across visits. It’s about building a pathway that reflects the actual clinical reality of a disease course.

Jane: And they also integrate textual evidence from clinical notes into this profile, essentially augmenting the patient's medical record with narrative context to make those trajectory predictions more informed by personal experience, as noted in their summary.

Lu: It’s interesting how they are combining the structured knowledge of the graph with the unstructured textual information from clinical notes to create this rich patient medical profile that drives the final assessment. That combination is where I see some really creative potential for future applications.

Meng: Practically, it means we're not just feeding an AI a list of diagnoses; we’re giving it a map of how those diagnoses typically interact and progress in real patients, which should lead to more reliable outputs when data is scarce.

Lalam: I think this focus on patient-specific progression paths derived from the severity-grounded graph is what gives this framework its real strength for handling those sparse data situations they mentioned.

Tom: So, to summarize this part: TRACER proposes a system that creates a severity-aware knowledge graph, finds patient progression paths within it, and uses clinical notes to enrich those paths for better risk prediction. That sets the stage for us to look at the broader implications next.

Jane: It’s certainly a framework designed to bridge the gap between raw EHR data and actionable, clinically grounded risk insights through this sophisticated combination of knowledge representation and retrieval techniques.

Conclusion: Tom: We’ve seen the framework outlined, so now we need to talk about what this whole endeavor really means for the field of clinical AI. Let's discuss the title, "Trajectory-Aware Clinical Risk Prediction via Severity-Grounded Knowledge Graphs and Retrieval-Augmented Generation," and who put this work together.

Jane: The authors are Kyunghoon Jeon, Youmin Ko, Woohwan Jung, and Hyunjoon Kim from Hanyang University Seoul. Their focus on combining knowledge graphs with retrieval augmentation is really key to their approach here.

Lu: The implications are huge because they show that we don't have to treat every piece of medical information equally; instead, we can structure the AI’s understanding around the actual severity of the condition. That structural change in how knowledge is organized has broad potential across all healthcare applications.

Meng: From an engineering perspective, this suggests that future clinical AI won't just be about training models on massive datasets; it will be about building systems with inherent, structured domain knowledge that guides the learning process more effectively.

Lalam: I think the cultural impact is profound because this work pushes us toward developing AI that is not just a statistical predictor but one that can provide reasoning steps based on documented clinical progression, which builds trust with medical professionals.

Tom: That trust aspect is vital, Jane. When an AI can show its work by tracing a specific path in the SMKG and citing relevant notes, it becomes a much more useful assistant rather than just a black box giving a number.

Jane: It shifts the goal from just getting high accuracy scores to achieving transparency in how that prediction was reached, which is something clinicians desperately need when making life-altering decisions based on AI input.

Lu: The paper’s success in demonstrating better Macro F1 scores and sensitivity improvements on datasets like MIMIC-III and MIMIC-IV validates the idea that capturing causal progression actually leads to more effective risk stratification in real-world clinical data.

Meng: So, if we translate those performance gains into a hospital setting, it means we could significantly reduce unnecessary interventions or improve early detection for high-risk patients by making the predictions much more reliable than what current baselines offer.

Lalam: I feel like the biggest impact here is in how it can enhance patient care by allowing for highly personalized risk assessments that go beyond standard demographic checks, tailoring the prediction to the specific severity and trajectory of that individual patient.

Tom: So we've seen how TRACER uses severity-grounded knowledge graphs and retrieval augmentation to tackle data sparsity and clinical narrative limitations, leading to performance gains in real datasets like MIMIC-III. That’s the big picture on this paper.

Hanyang University · Korea University

cs.AI, cs.CL

Submitted: 2026-06-02

Updated: 2026-06-02

Comments: Accepted to KDD 2026

Code: https://github.com/KyunghoonJeon/TRACER

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 90/100

The gist: Trajectory-Aware Clinical Risk Prediction via Severity-Grounded Knowledge Graphs and Retrieval-Augmented Generation proposes TRACER, a novel framework that enhances mortality and readmission

Key concepts

Severity-Weighted Medical Knowledge Graph (SMKG)
This graph assigns a severity score (1-20) to medical diagnoses based on curated literature. It integrates various medical sources and uses an LLM to generate these scores, ensuring the knowledge base reflects how severe a condition is clinically.
Trajectory Retrieval
A trajectory is defined as a sequence of medical concepts across visits in the SMKG. The system retrieves relevant paths using embeddings (BioClinicalBERT) based on specific risk or recovery queries, helping to map out a patient's disease progression over time.
Retrieval-Augmented Generation (RAG)
TRACER uses RAG to generate predictions by feeding the LLM three types of evidence: medical profiles, demographics, and key trajectories. This context repacking prevents 'lost-in-the-middle' issues, allowing the LLM to reason step-by-step before making a final prediction.

Terminology

Summary

Trajectory-Aware Clinical Risk Prediction via Severity-Grounded Knowledge Graphs and Retrieval-Augmented Generation proposes TRACER, a novel framework that enhances mortality and readmission prediction by integrating severity-grounded medical knowledge graphs, trajectory retrieval, and retrieval-augmented generation to provide clinically interpretable risk assessments. This method addresses limitations in existing models by explicitly accounting for disease severity, personal narrative evidence from clinical notes, and cohort-level experience to improve predictive performance under data sparsity.

Severity-Weighted Medical Knowledge Graph Construction

The framework constructs a novel medical knowledge graph called SMKG grounded on the severity of diagnoses by assigning clinically meaningful scores to diagnoses based on curated medical literature and clinical guidelines. This process involves:

  1. Constructing a unified medical KG integrating existing sources like UMLS, PubMed documents, and LLM-generated triples.

  2. For each diagnostic concept in the unified KG, a passage retriever searches external sources for top-l passages related to that diagnosis.

  3. An LLM generates a severity score (ranging from 1 to 20) based on the concatenated retrieved passages, where higher values indicate greater clinical severity.

Patient Medical Profile Retrieval

To address sparse visit histories, TRACER augments the target patient’s EHR with a Patient Medical Profile (PMP), which includes:

  1. Task-relevant progression of diseases and treatment responses.

  2. Task-relevant descriptions of health status.

  3. Demographics of the patient.

This profile retrieval involves two complementary factors:

(1) Trajectory Retrieval:

The framework defines a trajectory as a path in the SMKG that connects medical concepts across consecutive visits, following temporal order. It computes embeddings for candidate trajectories using BioClinicalBERT and retrieves paths based on risk queries (e.g., Clinical deterioration) and protective queries (e.g., Clinical recovery). The relevance scores are calculated using a formula:

(2) Trajectory Refinement:

Coarse first-pass retrieval is refined by reranking trajectories using Maximal Marginal Relevance (MMR) to balance query relevance and diversity. This involves iteratively selecting the trajectory with the highest gain score, defined as:

g(p i) = λMMR cos(e q, e p i) − (1 − λMMR) max j ∈ Si-1 cos(e p i, e j).

Clinical Note Retrieval and Context Augmentation

To leverage unstructured clinical notes, TRACER retrieves relevant passages from the set R of the patient’s historical clinical records. The passage pool B is constructed by splitting every clinical note in R into passages. For both risk and protective queries, embeddings are computed for the query and each passage, and cosine similarities are used to select the top-n passages most relevant to the query. These retrieved passages serve as note-level evidence in the patient medical profile.

Retrieval-Augmented Clinical Risk Prediction

The final prediction is generated by augmenting a single prompt with three types of supporting evidence: (A) risk and protective medical profiles, (B) demographics, and (C) the Key Supporting Trajectories (KSTs) and demographics of top-m similar patients. The framework employs dynamic context repacking, where KSTs and clinical notes are repacked in evidence (A) so that the most relevant information comes earlier in the prompt, mitigating the “lost-in-the-middle” problem. This integrated approach allows an LLM to generate a step-by-step reasoning process based on this comprehensive context before providing a binary prediction.

Key Contributions and Performance

The main contributions include:

(1) Severity-grounded medical knowledge graph (SMKG):

Assigning fine-grained and diagnosis-level severity scores derived from medical literature.

(2) Addressing data sparsity:

Augmenting sparse data with (a) patient-specific and severity-aware progression paths across visits from SMKG and (b) clinically similar patient cases, both of which enables the LLM reasoning based on nuanced patient status changes under data sparsity.

(3) Leveraging textual evidence:

Using salient and nuanced textual evidence in clinical notes, enabling robust modeling of red-flag events and personal experiences.

The experiments on MIMIC-III and MIMIC-IV demonstrate that TRACER consistently outperforms state-of-the-art baselines, showing significant gains in sensitivity (up to 84.1% improvement on MIMIC-III) and Macro F1 scores, validating the hypothesis that capturing causal progression and amplifying high-risk signals, rather than flattening medical history, leads to more effective risk stratification in realworld clinical data. The framework is also shown to be robust across different visit sequence lengths and LLM backbones.

Hyperparameter Sensitivity

A comprehensive sensitivity analysis indicates optimal performance for several key parameters:

Improvements for AI systems

As a fastidious and diligent researcher, I have analyzed the TRACER framework described in this paper. The core strength of TRACER lies in its integrated, multi-modal retrieval-augmented generation (RAG) approach that explicitly incorporates disease severity into trajectory modeling.

Here are the specific improvements to AI systems based on this research:


Improvement 1: Implement Severity-Grounding for Knowledge Graph Construction (SMKG)

Instead of using a flat knowledge graph where all medical concepts are treated equally, the system should construct a Severity-Weighted Medical Knowledge Graph (SMKG). This is achieved by using an LLM to assign fine-grained severity scores (1–20) to diagnosis nodes based on curated medical literature and clinical guidelines.

Improving AI Systems:

The resulting system can perform significantly better in high-stakes clinical decision support tasks where distinguishing between clinically critical conditions (e.g., sepsis vs. mild hypertension) is vital. It moves beyond simple concept presence to understanding the inherent risk of a diagnosis itself, which is crucial for accurate triage and risk stratification.

Improvement 2: Adopt Trajectory Retrieval with Risk/Protective Query Separation

The system must retrieve patient trajectories not just based on semantic similarity, but by querying two distinct types of paths: a Risk Path query (seeking deterioration signals like 'progresses to', 'exacerbates') and a Protective Path query (seeking recovery signals like 'resolves', 'stabilizes'). Furthermore, the retrieved trajectories should be filtered using Maximal Marginal Relevance (MMR) to ensure diversity and relevance.

Improving AI Systems:

This enables superior prediction robustness. The system can distinguish between transient fluctuations (noise) and sustained, clinically meaningful disease progression or recovery trends. This leads to more reliable predictions for long-term outcomes, as it actively seeks evidence of both worsening conditions and successful interventions, rather than relying on a single, potentially misleading historical sequence.

Improvement 3: Integrate Unstructured Clinical Note Evidence via Targeted Passage Retrieval

The system should retrieve relevant passages from unstructured clinical notes using embeddings based on specific risk/protective queries for each prediction task (mortality vs. readmission). These passages must then be repacked into the LLM prompt, ensuring that the most salient evidence appears first to mitigate lost-in-the-middle issues in long context windows.

Improving AI Systems:

This enhances interpretability and clinical trust. The system can generate explanations for its predictions by citing specific phrases from physician notes (e.g., The note mentioning uncontrolled infection led to the high risk prediction). This is essential for clinicians to validate the AI's reasoning, moving the model from a black-box predictor to an explainable decision support tool.

Improvement 4: Augment Patient Context with Similar Peer Cases Based on Recent Visit Patterns

Instead of using generic patient embeddings or static summaries, the system should retrieve top-m similar patients based on recent visit patterns (e.g., using Last-Visit Jaccard similarity). These peer cases should be represented by concise trajectory-based summaries to provide context for the LLM.

Improving AI Systems:

This allows the AI to perform robust case-based reasoning. When a patient's history is sparse, the system can leverage observed outcome patterns from historically similar cohorts to make more informed predictions. This significantly boosts performance in data-sparse scenarios common in real-world EHR data, ensuring that predictions are grounded not just in the individual record but also in population dynamics.

Improvement 5: Dynamic Context Repacking for LLM Input

The final prompt construction must employ a dynamic repacking strategy where the most relevant retrieved components (KSTs and clinical notes) are placed closest to the core task description, optimizing the LLM's attention mechanism.

Improving AI Systems:

This directly addresses the limitations of long-context LLMs. By prioritizing evidence, the system ensures that even with massive amounts of input data (trajectories, notes, peer context), the model focuses its reasoning on the most critical signals for making a conservative and clinically grounded prediction, minimizing noise and maximizing predictive accuracy under real-world constraints.


In summary, this improved AI system transforms from a simple sequence predictor into a sophisticated clinical reasoning engine capable of:

  1. Evaluating the inherent risk level of diagnoses (Severity Grounding).

  2. Identifying causal disease progression (Trajectory Retrieval).

  3. Synthesizing nuanced patient experiences (Clinical Note Retrieval).

  4. Leveraging population-level context for sparse data handling (Similar Patient Augmentation).

The resulting AI system will be highly accurate, robust to data sparsity, and clinically interpretable by providing traceable evidence for every prediction.

Abstract

While Electronic Health Records (EHRs) offer a wealth of clinical data, effectively augmenting a patient's records with heterogeneous external knowledge to predict the patient's clinical risk remains a significant challenge. Existing methods fail to capture disease severity, treatment responses, and nuanced clinical progression, due to data sparsity and the underutilization of unstructured clinical notes. To address these challenges, we propose TRACER (a trajectory-aware and clinically grounded prediction framework) that (1) constructs a medical knowledge graph enriched with severity information from medical literature, (2) retrieves clinically relevant, severity-weighted paths of a patient's progression from the knowledge graph, (3) extracts clinically relevant events from unstructured clinical notes, and (4) augments patient context with similar peer cases. Experiments on the MIMIC-III and MIMIC-IV datasets demonstrate large gains over state-of-the-art baselines, with up to 28.5% increase in Macro F1 score for the mortality prediction task, and 19.7% increase for the readmission prediction task.

Sources

Related papers