Hierarchical Modeling of ICD Codes in EHR Foundation Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Hierarchical Modeling of ICD Codes in EHR Foundation Models".
Jane: The paper was written by Megha Thukral, Dong Gyun Kang, Rudra Pratap Singh, Shruthi Kashinath Hiremath, Katrin Hänsel et al. from Georgia Institute of Technology and Optum AI.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: We are starting today with a really significant paper called Hierarchical Modeling of ICD Codes in EHR Foundation Models.
Jane: This research looks at how we can make medical AI much more intuitive, Tom.
Tom: The authors from Georgia Tech and Optum AI, including Megha Thukral and Thomas Plötz, realized that current models treat medical codes as if they were just random, flat symbols.
Jane: They're essentially saying these codes aren't just arbitrary identifiers with no connection to each other.
Tom: Exactly, because ICD codes are actually organized into families and subgroups.
Jane: It is like a computer seeing the word "apple" and the word "apricot" as completely unrelated objects without any shared category.
Lu: That is such a massive missed opportunity for learning!
Lu: If we give the AI a family tree for every medical term, it can understand that these things belong to broader clinical groups.
Meng: I wonder if this actually changes how the software has to be built in a hospital setting.
Jane: The authors suggest they can integrate this structure into existing transformer models without needing to overhaul everything.
Meng: That sounds much more practical than asking every hospital to scrap their current databases and start over.
Lalam: If we get this right, the AI will finally start to mirror the actual logic that clinicians use when they are diagnosing patients.
Tom: It moves us from simple pattern matching toward actual clinical understanding.
Jane: We should look at exactly how they implemented this "family tree" logic in their two different methods.
Summary: Tom: Moving into the specifics, Hierarchical Modeling of ICD Codes in EHR Foundation Models explains two distinct ways they tried to inject this hierarchy into the models.
Jane: The first method is called HICD-BERT, and it's a pretty clever way to use what's already in the code strings.
Tom: They basically take a code and break it down into its prefix parts, like taking "S72" and extracting "S" and "S7".
Jane: It’s like adding little sticky notes to every word that tell the model which family that word belongs to.
Lu: That's a brilliant way to use data-driven prefixes without needing a massive external medical dictionary!
Tom: It's definitely efficient, but their second method, HICD-Graph, takes things much further.
Jane: This one builds a giant web where diseases are connected based on how often they appear in the same patient.
Meng: Building that kind of web sounds like it could get incredibly messy from an engineering standpoint.
Tom: They actually used something called PMI to make sure they only kept the connections that were statistically meaningful.
Meng: So you aren't just connecting everything to everything else, which would crash the system?
Lu: And then they layer the official medical ontology on top of that web to create a hybrid structure!
Lalam: This allows the model to see both how doctors actually behave and how medical science is organized.
Jane: It's a perfect blend of real-world data and established medical knowledge.
Tom: We need to see if all this extra complexity actually results in better predictions for patients.
Improvements: Tom: Now we have to talk about whether this hierarchical approach actually works, because the results for Hierarchical Modeling of ICD Codes in EHR Foundation Models are quite striking.
Jane: They found that adding this structure improved performance in twenty-six out of twenty-eight different comparisons they ran!
Tom: That is a huge win across both the BERT-style and the graph-style models.
Jane: It wasn't just about getting higher scores on one task, either.
Lu: I was particularly impressed by how well the graph model transferred from one dataset to another!
Tom: Right, they tested it by training on the MIMIC-IV dataset and then seeing if it could work on the eICU database without changing anything.
Jane: The graph approach was much more robust during that transfer than the BERT approach was.
Meng: That makes sense because medical hierarchies are universal, while specific hospital coding habits can change.
Meng: If a model understands the underlying biology, it shouldn't care which hospital it's working in.
Lalam: This kind of stability is what we need to build tools that clinicians can actually trust in different parts of the world.
Tom: They even analyzed the actual embeddings and found that hierarchy makes the code clusters much more coherent.
Jane: It pulls related diagnoses closer together in the model's "mind," making everything much more organized.
Lu: It's like turning a pile of loose papers into a perfectly indexed library!
Meng: And if that organization helps with accuracy, then it's definitely worth the extra computation.
Tom: We should wrap this up and see what the big picture is here.
Conclusion: Tom: We've really covered a lot of ground today regarding Hierarchical Modeling of ICD Codes in EHR Foundation Models.
Jane: It's been so eye-opening to see how much we can learn just by respecting the structure that's already there.
Tom: They proved that those medical codes are far more than just arbitrary labels, and they did it with a very lightweight approach.
Lu: I can already see this being applied to even more complex systems, like mapping out every single protein interaction in the human body!
Meng: As long as we keep these implementations scalable for actual hospital production, this will be a major standard.
Lalam: It's a beautiful step toward technology that truly understands the profound nuances of human health.
Jane: Thanks for joining us on the show today!
Tom: We'll see you next time!
Georgia Institute of Technology · Optum AI
cs.AI
Submitted: 2026-06-13
Updated: 2026-06-13
Journal ref: Proceedings of the 11th Machine Learning for Healthcare Conference, PMLR 340:1998-2029, 2026
Code: https://github.com/yandexdataschool/roc_comparison
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 80/100
The gist: This paper investigates the use of ICD-10-CM hierarchy as a "general inductive bias for clinical representation learning" within Electronic Health Record (EHR) foundation models.
Key concepts
- Hierarchical Modeling
- Instead of treating medical codes as random symbols, hierarchical modeling organizes them into families and subgroups. This allows AI to understand that specific codes belong to broader clinical groups, helping the software mirror the actual logic and categorization used by clinicians when diagnosing patients.
- HICD-BERT
- This method integrates hierarchy by breaking ICD codes down into their prefix parts, such as extracting "S" and "S7" from "S72." This efficient approach uses data-driven prefixes to act like labels, informing the model which medical family a specific code belongs to.
- HICD-Graph
- This method builds a web of diseases connected by how often they appear in the same patient, using PMI to ensure connections are statistically meaningful. By layering official medical ontology on top, it creates a hybrid structure combining real-world clinical behavior with established medical knowledge.
Terminology
Summary
This paper investigates the use of ICD-10-CM hierarchy as a general inductive bias for clinical representation learning
within Electronic Health Record (EHR) foundation models. By moving beyond treating diagnosis codes as flat tokens,
the authors aim to bridge the gap between existing model structures and the clinically meaningful hierarchical structure
inherent in medical ontologies, which can improve prediction accuracy and robustness under domain shift.
The Problem of Flat Representations
Existing EHR models typically treat structured clinical codes as flat symbols,
ignoring the fact that ICD codes are not arbitrary identifiers
but organize diseases into clinically meaningful families, subgroups, and increasingly specific categories.
This oversight is particularly problematic because many diagnoses are sparse, long-tailed, and institution-specific.
The authors hypothesize that incorporating hierarchy can provide several benefits:
-
A natural inductive bias where rare codes benefit from data associated with
clinically similar codes.
-
The ability to capture
care pathways or anatomical systems
that may be more predictive than a specific code alone. -
Improved robustness under domain shift, as higher-level disease families are often
more stable across institutions.
Proposed Hierarchical Mechanisms
The researchers propose two complementary strategies for injecting ICD hierarchy into contemporary EHR representation learning:
-
HICD-BERT (Hierarchical ICD-BERT): This token-level approach injects hierarchy directly into token representations in a BERT-style encoder. It embeds each level of the ICD hierarchy as an
additive token embedding alongside the diagnosis code
usingstring-prefix truncation.
-
HICD-Graph (Hierarchical ICD-Graph): This relational approach encodes hierarchy via a
diagnosis co-occurrence graph augmented with ontology-derived edges.
It uses a graph convolutional network (GCN) to learnhierarchy-aware code embeddings
that subsequently initialize a patient-level Transformer.
Experimental Evaluation and Findings
The models were evaluated on two large-scale real-world clinical datasets, MIMIC-IV and eICU, across multiple prediction tasks. The study conducted a systematic ablation over three granularity levels (G 0, G 1, and G 2). The results demonstrated that:
-
Explicitly encoding ICD hierarchy improves performance in
26 of 28 comparisons
across both architectures and tasks. -
The most useful level of hierarchy is architecture-dependent: HICD-BERT benefits most from the
finest granularity level,
while HICD-Graph benefits progressively as more levels are added. -
Hierarchy improves cross-dataset transfer, specifically for the graph-based approach, which
transfers robustly from MIMIC-IV to eICU.
Embedding Geometry and Clinical Semantics
Through embedding analysis of the learned HICD-Graph representations, the authors show that hierarchy reshapes the learned embedding geometry
to better reflect clinical semantics. Specifically, hierarchy produces more coherent code clusters,
and configurations with the tightest clusters
also achieve the strongest downstream performance. The study found that:
-
Hierarchy strengthens
clinically meaningful similarity structure,
increasing cosine similarity among codes sharing ancestors. -
The magnitude of improvement is heterogeneous across chapters, with larger proportional gains in chapters characterized by
greater code diversity.
Improvements for AI systems
Improvement 1: Implementation of a Hybrid Multi-Granularity Embedding Layer
Integrate an additive hierarchical embedding mechanism into Transformer-based EHR encoders (e.g., BEHRT or Med-BERT). Instead of representing ICD codes as flat, atomic tokens, the embedding layer should compute the sum of the base diagnosis token embedding, temporal/positional embeddings, and three distinct hierarchy-specific embeddings (G 0: Chapter level; G 1: Block level; G 2: Category level) derived from string-prefix truncation or official ontology mappings.
- Improved System Capability: The system will significantly mitigate the
long-tail
problem of rare diagnoses. By inheriting semantic information from parent categories, the model can make accurate clinical predictions even for infrequent codes that lack sufficient individual occurrence data in a specific training corpus.
Improvement 2: Ontology-Augmented Graph-Relational Pretraining
Shift from purely data-driven co-occurrence embeddings to a hybrid graph construction strategy for initializing Transformer encoders. Construct a diagnosis graph that combines empirical, Pointwise Mutual Information (PMI)-weighted co-occurrence edges with uniform-weight ontological hierarchy edges. Use a Graph Convolutional Network (GCN) to perform message passing across these hybrid edges before using the resulting node embeddings to initialize the patient-level Transformer.
- Improved System Capability: The system will exhibit superior cross-dataset transferability and robustness under domain shift. Because the model relies on stable, ontology-anchored relational structures rather than institution-specific co-occurrence patterns, it can be deployed in new hospital environments (e.g., moving from MIMIC to eICU) with minimal loss in predictive accuracy for tasks like ICU readmission.
Improvement 3: Hierarchical Semantic Constraint on Latent Embedding Geometry
Apply hierarchical supervision during the pretraining phase to force the latent representation space to respect clinical taxonomy. This involves optimizing the model so that the cosine similarity between diagnosis embeddings is constrained by their shared ancestors in the ICD-10-CM hierarchy.
- Improved System Capability: The system will produce clinically coherent and interpretable embedding clusters. In downstream tasks, such as disease phenotyping or patient stratification, the model will naturally group related diseases (e.g., different types of heart failure) together in the vector space, ensuring that the learned representations reflect medical reality rather than mere statistical noise.
Abstract
Electronic health record foundation models typically treat ICD diagnosis codes as flat tokens, overlooking the clinically meaningful hierarchical structure that captures disease families, subcategories, and fine-grained diagnostic detail. As a result, existing EHR representation learning methods do not explicitly exploit the hierarchical structure already present in the coding system. In this work, we study ICD-10-CM hierarchy as a general inductive bias for clinical representation learning. We investigate two complementary mechanisms for incorporating hierarchy: first, by augmenting diagnosis sequences in a BERT-style transformer with tokens corresponding to different levels of the ICD hierarchy, and second, by injecting hierarchy into graph-based code representations through hierarchy-aware edges combined with diagnosis co-occurrence structure. Across these settings, we evaluate whether explicit hierarchy improves downstream prediction, which levels of the hierarchy are most useful, whether hierarchy encoding improves transfer across datasets, and how hierarchy reshapes embedding similarity structure. We conduct experiments on two large-scale real-world clinical datasets: MIMIC-IV, used for pretraining and in-domain evaluation, and eICU, used to assess cross-dataset transfer via frozen encoder probing. Our findings show that explicitly encoding ICD hierarchy improves over flat code representations in both in-domain and cross-dataset settings, while revealing that the most useful level of hierarchy depends on both the task and the modeling approach. More broadly, we focus on hierarchy-aware EHR representation learning and show that the benefits of encoding hierarchy are generalizable across modeling settings and hierarchy levels.
Sources
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection