Hierarchical Modeling of ICD Codes in EHR Foundation Models

summary

Video file (mp4)

The gist

This paper investigates the use of ICD-10-CM hierarchy as a "general inductive bias for clinical representation learning" within Electronic Health Record (EHR) foundation models.

In short

Researchers from Georgia Tech and Optum AI developed methods to integrate medical hierarchies into EHR foundation models. By treating ICD codes as organized families rather than random symbols, methods like HICD-BERT and HICD-Graph improved prediction accuracy and model robustness, allowing AI to better mirror clinical logic.

Key concepts

Hierarchical Modeling
Instead of treating medical codes as random symbols, hierarchical modeling organizes them into families and subgroups. This allows AI to understand that specific codes belong to broader clinical groups, helping the software mirror the actual logic and categorization used by clinicians when diagnosing patients.
HICD-BERT
This method integrates hierarchy by breaking ICD codes down into their prefix parts, such as extracting "S" and "S7" from "S72." This efficient approach uses data-driven prefixes to act like labels, informing the model which medical family a specific code belongs to.
HICD-Graph
This method builds a web of diseases connected by how often they appear in the same patient, using PMI to ensure connections are statistically meaningful. By layering official medical ontology on top, it creates a hybrid structure combining real-world clinical behavior with established medical knowledge.

Terminology used across episodes

This episode discusses

The paper

Hierarchical Modeling of ICD Codes in EHR Foundation Models · Read on arXiv

Georgia Institute of Technology · Optum AI

Electronic health record foundation models typically treat ICD diagnosis codes as flat tokens, overlooking the clinically meaningful hierarchical structure that captures disease families, subcategories, and fine-grained diagnostic detail. As a result, existing EHR representation learning methods do not explicitly exploit the hierarchical structure already present in the coding system. In this work, we study ICD-10-CM hierarchy as a general inductive bias for clinical representation learning. We investigate two complementary mechanisms for incorporating hierarchy: first, by augmenting diagnosis sequences in a BERT-style transformer with tokens corresponding to different levels of the ICD hierarchy, and second, by injecting hierarchy into graph-based code representations through hierarchy-aware edges combined with diagnosis co-occurrence structure. Across these settings, we evaluate whether explicit hierarchy improves downstream prediction, which levels of the hierarchy are most useful, whether hierarchy encoding improves transfer across datasets, and how hierarchy reshapes embedding similarity structure. We conduct experiments on two large-scale real-world clinical datasets: MIMIC-IV, used for pretraining and in-domain evaluation, and eICU, used to assess cross-dataset transfer via frozen encoder probing. Our findings show that explicitly encoding ICD hierarchy improves over flat code representations in both in-domain and cross-dataset settings, while revealing that the most useful level of hierarchy depends on both the task and the modeling approach. More broadly, we focus on hierarchy-aware EHR representation learning and show that the benefits of encoding hierarchy are generalizable across modeling settings and hierarchy levels.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Hierarchical Modeling of ICD Codes in EHR Foundation Models".

Jane: The paper was written by Megha Thukral, Dong Gyun Kang, Rudra Pratap Singh, Shruthi Kashinath Hiremath, Katrin Hänsel et al. from Georgia Institute of Technology and Optum AI.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: We are starting today with a really significant paper called Hierarchical Modeling of ICD Codes in EHR Foundation Models.

Jane: This research looks at how we can make medical AI much more intuitive, Tom.

Tom: The authors from Georgia Tech and Optum AI, including Megha Thukral and Thomas Plötz, realized that current models treat medical codes as if they were just random, flat symbols.

Jane: They're essentially saying these codes aren't just arbitrary identifiers with no connection to each other.

Tom: Exactly, because ICD codes are actually organized into families and subgroups.

Jane: It is like a computer seeing the word "apple" and the word "apricot" as completely unrelated objects without any shared category.

Lu: That is such a massive missed opportunity for learning!

Lu: If we give the AI a family tree for every medical term, it can understand that these things belong to broader clinical groups.

Meng: I wonder if this actually changes how the software has to be built in a hospital setting.

Jane: The authors suggest they can integrate this structure into existing transformer models without needing to overhaul everything.

Meng: That sounds much more practical than asking every hospital to scrap their current databases and start over.

Lalam: If we get this right, the AI will finally start to mirror the actual logic that clinicians use when they are diagnosing patients.

Tom: It moves us from simple pattern matching toward actual clinical understanding.

Jane: We should look at exactly how they implemented this "family tree" logic in their two different methods.

Summary: Tom: Moving into the specifics, Hierarchical Modeling of ICD Codes in EHR Foundation Models explains two distinct ways they tried to inject this hierarchy into the models.

Jane: The first method is called HICD-BERT, and it's a pretty clever way to use what's already in the code strings.

Tom: They basically take a code and break it down into its prefix parts, like taking "S72" and extracting "S" and "S7".

Jane: It’s like adding little sticky notes to every word that tell the model which family that word belongs to.

Lu: That's a brilliant way to use data-driven prefixes without needing a massive external medical dictionary!

Tom: It's definitely efficient, but their second method, HICD-Graph, takes things much further.

Jane: This one builds a giant web where diseases are connected based on how often they appear in the same patient.

Meng: Building that kind of web sounds like it could get incredibly messy from an engineering standpoint.

Tom: They actually used something called PMI to make sure they only kept the connections that were statistically meaningful.

Meng: So you aren't just connecting everything to everything else, which would crash the system?

Lu: And then they layer the official medical ontology on top of that web to create a hybrid structure!

Lalam: This allows the model to see both how doctors actually behave and how medical science is organized.

Jane: It's a perfect blend of real-world data and established medical knowledge.

Tom: We need to see if all this extra complexity actually results in better predictions for patients.

Improvements: Tom: Now we have to talk about whether this hierarchical approach actually works, because the results for Hierarchical Modeling of ICD Codes in EHR Foundation Models are quite striking.

Jane: They found that adding this structure improved performance in twenty-six out of twenty-eight different comparisons they ran!

Tom: That is a huge win across both the BERT-style and the graph-style models.

Jane: It wasn't just about getting higher scores on one task, either.

Lu: I was particularly impressed by how well the graph model transferred from one dataset to another!

Tom: Right, they tested it by training on the MIMIC-IV dataset and then seeing if it could work on the eICU database without changing anything.

Jane: The graph approach was much more robust during that transfer than the BERT approach was.

Meng: That makes sense because medical hierarchies are universal, while specific hospital coding habits can change.

Meng: If a model understands the underlying biology, it shouldn't care which hospital it's working in.

Lalam: This kind of stability is what we need to build tools that clinicians can actually trust in different parts of the world.

Tom: They even analyzed the actual embeddings and found that hierarchy makes the code clusters much more coherent.

Jane: It pulls related diagnoses closer together in the model's "mind," making everything much more organized.

Lu: It's like turning a pile of loose papers into a perfectly indexed library!

Meng: And if that organization helps with accuracy, then it's definitely worth the extra computation.

Tom: We should wrap this up and see what the big picture is here.

Conclusion: Tom: We've really covered a lot of ground today regarding Hierarchical Modeling of ICD Codes in EHR Foundation Models.

Jane: It's been so eye-opening to see how much we can learn just by respecting the structure that's already there.

Tom: They proved that those medical codes are far more than just arbitrary labels, and they did it with a very lightweight approach.

Lu: I can already see this being applied to even more complex systems, like mapping out every single protein interaction in the human body!

Meng: As long as we keep these implementations scalable for actual hospital production, this will be a major standard.

Lalam: It's a beautiful step toward technology that truly understands the profound nuances of human health.

Jane: Thanks for joining us on the show today!

Tom: We'll see you next time!

More episodes

← Home