Large Language Models are Powerful Electronic Health Record Encoders

summary

Video file (mp4)

The gist

The paper "Large Language Models are Powerful Electronic Health Record Encoders" by Stefan Hegselmann, Georg von Arnim, and colleagues investigates whether general-purpose Large Language Models

In short

The episode discusses 'Large Language Models are Powerful Electronic Health Record Encoders,' detailing how general-purpose LLMs can encode patient health data effectively. Hosts review the paper's findings, noting that LLMs match specialized models in performance and show superior portability across different hospital systems, though they face significant computational cost hurdles.

Key concepts

Electronic Health Records (EHRs)
Digital records containing a patient's medical history. The paper discusses how these records are complex and heterogeneous, often requiring specialized encoding methods to summarize a patient’s health state for predictive tools.
Large Language Models (LLMs)
General-purpose AI models trained on vast amounts of text (like the internet). The authors propose using LLMs to understand medical codes by translating them into plain English descriptions, allowing them to process data without needing private training sets.
Encoding
The process of converting raw clinical data, such as medical codes (e.g., LOINC), into a structured format that an AI model can understand and use for prediction or summarization of a patient's health status.
Portability
The ability of an AI tool to function across different healthcare settings or hospital systems. The episode highlights that LLMs are more portable because they rely on language descriptions rather than fixed, local coding vocabularies.

Terminology used across episodes

This episode discusses

The paper

Large Language Models are Powerful Electronic Health Record Encoders · Read on arXiv

Stefan Hegselmann, Georg von Arnim, Tillmann Rheude, Noel Kronenberg, David Sontag, Gerhard Hindricks, Roland Eils, Benjamin Wild

Berlin Institute of Health at Charité – Universitätsmedizin Berlin · Deutsches Herzzentrum der Charité – Medical Heart Center of Charité and German Heart Institute Berlin · Massachusetts Institute of Technology · Layer Health, Inc. · Fudan University

Electronic Health Records (EHRs) offer considerable potential for clinical prediction, but their complexity and heterogeneity challenge traditional machine learning. Domain-specific EHR foundation models trained on unlabeled EHR data have shown improved predictive accuracy and generalization. However, their development is constrained by limited data access and site-specific vocabularies. We convert EHR data into plain text by replacing medical codes with natural-language descriptions, enabling general-purpose Large Language Models (LLMs) to produce high-dimensional embeddings for downstream prediction tasks without access to private medical training data. LLM-based embeddings perform on par with a specialized EHR foundation model, CLMBR-T-Base, across 15 clinical tasks from the EHRSHOT benchmark. In an external validation using the UK Biobank, an LLM-based model shows statistically significant improvements for some tasks, which we attribute to higher vocabulary coverage and slightly better generalization. Overall, we reveal a trade-off between the computational efficiency of specialized EHR models and the portability and data independence of LLM-based embeddings.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Large Language Models are Powerful Electronic Health Record Encoders".

Jane: The paper was written by Stefan Hegselmann, Georg von Arnim, Tillmann Rheude, Noel Kronenberg, David Sontag et al. from Berlin Institute of Health at Charité – Universitätsmedizin Berlin and Deutsches Herzzentrum der Charité – Medical Heart Center of Charité and German Heart Institute Berlin and Massachusetts Institute of Technology and Layer Health, Inc. and Fudan University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back, everyone. Today we're digging into a paper that's got a title that really says it all: "Large Language Models are Powerful Electronic Health Record Encoders." Jane, I have to say, that title is a bold claim, right?

Jane: It is, Tom. And honestly, it's a claim that a lot of people in the medical AI world have been waiting to hear tested properly. For years, we've had these specialized models trained only on hospital data, and they work, but they're kind of locked in a box. This paper asks whether a general-purpose LLM can just step in and do the job without ever seeing a single private medical record during training.

Tom: And that's the part that gets me excited. The idea that you don't need to train on millions of patient charts to understand a patient's chart. The authors are basically saying, "Hey, we can take a model that learned language from the entire internet, and it can read a patient's history as text and produce a useful summary of their health state."

Jane: Exactly. And the way they do it is almost cheeky in its simplicity. They take all those medical codes, the cryptic stuff like "LOINC/seven hundred eighteen-seven" and they just replace them with plain English descriptions. So instead of a code, you get "Hemoglobin." Then they list all of those descriptions for a patient, and they feed that list to the LLM.

Tom: Just a big list of medical terms. No fancy formatting, no special structure.

Jane: Just a newline-separated list. And the model, because it understands language, can make sense of it. It's like the difference between giving someone a spreadsheet of numbers versus giving them a paragraph describing the same information. The LLM is really good at the paragraph.

Tom: So the title isn't just hype. It's a direct challenge to the status quo. And the implications are huge, because if this works, it means we can build predictive tools for hospitals that don't have the resources to train their own massive foundation models.

Jane: Right. And it means those tools can actually transfer between hospitals, because they're not tied to one specific coding system. That's the promise, anyway. We'll get into whether they actually proved it in the results, but for now, let's just appreciate the audacity of the title.

Tom: Audacity is the right word. And I want to know how they actually tested this. That's coming up next.

Summary: Jane: So, Tom, we've established the bold title. Now let's talk about what the paper actually did, because the summary is pretty impressive. They took this LLM embedding model, Qwen3-Emb-8B, and they tested it on a benchmark called EHRSHOT.

Tom: EHRSHOT, that's the big one. Fifteen different clinical prediction tasks, from predicting if a patient will be readmitted to the hospital to whether they'll develop a new disease like hypertension or pancreatic cancer.

Jane: And the headline result is that this general-purpose LLM performed just as well as CLMBR-T-Base, which is a specialized model trained on millions of private EHRs from Stanford. We're talking a macro AUROC of zero point seven six nine for the LLM versus zero point seven six nine for the specialized model. A dead tie.

Tom: A dead tie. That's the kind of result that makes you do a double take. And it wasn't just one task. It was across all four task groups. The LLM even did slightly better on lab test prediction and new diagnosis assignment.

Jane: And here's the kicker. They didn't just test it on the same hospital's data. They took it to the UK Biobank, a completely different healthcare system with different coding practices. And there, the LLM actually outperformed the specialized model, zero point seven five one to zero point seven three six.

Tom: Because the specialized model only recognized sixteen percent of the UK's medical codes. It was lost. But the LLM, since it reads descriptions, could handle all of them.

Jane: That's the portability argument in action. It's not just about being as good; it's about being useful where the other model simply can't go without a massive remapping effort.

Tom: And they didn't stop there. They also showed that combining the LLM's embedding with the specialized model's embedding gave an even bigger boost, zero point seven eight eight. So they're not just competitors, they're complementary.

Jane: Which is a really nice finding. It suggests the LLM is capturing something different, maybe the semantic meaning of the codes, while the specialized model is capturing the temporal patterns in the sequences.

Tom: So the summary is: LLMs are powerful, they're portable, and they can work with the old models. What's not to like? Well, there's always a catch, and I think we need to talk about the cost.

Improvements: Tom: Alright, Jane, so we've got this amazing result, but the paper also makes a really important point about what needs to improve. And it's not the accuracy. It's the efficiency.

Jane: Oh, the runtime numbers are brutal. To encode the entire EHRSHOT benchmark, the specialized model CLMBR-T-Base took about six minutes. The LLM, Qwen3-Emb-8B, took almost twenty-two hours. That's a two-hundred-fold difference.

Tom: Twenty-two hours versus six minutes. That's not a small gap, that's a chasm. And that's on a cluster of eight H200 GPUs. So the paper is essentially saying, "Yes, we can match your performance, but we're going to use a lot more electricity to do it."

Jane: And that's why they frame it as a trade-off. You get portability and you get to skip the private training data, but you pay for it in compute. The paper is very honest about this.

Tom: So what's the improvement they're suggesting? How do we close that gap?

Jane: They point to a few things. First, they note that not all LLMs are created equal when it comes to long inputs. They tested three different models, and only Qwen3-Emb-8B handled the full eight thousand one hundred ninety-two-token context without degrading. The others got worse as the input got longer.

Tom: So the improvement is in the model architecture itself. We need LLMs that are actually good at paying attention to a very long list of medical terms.

Jane: Exactly. And they also suggest that serialization-free approaches could help. Instead of turning the EHR into text, which is a lossy process, you could have a model that directly processes the raw tabular data. That would remove the bias introduced by our manual formatting.

Tom: And what about the temporal aspect? I noticed they said adding timestamps didn't help.

Jane: That was a surprising finding. You'd think knowing when something happened would be crucial. But in their simple list format, the current embedding models didn't seem to use that information effectively. So there's a clear opportunity for improvement in temporal reasoning.

Tom: So the path forward is about making these models faster, better at long contexts, and better at understanding time. That's a solid roadmap. But I'm curious about the broader implications. What does this mean for the future of clinical AI?

Jane: I think it means we're going to see a lot more work on making LLMs practical for healthcare. The accuracy is there, now it's about the engineering.

First Page: Jane: You know, Tom, when I look back at the first page of this paper, I'm struck by how clearly they frame the problem. They start with the fact that EHRs are a goldmine of data, but they're messy. They're complex, they're heterogeneous, and they're full of missing information.

Tom: And the traditional machine learning approaches, they've kind of hit a wall. The paper mentions that deep learning models often only get modest improvements over simple logistic regression.

Jane: Right. So the field moved to these foundation models, like CLMBR-T-Base, that are pretrained on massive amounts of unlabeled EHR data. And that helped. But the paper identifies a fundamental obstacle: these models are trained on a fixed vocabulary of codes from one specific hospital system.

Tom: It's like learning a language from only one book. You become an expert on that book, but you can't read anything else.

Jane: That's a perfect analogy. And the numbers back it up. CLMBR-T-Base knows twenty-six thousand two hundred forty-nine codes. The UK Biobank has fifty thousand seven hundred two. Only seven thousand nine hundred sixty-nine of those overlap. So eighty-four percent of the codes in the UK data are completely invisible to the Stanford-trained model.

Tom: And that's the gap the LLM fills. Because it doesn't learn a fixed vocabulary, it learns language. And medical codes, when you translate them to text, are just language.

Jane: The authors make a really elegant point here. They say LLMs benefit from pretraining on vast general-purpose text corpora. That means they've seen medical terminology, they understand concepts like "hypertension" and "chest x-ray," even if they've never seen a specific hospital's billing code for it.

Tom: So the first page sets up this David and Goliath story. The small, specialized model versus the giant, general-purpose model. And the giant wins on portability.

Jane: But the giant is also slow, which we talked about. So it's not a clean victory. It's a trade-off. And I think that's the most honest and useful way to frame it.

Tom: And it's a trade-off that the paper explores in detail with the few-shot experiments, where the LLM really shines with limited data. But that's a story for our final segment.

Conclusion: Tom: Well, Jane, we've reached the end of our discussion on "Large Language Models are Powerful Electronic Health Record Encoders." And I think the takeaway is that this paper is a landmark for showing what's possible.

Jane: Absolutely. It proves that you don't need access to private medical data to build a powerful EHR encoder. A general-purpose LLM, trained on public text, can match and sometimes beat a specialized model. And it can do it across different countries and coding systems.

Tom: The trade-off is the compute cost. Twenty-two hours versus six minutes is a real hurdle for everyday clinical use. But the paper gives us a clear roadmap for improvement.

Jane: Right. We need better long-context models, better temporal reasoning, and more efficient architectures. The accuracy is there, now it's about making it practical.

Tom: And let's not forget the few-shot results. The LLM was especially strong when there was very little labeled data. That's a huge deal for rare diseases or new clinical settings where you don't have years of historical data.

Jane: That could be the most impactful finding of all. For a small clinic or a hospital in a developing country, being able to say, "We have a model that works, and we only needed a hundred examples to tune it," that's transformative.

Tom: So we're saying goodbye to this paper, but the ideas it introduces are going to stick around. It's a benchmark for how we think about foundation models in healthcare.

Jane: And a reminder that sometimes the most powerful tool is the one you already have. You just need to give it the right input. We'll be back soon with more research that's shaping the future. Until then, keep asking questions.

Tom: See you on the next one.

More episodes

← Home