Large Language Models are Powerful Electronic Health Record Encoders
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Large Language Models are Powerful Electronic Health Record Encoders".
Jane: The paper was written by Stefan Hegselmann, Georg von Arnim, Tillmann Rheude, Noel Kronenberg, David Sontag et al. from Berlin Institute of Health at Charité – Universitätsmedizin Berlin and Deutsches Herzzentrum der Charité – Medical Heart Center of Charité and German Heart Institute Berlin and Massachusetts Institute of Technology and Layer Health, Inc. and Fudan University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back, everyone. Today we're digging into a paper that's got a title that really says it all: "Large Language Models are Powerful Electronic Health Record Encoders." Jane, I have to say, that title is a bold claim, right?
Jane: It is, Tom. And honestly, it's a claim that a lot of people in the medical AI world have been waiting to hear tested properly. For years, we've had these specialized models trained only on hospital data, and they work, but they're kind of locked in a box. This paper asks whether a general-purpose LLM can just step in and do the job without ever seeing a single private medical record during training.
Tom: And that's the part that gets me excited. The idea that you don't need to train on millions of patient charts to understand a patient's chart. The authors are basically saying, "Hey, we can take a model that learned language from the entire internet, and it can read a patient's history as text and produce a useful summary of their health state."
Jane: Exactly. And the way they do it is almost cheeky in its simplicity. They take all those medical codes, the cryptic stuff like "LOINC/seven hundred eighteen-seven" and they just replace them with plain English descriptions. So instead of a code, you get "Hemoglobin." Then they list all of those descriptions for a patient, and they feed that list to the LLM.
Tom: Just a big list of medical terms. No fancy formatting, no special structure.
Jane: Just a newline-separated list. And the model, because it understands language, can make sense of it. It's like the difference between giving someone a spreadsheet of numbers versus giving them a paragraph describing the same information. The LLM is really good at the paragraph.
Tom: So the title isn't just hype. It's a direct challenge to the status quo. And the implications are huge, because if this works, it means we can build predictive tools for hospitals that don't have the resources to train their own massive foundation models.
Jane: Right. And it means those tools can actually transfer between hospitals, because they're not tied to one specific coding system. That's the promise, anyway. We'll get into whether they actually proved it in the results, but for now, let's just appreciate the audacity of the title.
Tom: Audacity is the right word. And I want to know how they actually tested this. That's coming up next.
Summary: Jane: So, Tom, we've established the bold title. Now let's talk about what the paper actually did, because the summary is pretty impressive. They took this LLM embedding model, Qwen3-Emb-8B, and they tested it on a benchmark called EHRSHOT.
Tom: EHRSHOT, that's the big one. Fifteen different clinical prediction tasks, from predicting if a patient will be readmitted to the hospital to whether they'll develop a new disease like hypertension or pancreatic cancer.
Jane: And the headline result is that this general-purpose LLM performed just as well as CLMBR-T-Base, which is a specialized model trained on millions of private EHRs from Stanford. We're talking a macro AUROC of zero point seven six nine for the LLM versus zero point seven six nine for the specialized model. A dead tie.
Tom: A dead tie. That's the kind of result that makes you do a double take. And it wasn't just one task. It was across all four task groups. The LLM even did slightly better on lab test prediction and new diagnosis assignment.
Jane: And here's the kicker. They didn't just test it on the same hospital's data. They took it to the UK Biobank, a completely different healthcare system with different coding practices. And there, the LLM actually outperformed the specialized model, zero point seven five one to zero point seven three six.
Tom: Because the specialized model only recognized sixteen percent of the UK's medical codes. It was lost. But the LLM, since it reads descriptions, could handle all of them.
Jane: That's the portability argument in action. It's not just about being as good; it's about being useful where the other model simply can't go without a massive remapping effort.
Tom: And they didn't stop there. They also showed that combining the LLM's embedding with the specialized model's embedding gave an even bigger boost, zero point seven eight eight. So they're not just competitors, they're complementary.
Jane: Which is a really nice finding. It suggests the LLM is capturing something different, maybe the semantic meaning of the codes, while the specialized model is capturing the temporal patterns in the sequences.
Tom: So the summary is: LLMs are powerful, they're portable, and they can work with the old models. What's not to like? Well, there's always a catch, and I think we need to talk about the cost.
Improvements: Tom: Alright, Jane, so we've got this amazing result, but the paper also makes a really important point about what needs to improve. And it's not the accuracy. It's the efficiency.
Jane: Oh, the runtime numbers are brutal. To encode the entire EHRSHOT benchmark, the specialized model CLMBR-T-Base took about six minutes. The LLM, Qwen3-Emb-8B, took almost twenty-two hours. That's a two-hundred-fold difference.
Tom: Twenty-two hours versus six minutes. That's not a small gap, that's a chasm. And that's on a cluster of eight H200 GPUs. So the paper is essentially saying, "Yes, we can match your performance, but we're going to use a lot more electricity to do it."
Jane: And that's why they frame it as a trade-off. You get portability and you get to skip the private training data, but you pay for it in compute. The paper is very honest about this.
Tom: So what's the improvement they're suggesting? How do we close that gap?
Jane: They point to a few things. First, they note that not all LLMs are created equal when it comes to long inputs. They tested three different models, and only Qwen3-Emb-8B handled the full eight thousand one hundred ninety-two-token context without degrading. The others got worse as the input got longer.
Tom: So the improvement is in the model architecture itself. We need LLMs that are actually good at paying attention to a very long list of medical terms.
Jane: Exactly. And they also suggest that serialization-free approaches could help. Instead of turning the EHR into text, which is a lossy process, you could have a model that directly processes the raw tabular data. That would remove the bias introduced by our manual formatting.
Tom: And what about the temporal aspect? I noticed they said adding timestamps didn't help.
Jane: That was a surprising finding. You'd think knowing when something happened would be crucial. But in their simple list format, the current embedding models didn't seem to use that information effectively. So there's a clear opportunity for improvement in temporal reasoning.
Tom: So the path forward is about making these models faster, better at long contexts, and better at understanding time. That's a solid roadmap. But I'm curious about the broader implications. What does this mean for the future of clinical AI?
Jane: I think it means we're going to see a lot more work on making LLMs practical for healthcare. The accuracy is there, now it's about the engineering.
First Page: Jane: You know, Tom, when I look back at the first page of this paper, I'm struck by how clearly they frame the problem. They start with the fact that EHRs are a goldmine of data, but they're messy. They're complex, they're heterogeneous, and they're full of missing information.
Tom: And the traditional machine learning approaches, they've kind of hit a wall. The paper mentions that deep learning models often only get modest improvements over simple logistic regression.
Jane: Right. So the field moved to these foundation models, like CLMBR-T-Base, that are pretrained on massive amounts of unlabeled EHR data. And that helped. But the paper identifies a fundamental obstacle: these models are trained on a fixed vocabulary of codes from one specific hospital system.
Tom: It's like learning a language from only one book. You become an expert on that book, but you can't read anything else.
Jane: That's a perfect analogy. And the numbers back it up. CLMBR-T-Base knows twenty-six thousand two hundred forty-nine codes. The UK Biobank has fifty thousand seven hundred two. Only seven thousand nine hundred sixty-nine of those overlap. So eighty-four percent of the codes in the UK data are completely invisible to the Stanford-trained model.
Tom: And that's the gap the LLM fills. Because it doesn't learn a fixed vocabulary, it learns language. And medical codes, when you translate them to text, are just language.
Jane: The authors make a really elegant point here. They say LLMs benefit from pretraining on vast general-purpose text corpora. That means they've seen medical terminology, they understand concepts like "hypertension" and "chest x-ray," even if they've never seen a specific hospital's billing code for it.
Tom: So the first page sets up this David and Goliath story. The small, specialized model versus the giant, general-purpose model. And the giant wins on portability.
Jane: But the giant is also slow, which we talked about. So it's not a clean victory. It's a trade-off. And I think that's the most honest and useful way to frame it.
Tom: And it's a trade-off that the paper explores in detail with the few-shot experiments, where the LLM really shines with limited data. But that's a story for our final segment.
Conclusion: Tom: Well, Jane, we've reached the end of our discussion on "Large Language Models are Powerful Electronic Health Record Encoders." And I think the takeaway is that this paper is a landmark for showing what's possible.
Jane: Absolutely. It proves that you don't need access to private medical data to build a powerful EHR encoder. A general-purpose LLM, trained on public text, can match and sometimes beat a specialized model. And it can do it across different countries and coding systems.
Tom: The trade-off is the compute cost. Twenty-two hours versus six minutes is a real hurdle for everyday clinical use. But the paper gives us a clear roadmap for improvement.
Jane: Right. We need better long-context models, better temporal reasoning, and more efficient architectures. The accuracy is there, now it's about making it practical.
Tom: And let's not forget the few-shot results. The LLM was especially strong when there was very little labeled data. That's a huge deal for rare diseases or new clinical settings where you don't have years of historical data.
Jane: That could be the most impactful finding of all. For a small clinic or a hospital in a developing country, being able to say, "We have a model that works, and we only needed a hundred examples to tune it," that's transformative.
Tom: So we're saying goodbye to this paper, but the ideas it introduces are going to stick around. It's a benchmark for how we think about foundation models in healthcare.
Jane: And a reminder that sometimes the most powerful tool is the one you already have. You just need to give it the right input. We'll be back soon with more research that's shaping the future. Until then, keep asking questions.
Tom: See you on the next one.
Stefan Hegselmann, Georg von Arnim, Tillmann Rheude, Noel Kronenberg, David Sontag, Gerhard Hindricks, Roland Eils, Benjamin Wild
Berlin Institute of Health at Charité – Universitätsmedizin Berlin · Deutsches Herzzentrum der Charité – Medical Heart Center of Charité and German Heart Institute Berlin · Massachusetts Institute of Technology · Layer Health, Inc. · Fudan University
cs.LG, cs.AI, cs.CL
Submitted: 2026-08-07
Updated: 2026-08-11
Code: https://github.com/stefanhgm/ehrshot-benchmark
License: http://creativecommons.org/licenses/by-sa/4.0/
Importance score: 72/100
The gist: The paper "Large Language Models are Powerful Electronic Health Record Encoders" by Stefan Hegselmann, Georg von Arnim, and colleagues investigates whether general-purpose Large Language Models
Key concepts
- Electronic Health Records (EHRs)
- Digital records containing a patient's medical history. The paper discusses how these records are complex and heterogeneous, often requiring specialized encoding methods to summarize a patient’s health state for predictive tools.
- Large Language Models (LLMs)
- General-purpose AI models trained on vast amounts of text (like the internet). The authors propose using LLMs to understand medical codes by translating them into plain English descriptions, allowing them to process data without needing private training sets.
- Encoding
- The process of converting raw clinical data, such as medical codes (e.g., LOINC), into a structured format that an AI model can understand and use for prediction or summarization of a patient's health status.
- Portability
- The ability of an AI tool to function across different healthcare settings or hospital systems. The episode highlights that LLMs are more portable because they rely on language descriptions rather than fixed, local coding vocabularies.
Terminology
Summary
The paper Large Language Models are Powerful Electronic Health Record Encoders
by Stefan Hegselmann, Georg von Arnim, and colleagues investigates whether general-purpose Large Language Models (LLMs) can serve as effective encoders of longitudinal Electronic Health Record (EHR) data for clinical prediction tasks.
The authors convert EHR data into plain text by replacing medical codes with natural-language descriptions, enabling general-purpose LLMs to produce high-dimensional embeddings for downstream prediction tasks without access to private medical training data.
They evaluate three instruction-tuned LLM embedding models: Qwen3-Embedding-8B (Qwen3-Emb-8B), GTE-Qwen2-7B-Instruct (Qwen2-Emb-7B), and LLM2Vec-Llama-3.1-8B-Instruct (Llama3.1-LLM2Vec-8B), with Qwen3-Emb-8B as the primary focus due to its recency and stronger long-context performance.
The default serialization is a simple newline-separated list of medical code descriptions, including units and values when available, with minimal preprocessing,
retaining only the most recent occurrence of each medical code
to stay within the 8,192-token context window. The authors use a logistic regression classifier on top of the embeddings, following the EHRSHOT protocol to isolate embedding quality from downstream model complexity.
Performance on EHRSHOT Benchmark: Qwen3-Emb-8B matched the in-domain EHR foundation model CLMBR-T-Base on EHRSHOT, with an overall macro-AUROC of 0.769 (0.744–0.794) versus 0.769 (0.746–0.792).
Qwen3-Emb-8B performed slightly better in three of four task categories (lab prediction, assignment of new diagnoses, and chest X-ray prediction). Task-level statistical testing showed Qwen3-Emb-8B significantly outperformed CLMBR-T-Base on thrombocytopenia, hyponatremia, and hyperkalemia, whereas CLMBR-T-Base performed significantly better on anemia and hypoglycemia.
Concatenating embeddings from both models improved performance to 0.788 (0.764–0.812), suggesting that the two models capture complementary information.
External Validation on UK Biobank: Qwen3-Emb-8B achieved slightly higher overall performance than CLMBR-T-Base, with 0.751 (0.740–0.761) compared to 0.736 (0.726–0.747).
Statistical testing showed significant improvements on six of 25 tasks and no significant differences otherwise.
A sensitivity analysis restricting Qwen3-Emb-8B to codes mappable to CLMBR-T-Base's vocabulary decreased performance to 0.743 (0.732–0.754), indicating that the gains on UKB can be explained by both broader vocabulary coverage and slightly improved generalization.
Few-Shot Performance: Qwen3-Emb-8B maintained slightly higher performance than CLMBR-T-Base for new diagnoses and chest X-ray tasks across most shot settings,
with statistical testing confirming significant improvements over CLMBR-T-Base on four versus one tasks in the 8-shot setting and on three versus two tasks in the 64-shot setting.
The authors conclude that the advantage of LLM embeddings is most apparent with limited labeled data.
Comparison with Baselines: The count-based model with ontology expansion, string and numeric values, and time binning achieved 0.777 (0.756–0.799), slightly exceeding both Qwen3-Emb-8B and CLMBR-T-Base
when using all data, but deteriorated in few-shot settings, indicating lower data efficiency than pretrained representations.
BioClinicalBERT, the best encoder-only model, achieved 0.705 (0.680–0.730), with Qwen3-Emb-8B significantly outperforming BioClinicalBERT on eight of 15 tasks.
Serialization Effects: Using first rather than most recent occurrences substantially reduced performance, especially for lab-test prediction and chest X-ray findings.
Adding date and time information did not improve performance.
A handcrafted Markdown serialization did not improve performance for Qwen3-Emb-8B relative to the simpler list-based serialization,
though it yielded small gains for lab-test prediction (0.006 to 0.021 AUROC).
Context Length and Temporal Scope: Qwen3-Emb-8B performed best at 4,096 tokens,
while Qwen2-Emb-7B performed best at 2,048 tokens and Llama3.1-LLM2Vec-8B at 1,024 to 2,048 tokens. For Qwen2-Emb-7B and Llama3.1-LLM2Vec-8B, performance dropped substantially at 8,192 tokens,
indicating only Qwen3-Emb-8B handled unstructured long-context EHR input robustly.
Qwen3-Emb-8B performed best with a one-year window,
while the count-based baseline improved with larger time windows.
Content Ablations: Task-specific instructions improved performance, with the largest drop occur[ing] for lab-test prediction
when instructions were removed. Removing individual code categories had limited impact overall except for lab results,
and no single modality matched the full-EHR representation.
The authors evaluated decoder-style LLMs (Qwen3-8B) with in-context learning and LoRA fine-tuning for both encoder and decoder variants. Decoder ICL achieved its best performance at 2-shot, with a macro-AUROC of 0.636 (0.588–0.685).
The fine-tuned decoder improved steadily, reaching macro-AUROC values of 0.632 (0.584–0.679) at k = 8, 0.708 (0.664–0.752) at k = 64, and 0.719 (0.673–0.764) at k = 128.
However, neither surpassed the frozen Qwen3-Emb-8B baseline from the main experiments, which reached 0.733 (0.686–0.777) at k = 128.
CLMBR-T-Base required 6:04 minutes to encode all examples of the EHRSHOT benchmark, whereas Qwen3-Emb-8B required 21:48:56 hours.
The authors note that LLM-based methods achieved similar predictive performance at the cost of a substantially larger memory footprint and higher computational cost.
The authors conclude that general-purpose LLM embedding models pretrained on large-scale natural-language corpora can serve as effective encoders of longitudinal EHR data for clinical prediction.
They highlight a trade-off between the flexibility of LLM embeddings and the efficiency of specialized EHR foundation models,
noting that LLM embeddings offer greater flexibility and portability across datasets
while domain-specific models remain substantially more computationally efficient.
Improvements for AI systems
Based on the paper, here are specific improvements I can implement in AI systems:
Improvement: Add a module that converts structured EHR data (medical codes, lab values, visits) into plain-text, newline-separated lists with task-specific instructions, retaining only the most recent occurrence of each code.
Capability: The system can now process any EHR dataset without requiring site-specific code mappings or fixed vocabularies, enabling cross-institutional deployment without retraining.
Improvement: Implement a pipeline that:
-
Uses Qwen3-Embedding-8B (or similar contrastive-trained LLM) to generate 4096-dimensional patient embeddings from serialized EHR text
-
Feeds embeddings into a logistic regression classifier
-
Supports few-shot learning (8-128 examples per class)
Improvement: Prepend task-specific prompts (e.g., will the patient stay in the hospital for more than 7 days
) to the EHR serialization before embedding extraction.
Improvement: Concatenate LLM embeddings with domain-specific EHR foundation model embeddings (e.g., CLMBR-T-Base) before classification.
Improvement: Implement a serialization strategy that:
-
Truncates inputs to 8,192 tokens
-
Prioritizes most recent medical events
-
Uses list format rather than structured formats (JSON/XML/YAML)
Improvement: Build a system that:
-
Uses frozen LLM embeddings (no fine-tuning)
-
Trains logistic regression on k positive and k negative examples (k=1 to 128)
-
Handles class imbalance via balanced sampling
Improvement: Add a vocabulary-agnostic encoding layer that maps any clinical code to its natural-language description, eliminating the need for ontology mapping or code harmonization.
Improvement: Implement adaptive time-window selection (1 hour to 3 years) based on task type, with automatic detection of optimal window length.
Improvement: Add a diagnostic module that tests model sensitivity to instruction variations (generic vs. task-specific vs. empty), identifying tasks where instructions matter most.
Improvement: Implement a model-selection framework that balances parameter count (0.6B to 8B) against performance, with runtime profiling (5-22 hours for EHRSHOT encoding).
Abstract
Electronic Health Records (EHRs) offer considerable potential for clinical prediction, but their complexity and heterogeneity challenge traditional machine learning. Domain-specific EHR foundation models trained on unlabeled EHR data have shown improved predictive accuracy and generalization. However, their development is constrained by limited data access and site-specific vocabularies. We convert EHR data into plain text by replacing medical codes with natural-language descriptions, enabling general-purpose Large Language Models (LLMs) to produce high-dimensional embeddings for downstream prediction tasks without access to private medical training data. LLM-based embeddings perform on par with a specialized EHR foundation model, CLMBR-T-Base, across 15 clinical tasks from the EHRSHOT benchmark. In an external validation using the UK Biobank, an LLM-based model shows statistically significant improvements for some tasks, which we attribute to higher vocabulary coverage and slightly better generalization. Overall, we reveal a trade-off between the computational efficiency of specialized EHR models and the portability and data independence of LLM-based embeddings.
Sources
- On the Opportunities and Risks of Foundation Models
- CORE-BEHRT: A Carefully Optimized and Rigorously Evaluated BEHRT
- CEHR-BERT: Incorporating temporal information from structured EHR data to improve prediction tasks
- Generative Medical Event Models Improve with Scale
- Prompting Large Language Models for Zero-Shot Clinical Prediction with Structured Longitudinal Electronic Health Record Data
- LLMs-based Few-Shot Disease Predictions using EHR: A Novel Approach Combining Predictive Agent Reasoning and Critical Agent Instruction
- ClinicalBench: Can LLMs Beat Traditional ML Models in Clinical Prediction?
- DeLLiriuM: A large language model for delirium prediction in the ICU using structured EHR
- Qwen2 Technical Report
- Qwen3 Technical Report
- NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models
- Generative Representational Instruction Tuning
- Towards General Text Embeddings with Multi-stage Contrastive Learning
- Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
- The Llama 3 Herd of Models
- Repetition Improves Language Model Embeddings
- LoRA: Low-Rank Adaptation of Large Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks