Scaling Electronic Health Record Foundation Models for Population Health Management

arXiv:2506.00209 · cs.LG, cs.CL · Submitted 2026-08-24 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Scaling Electronic Health Record Foundation Models for Population Health Management".

Jane: The paper was written by Liwen Sun, Hao-Ren Yao, Gary Gao, Ophir Frieder and Chenyan Xiong from Carnegie Mellon University and Georgetown University, Department of Computer Science (implied by email domain).

Tom: Stay tuned as we take you through the paper and discuss its implications.

Summary: Jane: So, we know CATCH-FM is a prescreening tool based on existing records, but how does it actually function? The authors detail that it's pretrained on large-scale longitudinal EHR data.

Lu: They use the massive Taiwanese National Health Insurance Research Database to build this benchmark, which provides over three million patients and billions of medical events for training.

Meng: Training a foundation model on such a vast dataset is computationally demanding, but the researchers successfully managed it at scale, which suggests that the technical hurdles are manageable with appropriate hardware.

Jane: It's not just any random data; they structure the patient record into a sequence of medical codes, which is quite detailed.

Lu: They are essentially mapping out a complex health trajectory by treating every medical event as an atomic token in a sequence, allowing us to see patterns emerge over time.

Meng: This sequential approach means we're not just looking at snapshots; we're modeling the entire history of the patient, which is crucial for understanding risk factors.

Lalam: The implications of this summary are profound because it shows how AI can act as a global equalizer, allowing high-quality risk assessment in regions where expensive medical imaging might not be available.

Tom: It sounds like a really powerful combination of data scale and structure, but what’s the next major point they found?

Improvements: Jane: The authors show that CATCH-FM significantly outperforms traditional feature-based models and also beat general large language models, which is a huge win for the clinical community.

Lu: It’s proving that using medical codes as their own language—treating them like tokens—is much more effective than treating them like simple text strings in the way standard LLMs do.

Meng: I’m impressed that CATCH-FM achieves state-of-the-art performance on the EHRSHOT benchmark, which is a massive hurdle because of distribution shifts between different healthcare systems.

Lu: My initial thoughts on the scaling laws presented in the paper are that they provide a clear blueprint for future research; seeing how FLOPs and model size relate to loss optimization gives us a very concrete path forward for building even bigger systems.

Jane: It’s not just beating other models, though; it' is achieving high sensitivity—like fifty percent or seventy percent—while maintaining a very high level of specificity at the ninety-nine percent cutoff.

Meng: That balance between finding cases and avoiding false alarms is what makes this actually useful for clinicians, which is something I find compelling.

Lalam: The ability to generalize across diverse coding systems and clinical settings shows that this AI can transcend borders, suggesting a future where risk assessment isn't constrained by geographical or economic limitations.

Tom: That’s an impressive performance profile; it's not just a theoretical improvement, but a practical one, but what does this model handle beyond the primary cancer types?

Generalization and Application: Jane: The paper really shows that by providing this low-risk, efficient pre-screening tool to healthcare providers, we can help them decide who needs further attention and when they need it most.

Lu: I see the potential for this as a paradigm shift where our AI's role is not just to diagnose but to guide the entire clinical workflow based on population health trends.

Meng: The generalization ability is impressive too, especially how CATCH-FM performs robustly on the EHRSHOT dataset, which contains patients from a completely different healthcare system than NHIRD.

Jane: It’s not just about the big data; it' also about the "why" behind it—the model clearly captures non-trivial risk factors that were recently discovered in medical research.

Lu: I think this ability to find subtle, hidden patterns is where the true power of data-driven AI shines, moving beyond simple correlation.

Meng: It’s a practical tool that works, and my focus is on ensuring that its deployment—the "how"—is robust enough to handle real-world data variability without losing accuracy across different hospital sites.

Lalam: The most impactful vision is one where this technology fosters a culture of sustained vigilance, empowering patients and providers alike with actionable insights into their health trajectory.

Tom: That’s a powerful way to look at it all, Jane. We've seen how this model handles complex data and what its future could look like, but we need to wrap up our discussion on "Scaling Electronic Health Record Foundation Models for Population Health Management."

Conclusion: Jane: We have seen that CATCH-FM is a low-cost, high-impact solution that provides crucial support to those who lack access to invasive procedures.

Lu: I think the power of seeing those scaling laws is that it proves we are ready for much larger, more sophisticated AI systems in medicine, moving beyond just handling data toward actually solving complex health problems.

Meng: It’s a viable solution we can actually start building towards reality, and the complexity of coding standards is something we've successfully navigated with this model.

Lalam: This technology fosters a culture where proactive care becomes the standard, enabling people to manage their health journey before symptoms even start.

Tom: That's exactly what I mean; we’re shifting from a reactive model of treatment to a predictive one, which is such a powerful change for sure.

Jane: It feels like an enormous step forward in making healthcare more equitable and efficient for everyone who needs it.

Lu: Indeed, it sets the stage for so many more creative applications of AI that we're only beginning to imagine in large-scale medical contexts.

Meng: We can start thinking about how this running system might scale up to a massive global deployment now, which is the next big hurdle for operational planning.

Lalam: I hope this work inspires a focus on preventative health globally, moving away from crisis management toward a culture of sustained vigilance.

Tom: That's an incredible vision to end on; we've learned so much about "Scaling Electronic Health Record Foundation Models for Population Health Management" today, and it’s exciting to see the future possibilities ahead.

Jane: We have a lot more ground to cover next time, but we hope you enjoyed this deep dive into the paper.

Tom: Absolutely everyone; this is a truly groundbreaking piece of work, and we'll be back soon with more AI research!

Liwen Sun, Hao-Ren Yao, Gary Gao, Ophir Frieder, Chenyan Xiong

Carnegie Mellon University · Georgetown University, Department of Computer Science (implied by email domain)

cs.LG, cs.CL

Submitted: 2026-08-24

Updated: 2026-08-25

Importance score: 90/100

The gist: The paper details an evaluation of scaling Electronic Health Record (EHR) foundation models for population health management, focusing specifically on cancer screening tasks across multiple cohorts,

Key concepts

CATCH-FM
CATCH-FM is a pre-screening tool designed to analyze patient records. It was trained on the massive Taiwanese National Health Insurance Research Database, containing over three million patients and billions of medical events. It processes these records by structuring them as sequences of medical codes.
EHRSHOT Benchmark
EHRSHOT is a benchmark used to test AI models on Electronic Health Records. The model's ability to perform well here demonstrates its generalization across different healthcare systems, showing that the AI can be robust even when dealing with data from a completely different system than the one it was originally trained on.
Population Health Management
This concept involves using AI to guide clinical workflows based on population health trends. CATCH-FM helps healthcare providers identify who needs attention by analyzing complex data, moving away from a reactive treatment model to a predictive one.

Terminology

Summary

The paper details an evaluation of scaling Electronic Health Record (EHR) foundation models for population health management, focusing specifically on cancer screening tasks across multiple cohorts, including Pancreatic, Liver, and Lung cancers.

Model Performance and Comparisons:

The study evaluates the performance of several models—including CATCH-FM, Qwen2.5, and XGBoost—using various metrics such as AUROC (Area Under the Receiver Operating Characteristic curve), AUPRC (Area Under the Precision-Recall Curve), F1 score, Specificity, and Sensitivity. The evaluation demonstrates that model performance varies depending on the specific cancer cohort and the prediction setting used.

Methodological Deep Dive:

The research includes detailed analyses of model interpretability and data scaling:

  • Interpretability Experiments (Sparse Autoencoder - SAE): To process positive patients’ event token sequences, the fine-tuned CATCH-FM-1b model is used to obtain hidden states h for every token. An SAE is then trained on the hidden states of the [EOS] token, h[EOS], which serves as an aggregated representation of patient trajectories. The SAE implementation involves an encoder parameterized by W enc and b enc and a decoder parameterized by W dec and b dec, where the latent representation is calculated as:

z = TopK(W enc (h[EOS] - b dec) + b enc)

[EOS] = W dec z + b dec

The SAE is trained using the mean squared error (MSE) loss for reconstruction. The findings indicate that as the number of active latent features increases, both cancer screening performance and reconstruction quality improve. Crucially, with just 16 active latent features, the SAE’s reconstructed embeddings already capture enough information to match the original performance, suggesting that the latent cancer signal extracted from the fine-tuned CATCH-FM is inherently low-dimensional.

  • Supervised Data Scale Analysis: The paper investigates CATCH-FM's robustness when trained with varying amounts of supervised finetuning labels. The results show that CATCH-FM maintains a high specificity of 99% with as few as 10k training labels (positives). Furthermore, its sensitivity increases significantly with more finetuning data, crossing the 50% threshold with only 20k total patient data across two decades, which is fewer than a typical hospital.

  • Adaptability and Prediction Settings: The model's adaptability is highlighted by noting that CATCH-FM can be conveniently adapted to different prediction settings based on healthcare professional’s preference, such as performing better in setups where more short term risk factors may be observable.

Conclusion:

Overall, the research demonstrates the capability of foundation models like CATCH-FM to perform complex tasks such as cancer screening across diverse cohorts. The model shows strong performance metrics and robustness even when trained with limited data, while its interpretability analysis confirms that the underlying cancer signal is efficiently captured in a low-dimensional latent space.

Improvements for AI systems

Based on a thorough analysis of the CATCH-FM framework presented in this paper, I have identified several critical areas for improvement and expansion across existing AI systems. These improvements are designed to enhance efficiency, clinical trust, and cross-system generalization.


Current Limitation: While CATCH-FM demonstrated robustness in distribution shifts (NHIRD vs EHRSHOT), the reliance on specific ICD/NHI codes remains a bottleneck for seamless global deployment.

Improvement: We must develop an advanced mapping layer, extending the semantic text matching used in Table 12 of the paper, into a Graph-Based Cross-Ontology Alignment Module. This module would use techniques beyond simple cosine similarity (e.g., knowledge graph embedding) to map codes from disparate systems (SNOMED, CPT, local proprietary codes) into a standardized latent space, independent of the original code's literal structure.

What the Improved System Can Do:

  • Achieve True Zero-Shot Generalization: The system can ingest EHR data from completely unknown or non-standardized healthcare systems (e.g., emerging market databases), map them into the CATCH-FM latent space, and provide robust risk predictions without needing manual code translation, enabling deployment in underserved populations globally.

Current Limitation: The current interpretation methods (SAE and N2G) are retrospective—they explain why the model made a prediction after the fact. This is insufficient for dynamic clinical decision support.

Improvement: Integrate the Sparse Autoencoder (SAE) into a Real-Time Feature Attribution Engine. Instead of just analyzing h[EOS] post-hoc, we must modify the forward pass to calculate and output which latent features are contributing most strongly to the risk score at every decision point in real time. This requires modifying the SAE training objective to include a localized gradient tracking mechanism (e.g).

What the Improved System Can Do:

  • Provide Justification for Treatment: When CATCH-FM flags a patient as high risk, it will not only provide a probability but also deliver a concise, clinically relevant list of the top 5 most influential factors (e.g., High Risk due to: Type 2 Diabetes [Factor A], Persistent Liver Enzyme Elevation [Factor B]). This allows clinicians to immediately verify the prediction against known patient conditions, increasing clinical trust and reducing diagnostic uncertainty.

Current Limitation: The paper established scaling laws for fixed FLOP budgets (Table 3). However, real-world hospital resources vary wildly in computational power and budget.

Improvement: Implement an Automated Resource Allocation Scheduler (ARAS) that allows the model to dynamically select the optimal scale (CATCH-FM-160m vs CATCH-FM-2.4b) based on a user-defined constraint (e.g., maximum allowable GPU hours or maximum latency for inference). This requires generalizing the IsoFLOP curves into a practical, operational framework.

What the Improved System Can Do:

  • Optimize Deployment Cost: A healthcare provider can set their budget (e.g., I have 128 GPU hours available) and the system will automatically train/deploy the most effective model scale within that constraint, ensuring maximum predictive power without wasteful over-provisioning of resources.

Current Limitation: CATCH-FM was highly successful in specific cancer risk prediction (pancreatic, liver, lung). The subsequent target cancer definition is broad but not fully leveraged for cross-disease prediction.

Improvement: Develop a Multi-Task Unified Prediction Head. Instead of training separate finetuning layers for each cancer type, we will train a single shared set of weights (theta) and utilize multiple output heads (W i) corresponding to different target cancers. This is achieved by adapting the cross-entropy loss to include a penalty term that encourages shared feature representation across all diseases (a form of parameter sharing).

What the Improved System Can Do:

  • Identify Co-morbid Risk Clusters: The system can predict not just if a patient has cancer, but which specific combination of multiple cancers is most likely (e.g., High risk for both Liver and Lung cancers within 5 years), providing holistic risk assessment that single-task models cannot achieve.

Sources

Related papers