Digitally enriching a high-risk population for pancreatic cancer using routine blood-based measures and clinical histories
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Digitally enriching a high-risk population for pancreatic cancer using routine blood-based measures and clinical histories".
Tom: Longitudinal sequences of coded diagnoses and blood test values accrued by patients throughout their clinical interactions were used to train a custom Transformer-based neural network with a multi-head attention mechanism…
Jane: First, who's behind it and why it matters.
Paper summary: Tom: So, to summarize what we just heard about "Digitally enriching a high-risk population for pancreatic cancer using routine blood-based measures and clinical histories," the thesis is that latent indicators of pancreatic pathology are visible in an individual’s disease and blood test trajectories, and these sequences can predict future development of pancreatic cancer Tom The authors claimed they could use a custom Transformer-based neural network with a multi-head attention mechanism for this prediction over several years.
Jane: It claims that this model can do two main things: first, it predicts the risk of pancreatic cancer with a lead time spanning one, two, or three years ahead of diagnosis, and second, it helps to stratify populations so doctors know who needs targeted screening.
Tom: And what really makes this relevant for us is the cohort size they used; they trained it on six thousand seventeen adults diagnosed with pancreatic cancer and one hundred seventy-seven thousand eighty-one controls who had a median medical history of about twelve years before getting their diagnosis Tom That large dataset suggests the model has been tested in a robust setting.
Lu: And those results are quite compelling when you look at the external validation they performed using leave-one-site-out testing on out-of-sample data, showing mean area under the receiver operating characteristic values of zero point eight three seven for a one-year lead time, zero point seven nine seven for a two-year lead time, and some value for a three-year lead time Lu Those AUROC numbers give us a concrete measure of how well it performs in predicting those specific future risk windows.
Meng: So the practical implication is that if this model proves accurate outside of the original study, we could start using it to flag patients who are at high risk for pancreatic cancer based on their existing clinical history and routine blood work.
Jane: That’s precisely what the paper suggests—moving from reactive treatment after symptoms appear to a more proactive screening strategy driven by AI prediction.
Conclusion: Tom: Looking at the title, "Digitally enriching a high-risk population for pancreatic cancer using routine blood-based measures and clinical histories," it really paints a picture of how we can use existing healthcare data to build better health systems Tom The authors are Chris Varghese, Leo Y. Li-Han, Richa Bisht, Ellen Larson, Frank Lee, Ryan M. Carr, Tanios S. Bekaii-Saab3, Shounak Majumder4, John D. Halamka5, Mark Truty1, Ajit H. Goenka6, Hojjat Salehinejad7,8 and Cornelius A Thiels1 Tom The implication is that we are setting the stage for a first population-level digital enrichment tool designed to widen access to curative-intent management of pancreatic cancer Tom It’s about using AI not just to diagnose, but to manage the entire journey of risk assessment before the disease progresses significantly.
Jane: I think what this means in simple terms is that we are moving toward a system where we can identify those at highest risk much earlier than current methods allow, allowing for intervention when it has the best chance of success.
Jane: This paper lays a foundation for a tool that can help tailor screening efforts to specific groups based on their unique longitudinal data patterns.
Lu: And from a theoretical view, this work establishes how temporal patterns in healthcare interactions are essential features for accurate risk prediction in complex diseases like pancreatic cancer Lu It shows that the order and sequence of medical events matters immensely when training these kinds of models.
Meng: Practically, this suggests a future where routine clinical data feeds into an intelligent system that helps coordinate care across different stages of a patient's journey.
Tom: So, to wrap up, the implication is clear: this work lays the foundation for a first population-level digital enrichment tool to widen access to curative-intent management of pancreatic cancer Tom It’s about using AI not just to diagnose, but to manage the entire journey of risk assessment before the disease progresses significantly.
Jane: And it’s exciting because it shows that routine clinical data, when processed correctly, can give us a better predictive edge in oncology.
Lalam: From my perspective as an LLM, this advance has the potential to improve culture by helping healthcare providers focus their efforts where they are most needed based on these precise risk stratifications Lalam It means shifting from broad screening approaches to highly personalized pathways guided by robust data.
Tom: Exactly! We’ve seen how powerful this predictive modeling can be when it’s grounded in real patient trajectories, and I think we need to keep an eye on how this technology gets integrated into actual clinical workflows.
Department of Surgery, Mayo Clinic, Rochester, MN, USA · Department of Surgery, University of Auckland, Auckland, NZ · Department of Hematology and Oncology, Mayo Clinic Phoenix/Rochester/Minnesota · Division of Gastroenterology and Hepatology at Mayo Clinic Rochester · Mayo Clinic Platform
cs.LG, q-bio.QM
Submitted: 2026-05-28
Updated: 2026-09-27
Code: https://github.com/lcapacitor/premod
Project page: https://ohdsi.github.io/CommonDataModel/cdm53.html
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 77/100
The gist: Longitudinal sequences of coded diagnoses and blood test values accrued by patients throughout their clinical interactions were used to train a custom Transformer-based neural network with a
Key concepts
- Transformer Encoder Architecture
- This is the core neural network structure used, similar to advanced language models, designed to process sequences of time-ordered medical data. It uses 'multi-head attention' to weigh the importance of different past events when making a future prediction about cancer risk.
- Unique Time-Bucket Encoding
- Instead of treating every visit separately, this method groups diagnostic codes and blood tests that occurred within the same month into one single data point. This simplifies the input data while ensuring that recent, related clinical events are not lost in the encoding process.
- Bayesian Update Approach to Calibration
- This technique adjusts the model's raw risk score to provide more accurate, context-specific probabilities. It allows the model to maintain its predictive accuracy (AUROC) while giving clinicians a probability that is meaningful for different types of healthcare environments.
Terminology
Summary
Longitudinal sequences of coded diagnoses and blood test values accrued by patients throughout their clinical interactions were used to train a custom Transformer-based neural network with a multi-head attention mechanism to predict risk of pancreatic cancer with a multiyear lead time and risk-stratify populations for targeted screening.
Model Architecture and Data Representation
The study utilizes a Transformer encoder architecture, which comprises four components: an input encoder, a feature extractor, a feature aggregator, and a classifier. To improve data representation portability compared to previous efforts, the researchers applied a unique time-bucketed visit encoding,
where diagnostic codes and blood test values acquired within a month are counted together in the encoding space to simplify integration while minimizing short-scale noise. The input encoder linearly projects the row-wise concatenated encoding matrix of diagnosis codes and blood tests into a lower-dimensional feature space, which is then augmented with a sinusoidal temporal encoding generated using positional encoding to incorporate the temporal order of time buckets.
Training and Performance Metrics
The model was trained using a leave-one-site-out (LOSO) strategy for development and performance evaluation, while 10-fold cross-validation was applied to the remaining data for training, validation, and hyperparameter tuning. To address data imbalance during training, the researchers employed two methods: first, a down-sampling strategy with a 1: 10 case-to-control ratio,
and second, Focal Loss as the objective for the model to focus on minority classes.
Performance was assessed using metrics such as Area Under the Receiver Operating Characteristic curve (AUROC), diagnostic odds ratio (DOR), and calibration analysis. The model showed good to excellent generalizable performance
across geographic settings, with AUROC values ranging from 0.760 to 0.837 depending on the prediction lead time (1-, 2-, or 3-year).
Biomedical Plausibility and Feature Importance
The model's learning is grounded in temporal patterns of healthcare interactions, specifically the acquisition of new and repeated diagnoses and blood test values.
The top 20 most important features, identified through Shapley additive explanation values, were largely consistent with known risk features from electronic medical records for pancreatic cancer. Notably, complete blood count parameters were particularly informative for pancreatic cancer risk,
alongside diagnostic codes associated with metabolic disorders (e.g., diabetes), health system interactions, and benign pancreatic conditions. The analysis also visualized the Overall importance and temporal contributions of all international classification of disease chapter codes and blood test types.
Population-Level Screening Feasibility
The study operationalizes the model as a high-sensitivity, population-level digital screen amongst those aged 50–74 years, where 90% of pancreatic cancers occur.
This involves a multi-stage Bayesian pipeline:
-
Stage 1: High-sensitivity digital screen using the PREMOD model.
-
Stage 2: High-specificity filter using non-invasive CT with the REDMOD AI algorithm, which is externally validated to detect visually occult pancreatic cancer with a
475 day lead-time.
-
Stage 3: Diagnostic confirmation via EUS and biopsy.
This pipeline improves screening efficiency by 8.2x compared to if EUS biopsy was used in an unenriched population,
aiming to screen 100,000 individuals and identify 21 cancers with a calculated Number Needed to Screen (NNS) of 406. The proposed implementation framework is designed to be more economically feasible
by including the ">60% of pancreas cancer patients without a diabetes signature."
Calibration and Portability
A key strength of the work is the use of standardized data inputs concordant with Observational Medical Outcomes Partnership (OMOP) Common Data Models,
which supports model transportability across health systems. Furthermore, a Bayesian update approach to calibration accuracy
allows for model portability across different incidence settings without retraining. This framework derives calibrated posterior odds by adjusting the raw logit output using an adjustment term, allowing the model to remain discriminatively robust (i.e., preserving AUROC) while providing context-specific probabilities suitable for diverse clinical environments.
The final conclusion is that this work lays the foundation for a first population-level digital enrichment tool to widen access to curative-intent management of pancreatic cancer.
The gist
Longitudinal sequences of coded diagnoses and blood test values accrued by patients throughout their clinical interactions were used to train a custom Transformer-based neural network with a multi-head attention mechanism to predict risk of pancreatic cancer with a multiyear lead time and risk-stratify populations for targeted screening.
How it works
- The input data, consisting of longitudinal diagnostic codes (ICD chapters) and blood test values, is first transformed into structured matrices using a
time-bucketed frequency encoding paradigm,
where each bucket represents a 30-day duration.
Improvements for AI systems
As a diligent AI researcher, I have analyzed this paper, Digitally enriching a screening population for pancreatic cancer using routine blood-based measures and clinical histories,
which presents a novel Transformer-based neural network for predicting pancreatic cancer risk from longitudinal patient data.
Here are the specific improvements that can be made to existing AI systems and what those improved systems can achieve:
)Improved AI System Capabilities: Population-Level, Early, Multi-Modal Risk Stratification
The core improvement lies in moving beyond single-feature or static risk stratification toward a dynamic, population-level prediction system capable of identifying pre-diagnostic signatures years in advance.
- 20+ Year Predictive Lead Time for High-Risk Identification:
A current limitation is that most models predict risk within 1–3 years before diagnosis. This model achieves robust prediction up to a 3-year lead time, and the simulation framework demonstrates the potential to capture signals up to 5 years prior.
The improved system will be able to flag individuals for proactive intervention (e.g., intensive surveillance, specialized clinical trials) with a significantly extended window before overt disease manifestation, fundamentally shifting management from reactive diagnosis to predictive interception.
- Dynamic Risk Trajectory Visualization:
The model doesn't just output a single risk score; it generates continuous risk probability scores and individual risk trajectories over time (Figure 3).
The improved system can provide clinicians with a risk curve
for any patient, showing how their predicted risk evolves based on changes in blood biomarkers or clinical history. This allows for the identification of critical inflection points—moments where a trajectory shifts rapidly—enabling ultra-personalized, time-sensitive interventions.
- Multi-Modal Data Fusion Robustness:
The architecture successfully integrates two distinct data streams (longitudinal diagnosis codes and routine blood test values) using a Transformer encoder with multi-head self-attention. The supplementary results show that combining these modalities yields superior performance compared to analyzing them separately (0.829 AUROC vs 0.638).
The improved system will be a truly multi-modal
predictor, ensuring that risk assessments are not biased by the limitations of any single data source (e.g., relying solely on medical records or solely on lab results).
- Population Portability via Bayesian Recalibration:
The system utilizes a Bayesian post-hoc recalibration strategy to adjust raw model logits based on the target population's actual prevalence, allowing the model to be deployed in any new clinical setting without full retraining.
The improved system can be rapidly deployed across diverse healthcare systems (e.g., community clinics vs. quaternary centers) with minimal local fine-tuning, overcoming the data silo
problem prevalent in current medical AI applications and ensuring consistent performance regardless of where it is used.
- Optimized Screening Pipeline for Low-Prevalence Cancers:
The system is specifically tuned to operate as a high-sensitivity filter (Stage 1) feeding into a downstream diagnostic tool (like REDMOD for CT screening). The simulation shows this pipeline can achieve an 8.2x improvement in screening efficiency over traditional methods.
The improved system will be the foundation of an automated, AI-enabled triage pipeline:
a) It identifies the highest-risk subgroup in a general population (Stage 1).
b) It directs these individuals to a highly efficient, low-resource diagnostic modality (Stage 2, e.g., non-invasive CT).
c) Only the few candidates identified by Stage 2 proceed to expensive, invasive confirmation (Stage 3, EUS/biopsy), dramatically reducing unnecessary procedures while maximizing cancer detection yield.
- Biologically Plausible Feature Importance:
By utilizing SHAP analysis, the model identifies which specific features (e.g., complete blood count parameters like MCV or specific ICD chapters related to metabolic disorders) are driving a prediction at any given time point.
The improved system moves beyond a black box
score by providing actionable insights: Risk is increasing due to a 15% rise in MCV and the recurrence of ICD code X.
This level of explainability ensures clinical trust and guides targeted biomarker monitoring rather than just flagging an outcome.
Abstract
Earlier detection of pancreatic cancer is key to enabling wider access to curative treatment and reducing cancer deaths; however, screening is presently not viable. Latent digital indicators of pathology are evident in an individual's disease and blood test trajectories and may predict the development of pancreatic cancer. Longitudinal sequences of coded diagnoses and blood test values accrued by patients throughout their clinical interactions were used to train a custom Transformer-based neural network with a multi-head attention mechanism to predict risk of pancreatic cancer with a multi-year lead time and risk-stratify populations for targeted screening. Mayo Clinic Platform with trained model from Mayo Clinic Rochester validated at Mayo Clinic Arizona, Mayo Clinic Florida, and Mayo Clinic Health Systems. The cohort comprised 6,017 adults with pancreatic cancer and 177,081 controls (median age 75, 45% female) with median 12 years (interquartile range 6.9-16.2) of medical history prior to pancreatic cancer diagnosis. External validation via leave-one-site-out, out-of-sample testing predicting pancreatic cancer 1-, 2-, and 3-years prior to diagnosis demonstrated mean area under the receiver operating characteristic of 0.837 (95% confidence interval 0.827-0.848), 0.797 (95% confidence interval 0.782-0.813), and 0.760 (95% confidence interval 0.745-0.776), respectively. Estimated pancreatic cancer risks were well-calibrated (calibration plot slope 1.08, intercept of-0.077; Brier score 0.025), and a Bayesian population pancreatic cancer prevalence update allows estimated cancer risk outputs to be transportable across settings. At testing, a screening threshold of >3.3% risk of pancreatic cancer in 1-year offered a diagnostic odds ratio of 18.2. Our work therefore lays the foundation for a digital population-level risk enrichment tool that could widen access to curative-intent management.
Sources
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks