Enhancing In-Hospital Mortality Prediction Using Multi-Representational Learning with LLM-Generated Expert Summaries
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Enhancing In-Hospital Mortality Prediction Using Multi-Representational Learning with LLM-Generated Expert Summaries".
Jane: The paper was written by Harshavardhan Battula, Jiacheng Liu and Jaideep Srivastava from University of Minnesota Twin Cities.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the channel, everyone. Today we're digging into a paper that's been making the rounds — "Enhancing In-Hospital Mortality Prediction Using Multi-Representational Learning with LLM-Generated Expert Summaries." Jane, this title is a mouthful, but the idea underneath is pretty straightforward, right?
Jane: Absolutely, Tom. So picture this — you're in an ICU, and the team has two sources of information about a patient. You've got the numbers — heart rate, blood pressure, oxygen levels — and you've got the notes — what the doctors and nurses actually wrote down. This paper from the University of Minnesota is trying to combine both to predict whether a patient will survive their hospital stay.
Tom: And here's the twist — they're not just feeding the raw notes into the model. They're using a large language model to first summarize those notes into a structured expert-style summary, and then they feed that summary in. Like having a super-smart resident read all the charts and give you the highlights before you make a call.
Jane: Exactly. And the authors — Harshavardhan Battula, Jiacheng Liu, and Jaideep Srivastava — they wanted to know something really specific. When the summary helps improve prediction, is it because the LLM is adding new medical knowledge, or is it just reorganizing what was already in the notes into a form the model can actually use?
Tom: That's the million-dollar question, and honestly, it's a question most papers in this space just skip. They show the summary helps, they report the numbers, and they move on. This team actually built experiments to separate those two possibilities.
Jane: Right, and we should say — they used MIMIC-III, which is this massive public ICU database. Over nineteen thousand patient stays, with about twelve point eight percent mortality. And they were really careful about something called leakage — making sure the notes and data they used only came from the first forty-eight hours, before the prediction point, so the model isn't accidentally peeking at the future.
Tom: Yeah, that leakage control is huge. A lot of earlier work in this area has been sloppy about it. If a patient dies on day three, and the notes from day two describe the deterioration, the model might just be reading the outcome off the page. This team excluded those cases and made sure any notes charted after hour forty-eight were thrown out.
Jane: And that's what makes their headline numbers credible — the full model hit an AUPRC of about zero point five zero, which is a big jump from the physiology-only baseline of zero point three six. But the real story, Tom, is what happens when they start pulling the summary apart to see where that gain comes from.
Tom: And that's exactly where we're headed next. Stay with us.
Abstract: Jane: So we're back with "Enhancing In-Hospital Mortality Prediction Using Multi-Representational Learning with LLM-Generated Expert Summaries," and we've set the stage — they've got physiology, they've got notes, they've got LLM-generated summaries. Now let's talk about what the abstract actually claims, because it's subtle.
Tom: The abstract says the gain is real, but the mechanism is not what you'd hope. They found that when you take the summary embedding — that's the numerical representation of the summary — about forty-one percent of its variance can be predicted from the notes alone using a simple linear model. So a big chunk of what the summary says is just a re-encoding of the notes.
Jane: Right, and that's the redundancy check. But then they did something clever — they subtracted out that predictable part, leaving what they call a "note-orthogonal residual." That's the part of the summary that genuinely cannot be derived from the notes. And when they fed that residual into the model instead of the full summary, it still improved prediction, but only by about twenty-eight percent of the original gain.
Tom: So the headline is — the summary helps, but most of the help comes from reorganizing information the notes already had, not from the LLM injecting new medical knowledge. That's a really important finding for anyone building these systems.
Meng: Can I jump in here? I'm Meng, by the way — I work on deploying models in production. The practical question I have is — does that mean the LLM is just a fancy compression algorithm? Because if so, maybe a simpler text summarizer would do the same job for a fraction of the compute cost.
Jane: That's a great point, Meng. And the paper actually supports that reading. The summary is about two hundred sixty-four words on average, while the raw notes span thousands of characters across multiple documents. The encoder — Bio ClinicalBERT — can only handle five hundred twelve tokens per note, and they mean-pool across notes. So the LLM is essentially condensing a huge, messy record into a compact form the encoder can actually digest.
Lu: But I'd push back slightly. The residual still carries real signal — twenty-eight percent of the gain is not nothing. And the patient-shuffled control — where they randomly reassign summaries to patients — actually hurt performance. So the summary is not just a scale artifact. It's patient-specific. The LLM is doing something, even if most of it is reformatting.
Tom: And that's the nuance, right? The paper doesn't say the LLM is useless. It says the LLM is valuable, but for a different reason than people assume. It's not a knowledge oracle — it's a representation-prep step. It takes a lossy channel and feeds it a better-prepared input.
Jane: And that reframing matters because it tells researchers where to focus. If the bottleneck is the narrative-to-encoder interface, then improving that interface — better pooling, better summarization, maybe a different encoder — is where the next gains will come from.
Lu: Exactly. And it also warns against the hype cycle where people assume LLMs are adding hidden medical expertise. They might just be really good at cleaning up text. Which is useful, but it's a different claim.
Meng: So if I'm building this in a hospital, I should budget for the LLM as a preprocessing step, not as a clinical reasoning engine. That changes the risk profile too — you're not relying on the LLM to be right about medicine, just to be good at summarizing.
Tom: And that's a much safer deployment story. We'll get into the actual ablation results and the numbers in more detail next.
Improvements: Tom: Alright, we're deep into "Enhancing In-Hospital Mortality Prediction Using Multi-Representational Learning with LLM-Generated Expert Summaries" now, and I want to get into the improvements — not just what they found, but what they suggest the field should do differently.
Jane: Right, and the biggest suggestion is about where to put your effort. The paper's core argument is that the LLM's contribution is mostly reorganization, not new knowledge. So the authors are saying — stop treating the LLM as a black box that adds clinical wisdom, and start treating it as part of the representation pipeline.
Lu: And that's a genuinely useful redirect. Because if you believe the LLM is adding knowledge, you'd invest in better prompts, bigger models, maybe even fine-tuning on medical corpora. But if the gain is about compression and structure, then the real lever is the encoder — how you turn text into vectors. The paper shows the note branch underperforms its own inputs — a simple logistic probe on the raw notes got AUROC zero point eight one, while the trained notes-only network only got zero point seven five. So the encoder is the weak link.
Meng: That's a striking number, actually. The notes contain more signal than the network is extracting. So the improvement isn't just about the summary — it's about the fact that the summary happens to be in a format the encoder can handle. If you fixed the encoder, you might not need the LLM at all.
Jane: And that's the actionable insight. The authors even suggest that a better text encoder, or a better pooling strategy, could capture what the LLM is providing. The summary is a workaround for a bottleneck, not a fundamental addition.
Tom: But let's not undersell the residual. The note-orthogonal part still gave a real boost — about zero point zero two six AUPRC over the physiology-plus-notes baseline. And the patient-shuffled control dropped performance, which means the summary is tied to the actual patient, not just adding noise or scale.
Lu: Right, and that's the part that keeps the LLM in the picture. There's something in the summary that the notes don't linearly contain — maybe the LLM is re-weighting important details, or connecting pieces across notes that the mean-pooling destroys. The paper can't fully separate that, and they're honest about it.
Meng: And they're also honest about the limitations — the ridge regression they used is linear, so forty-one percent redundancy is a lower bound. A non-linear map could explain more. And the residual isn't perfectly orthogonal to the notes, so some of that twenty-eight percent gain might still be note content the ridge missed.
Jane: That's the thing I appreciate about this paper — they don't overclaim. They give you the numbers, they tell you where the uncertainty is, and they explicitly say the residual gain is a lower bound on the summary's value, while the redundancy is also a lower bound. So the truth is somewhere in between.
Tom: And the practical improvement they're advocating is — redesign the interface between clinical narrative and the model. Maybe that means hierarchical encoding, maybe it means training a summarizer specifically for this task, maybe it means a different pooling scheme. But the target is clear.
Lu: And I'd add — the leakage controls they used should become standard. Excluding deaths inside the window, excluding notes charted after hour forty-eight checking DEATHTIME — these are the kind of details that make results trustworthy. If more papers did this, the field would have fewer inflated numbers.
Meng: Agreed. And for deployment, the fact that they self-hosted the LLM on local GPUs is a big deal for privacy. No patient text left their environment. That's a template for how to do this responsibly.
Jane: So the improvements here are both technical and methodological — better encoding, better evaluation, better privacy practices. We'll wrap up with the big picture next.
Conclusion: Tom: We're closing out our discussion of "Enhancing In-Hospital Mortality Prediction Using Multi-Representational Learning with LLM-Generated Expert Summaries," and I want to step back and think about what this means for the world.
Jane: For me, the big takeaway is that the paper gives us a more honest picture of what LLMs are doing in clinical prediction. They're not magic — they're really good at taking messy, scattered information and making it usable. That's valuable, but it's a different value than we've been told.
Lu: And that honesty is actually empowering. If the bottleneck is the encoder, then we know what to fix. If the bottleneck were the LLM's knowledge, we'd be stuck waiting for bigger models. Instead, we can work on better ways to read clinical text — and that's a problem we know how to solve.
Meng: From my side, the practical impact is clear. Hospitals can adopt this framework with a clear understanding of what it does and doesn't do. The LLM is a preprocessing step, not a clinical oracle. That lowers the risk and makes the deployment case much cleaner.
Tom: And the patient-shuffled control — that's the detail I keep coming back to. It proves the summary is tied to the actual patient. It's not just adding generic medical language that happens to correlate with outcomes. The summary is about this person, in this bed, at this time.
Jane: Exactly. And that's why the residual signal, even at twenty-eight percent, matters. It means the LLM is capturing something about the individual case that the raw notes, as currently encoded, are losing. Maybe it's the way the LLM connects a lab value to a note written six hours earlier. Maybe it's the gestalt of the trajectory.
Lu: And the authors are careful to say they can't fully separate those mechanisms. But they've built the tools to ask the question, and that's the contribution — a methodology for measuring redundancy and complementarity in LLM-generated representations. That's going to be useful far beyond this one task.
Meng: And the leakage controls — I hope those become the floor, not the ceiling. If you're going to publish numbers on MIMIC-III, you should be doing what they did. Otherwise, you're just reporting artifacts.
Tom: So where does that leave us? The paper says — use LLMs to prepare representations, not to reason about patients. And measure what they're actually contributing, instead of assuming it's knowledge. That's a mature, grounded way to build clinical AI.
Jane: And it's a good note to end on. We've got the numbers, we've got the mechanism, we've got the limitations. That's a complete piece of work. Thanks to Battula, Liu, and Srivastava for putting it together.
Tom: And thanks to all of you for listening. We'll be back with the next paper soon — until then, keep reading, keep questioning, and keep the signal separate from the noise.
Harshavardhan Battula, Jiacheng Liu, Jaideep Srivastava
University of Minnesota Twin Cities
cs.CL, cs.AI, cs.LG
Submitted: 2026-08-13
Updated: 2026-08-14
Code: https://github.com/HarBatt/MultiRep-IHM-Prediction
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 70/100
Key concepts
- Multi-Representational Learning
- This approach combines two types of patient information: the raw physiological data (like heart rate and blood pressure) and the unstructured clinical notes. The goal is to use both sources together to improve the accuracy of predicting a patient's survival.
- LLM-Generated Expert Summaries
- A large language model (LLM) is used to read extensive, messy patient notes and condense them into a structured, expert-style summary. This summary is then fed into the prediction model to help capture key information efficiently.
- Leakage Control
- The study strictly controls for data leakage by ensuring the model only uses patient notes and data from the first 48 hours, preventing it from accidentally seeing information that would only be available after a patient's death.
Terminology
Summary
Summary
This paper evaluates a multi-representational framework for predicting in-hospital mortality (IHM) in intensive care unit (ICU) patients, in which large language model (LLM)-generated expert summaries of clinical notes are fused with physiological time-series data. The authors quantify how much of the resulting predictive gain is non-redundant with the information already contained in the clinical notes themselves.
Objective: To evaluate the multi-representational framework and determine how much of the gain from LLM-generated expert summaries is non-redundant with the notes they were generated from.
Data and Cohort: Using MIMIC-III, the analytic cohort comprises 19,211 first ICU stays (12.83% mortality, 2,465 deaths). Inclusion criteria: first ICU stay per patient, age ≥18, ICU length of stay exceeding the 48-hour observation window, survival beyond the prediction point, at least one usable in-window note, and at least one in-window physiological measurement. Three leakage controls were applied: (1) stays with death inside the 48-hour window were excluded; (2) positive labels with null DEATHTIME were excluded; (3) notes whose STORETIME fell after hour 48 were excluded. The cohort has median age 66.9 years, 56.2% male, 70.9% White, median ICU length of stay 3.98 days. Splits were stratified 60/20/20 (11,526/3,842/3,843), patient-disjoint.
Representations: Three branches were constructed. (1) Physiology: ten physiological variables (blood pressures, heart rate, temperature, respiratory rate, SpO2, FiO2, pH, glucose) extracted from 61 item IDs following the Harutyunyan et al. benchmark, averaged within each of 48 hours, with missing hours forward-filled and imputed with benchmark normal values plus a binary mask; encoded by a single-layer LSTM to a 256-dimensional vector HT. (2) Clinical notes: 200,145 in-window notes (10.4 per stay) encoded independently with Bio ClinicalBERT and mean-pooled (768-d); the 32 most recent notes per stay were encoded; representation applies exponential temporal decay with λ tuned on validation AUPRC, with λ=0 selected. (3) Expert summary: all in-window notes concatenated chronologically, truncated to 4,096 ClinicalBERT tokens, summarized by deepseek.v3.2 (self-hosted on eight NVIDIA H100 GPUs) under six fixed headings (presentation, active problems, trajectory, organ support, concerning findings, clinical gestalt), constrained to use only information in the notes, not to state/predict/imply survival or death, temperature 0, 700-token cap, median summary length 264 words; encoded by the same encoder to give V. Fusion concatenates enabled branches (1,792 → 1 for full model) with a single linear layer, dropout 0.3, Adam (lr 1e-4, weight decay 1e-5), batch 32, gradient clipping 5.0, up to 50 epochs with early stopping on validation loss, seed 42.
Results: The full multi-representational model (physiology + notes + standardized summary) reached AUPRC 0.4977 and AUROC 0.8429, 37.3% and 8.5% above the physiology-only baseline (AUPRC 0.3625, AUROC 0.7770). Clinical notes alone were the weakest configuration (AUPRC 0.3010, AUROC 0.7473); expert summaries alone matched physiology on AUPRC (0.3626) and were numerically higher on AUROC (0.7969). Standardizing the expert embedding on training statistics was worth +0.0427 AUPRC and +0.0198 AUROC over unstandardized fusion, and was also superior on validation.
Redundancy Analysis: (1) Geometry: matched pairs (Vi, Ui) had mean cosine 0.2734 (SD 0.1322) versus 0.0025 for mismatched pairs; retrieval recovered the correct stay at rank 1 in 7.1% of cases (chance 0.03%) and within top 5 in 16.2% (chance 0.13%), median rank 82; linear CKA was 0.5621 and leading canonical correlation 0.927, with only 2 of 50 directions exceeding 0.9. (2) Recoverability: ridge regression from UT to V achieved test R2 = 0.4078; adding physiological summary statistics raised this only to 0.4120. The cross-fitted residual retained 77.0% of the norm of centered V. (3) Probes: L2-regularized logistic probes showed V alone at AUROC 0.8293/AUPRC 0.4788; the note-explainable part V̂ probed at AUROC 0.8157/AUPRC 0.4228; the note-orthogonal residual V⊥ reached AUROC 0.6906/AUPRC 0.2565; concatenating [UT; V] improved on V alone by only 0.005 AUROC and 0.009 AUPRC. (4) Retrained ablations: substituting V⊥ for the expert branch improved AUPRC over the physiology-plus-notes reference by 0.0258 (95% CI 0.005–0.047), which is 28% of the full summary gain (+0.0928 AUPRC); the AUROC delta for the residual arm included zero. The patient-shuffled control degraded the reference by 0.0219 AUPRC, confirming the contribution is patient-specific.
Conclusion: The framework's core proposition holds: encoding a structured summary yields a more useful representation than encoding the notes directly, and adds value over physiology and notes together. However, the mechanism is predominantly reorganization of information the notes already contain into a form the encoder and linear head can use, rather than knowledge the LLM supplies. The authors state: The summary's contribution is predominantly reorganization of information the notes already contain into a form the encoder and linear head can use.
They conclude that framing the LLM as a representation-preparation step rather than a knowledge source is more defensible and more actionable, identifying the narrative-to-encoder interface as the component worth optimizing.
Limitations: Recoverability was tested with a linear map, so 40.8% is a lower bound on redundancy; α was selected for predictive R2 rather than orthogonality, so the residual is not exactly orthogonal to UT; all arms were trained once at a fixed seed, so intervals reflect test-set sampling variance only; a logistic probe on UT reached AUROC 0.8118, well above the trained notes-only network's 0.7473, so the note branch under-performs its own inputs; input budgets are asymmetric (32 notes at 512 tokens versus a 4,096-token whole-record view); calibration was not assessed; a single locally hosted LLM was used at temperature 0 without auditing compliance with the no-prognostication constraint; stays excluded for lack of a usable note had higher mortality (19.1% vs 12.8%); subgroup metrics carry no intervals; and this is one institution and era with concatenative fusion precluding per-modality attribution.
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems, along with what the improved system can do:
Improvement: Before fusing LLM-generated summaries with raw notes and physiology, the system explicitly computes the linear recoverability of the summary embedding from the note embedding (e.g., ridge regression R squared) and separates the summary into a note-explainable component and a note-orthogonal residual (V). The fusion layer then weights these components differently, rather than treating the summary as a single homogeneous input.
What the improved system can do:
-
Automatically detect when an LLM summary is mostly rephrasing the notes (high R squared) and down-weight that redundant portion, preventing overfitting and wasted capacity.
-
In cases where the residual carries genuine new signal (low R squared), the system will amplify that component, improving precision-recall performance (AUPRC) on imbalanced outcomes like mortality.
-
Provide a per-patient interpretability score: “how much of this prediction came from note-reorganization vs. genuinely new information.”
Improvement: Before using a summary embedding V in prediction, the system runs a fast retrieval check: compute cosine similarity between V and the corresponding note embedding U, and compare it against a threshold calibrated on training data. If the similarity is anomalously low (e.g., below the 5th percentile of matched pairs), flag the summary as potentially misaligned (e.g., wrong patient, corrupted text, or LLM hallucination) and fall back to a notes-only or physiology-only prediction.
Improvement: The paper found that a fixed exponential decay (lambda = 0) was optimal, but this is an artifact of the normalization scheme. I will replace the mean-pooling over notes with a learned attention mechanism that weights each note by its predicted relevance to the outcome, conditioned on the note’s category, time offset, and the patient’s current physiology state. This attention is trained jointly with the fusion head.
Improvement: The paper explicitly notes that calibration was not assessed. I will add a temperature-scaling or isotonic-regression calibration layer specifically for the summary branch’s contribution, trained on validation data. This layer is applied after fusion, but it is conditioned on the norm of V (the note-orthogonal residual), because the paper shows that residual carries the most novel signal and is most prone to miscalibration.
Improvement: The paper’s single-seed results are a limitation. I will train the full fusion model with 10 different seeds and report the mean and 95% confidence interval for AUROC and AUPRC. Additionally, I will use the variance across seeds to compute a “stability score” for each patient subgroup, flagging subgroups where predictions are highly seed-dependent.
Improvement: The paper admits it did not audit whether the LLM violated the “no prognostication” constraint. I will add a lightweight classifier (e.g., a fine-tuned Bio ClinicalBERT) that scans each generated summary for outcome-directed language (e.g., “likely to die,” “poor prognosis,” “will not survive”). If the classifier flags the summary, the system either regenerates it with a stricter prompt or excludes that summary from the fusion and falls back to notes-only.
Improvement: The paper uses a single linear layer (1,792 → 1) for fusion. I will replace this with a two-layer MLP (1,792 → 256 → 1) with ReLU and dropout, while keeping the same training protocol. This allows the model to learn non-linear interactions between physiology, notes, and the summary residual, which the paper’s own probe results suggest exist (the linear probe on [U;V] barely improved over V alone, but a non-linear head may capture more).
Improvement: The paper uses a fixed budget (32 notes, 4,096 tokens for the LLM). I will make this adaptive: the system first encodes all notes with a fast, cheap encoder (e.g., TF-IDF or a small BiLSTM) to estimate each note’s relevance, then selects the top-K notes (K dynamically chosen per patient) to feed to the expensive LLM and the clinical BERT encoder. This reduces compute and allows the LLM to see a longer, more relevant context for patients with many notes.
Improvement: The paper reports that the Black/African American subgroup gained only +0.004 AUROC, and the Hispanic subgroup had chance-level AUROC (0.532). I will add a loss reweighting term that up-weights underrepresented subgroups during training, and I will monitor subgroup AUROC/AUPRC at every epoch, early-stopping if any subgroup’s performance degrades by more than 5% relative to the majority group.
Improvement: Instead of only using the full summary embedding V, I will explicitly compute V (the residual after ridge regression from notes) and feed it as a separate input channel to the fusion head, alongside V and the notes. This forces the model to learn to use the residual directly, rather than having to discover it implicitly.
Summary of what the improved AI system can do overall:
It will predict in-hospital mortality with higher AUPRC (target: >0.52 vs. 0.4977 in the paper), better calibration (Brier score <0.10), lower variance across seeds, and explicit fairness guarantees. It will also be more computationally efficient, more robust to data pipeline errors, and able to explain to clinicians why a prediction was made (note-reorganization vs. new information). Most importantly, it will be safe for prospective deployment because it audits LLM compliance, checks patient-summary alignment, and reports uncertainty honestly.
Abstract
To evaluate a multi-representational framework in which large language model (LLM)-generated expert summaries of intensive care unit (ICU) notes are fused with physiology for in-hospital mortality (IHM) prediction, and to determine how much of the resulting gain is non-redundant with the notes themselves. Using MIMIC-III (19,211 first ICU stays, 12.83% mortality), we encoded 48-hour physiology, clinical notes, and LLM summaries generated under a prompt forbidding prognostication, then fused them. Redundancy was assessed by ridge recoverability, linear probes, and retrained ablations substituting a note-orthogonal residual or a patient-shuffled summary embedding. On 3,843 held-out stays, fusion reached AUPRC 0.4977/AUROC 0.8429 versus 0.3625/0.7770 for physiology alone. Ridge regression from note embeddings explained 40.8% of summary-embedding variance. The note-orthogonal residual retained a minority of the gain (+0.0258 AUPRC, 95% CI 0.005--0.047; 28%), while patient-shuffled summaries fell below the reference. Summaries improve prediction patient-specifically, but predominantly by reorganizing information the notes already contain.
Sources
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering