The Next Layer: Augmenting Foundation Models with Structure-Preserving and Attention-Guided Learning for Local Patches to Global Context Awareness in Computational Pathology

arXiv:2508.19914 · q-bio.QM, cs.AI, stat.ML · Submitted 2025-08-27 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "The Next Layer: Augmenting Foundation Models with Structure-Preserving and Attention-Guided Learning for Local Patches to Global Context Awareness in Computational Pathology".

Jane: The paper was written by Muhammad Waqas, Rukhmini Bandyopadhyay, Eman Showkatian, Amgad Muneer, Anas Zafar et al. from MD Anderson Cancer Center and Northwestern University and The University of Texas MD Anderson Cancer Center.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title and Authors: Tom: Jane, we've got a paper today that's going to make your head spin — it's called "The Next Layer: Augmenting Foundation Models with Structure-Preserving and Attention-Guided Learning for Local Patches to Global Context Awareness in Computational Pathology." That's a mouthful, but the idea behind it is actually pretty beautiful.

Jane: It really is, Tom. So this is a team from MD Anderson Cancer Center, led by Jia Wu, and they're tackling a huge problem in digital pathology. When a pathologist looks at a biopsy slide, they see the whole tissue architecture — how tumor cells sit next to blood vessels, where immune cells are clustering, how necrosis spreads. But when we train AI on these slides, we usually chop them into tiny patches and treat each patch as if it's isolated from its neighbors.

Tom: Right, and that's the gap they're trying to close. The foundation models — like the ones that extract features from these patches — are incredibly powerful, but they don't inherently understand where each patch sits in the bigger picture. So you get a model that can recognize a tumor cell but has no idea that it's sitting right at the invasive front, which is exactly the kind of context a pathologist would care about.

Jane: Exactly. And that's why the title uses that phrase "structure-preserving." They want to keep the spatial layout of the tissue intact while the model learns. It's like the difference between reading a single sentence from a novel versus reading the whole chapter — you need both to understand the story.

Tom: And the authors here are a big deal. This is a collaboration between imaging physicists, pathologists, and oncologists. They're not just computer scientists in a vacuum; they have people like John Heymach and Natalie Vokes from thoracic oncology, which means they're deeply connected to real clinical questions about lung cancer and treatment response.

Jane: That clinical grounding shows up in how they designed the whole thing. They didn't just build a model that works on one dataset. They tested it across seven cancer types for survival prediction and three for classification — that's over fourteen thousand whole slide images. That's not a toy experiment.

Tom: Fourteen thousand slides, Jane. And they're using three different foundation models as backbones, which is smart because it shows their method isn't tied to one specific feature extractor. It's like building a better engine that works in any car, not just one brand.

Jane: So the big promise here is that we can get AI to think more like a pathologist — seeing both the trees and the forest. And that could mean better predictions for which patients need aggressive treatment and which ones don't. That's the kind of impact that actually changes lives.

Tom: And that's just the title and the team. Wait until we get into what they actually built — the architecture has some clever tricks that I think are going to surprise you.

Jane: I'm ready. Let's dig into the summary and see how they made this work.

Summary: Tom: So Jane, let's talk about what EAGLE-Net actually does. The paper's summary lays it out pretty clearly — this is a multiple instance learning framework, which sounds complicated, but it's really about how you handle a whole slide image that's way too big for any AI model to process at once.

Jane: Right, so you chop the slide into thousands of small patches, extract features from each one using a foundation model, and then you need a way to combine all those patch-level features into a single prediction for the whole slide. That's what multiple instance learning does — it learns which patches matter most and aggregates them.

Tom: And the standard approaches do that with something called attention — the model learns to weight each patch by how relevant it is. But here's the problem: most of those methods treat patches as if they're scattered points on a map with no relationship to each other. They ignore the spatial layout entirely.

Jane: Which is biologically wrong, because the arrangement of cells matters. A tumor patch sitting next to a necrotic region tells you something different than a tumor patch sitting next to healthy lung tissue. So EAGLE-Net adds three key ingredients to fix this.

Tom: First, they have this Multi-scale Absolute Spatial Encoding — MASE for short. It takes the actual coordinates of each patch in the original slide and builds a three dee matrix that preserves the tissue structure. Then they run convolutions with different kernel sizes — small ones to capture fine detail, larger ones to capture broader context — and combine them.

Jane: So it's like having both a magnifying glass and a wide-angle lens at the same time. The small kernels see individual cells, the larger kernels see how those cells relate to their surroundings, and the model learns to use both.

Tom: Second ingredient: a neighborhood-aware loss. Instead of just focusing on the top-attended patches in isolation, they look at the local neighborhood around those patches — the patches right next to them — and make sure the model is learning from that whole region. Because in pathology, the interesting stuff often happens at boundaries, where tumor meets stroma, or where immune cells are infiltrating.

Jane: And that's biologically meaningful. A single patch might look ambiguous, but if you see it's surrounded by immune cells and necrosis, that tells you a lot about what's happening in that microenvironment.

Tom: Third, they add a background suppression loss. When you build that spatial matrix, you inevitably include some empty or non-tissue patches. They explicitly penalize the model for paying attention to those, so it doesn't waste its capacity on glass and white space.

Jane: And the results? They benchmarked this on survival prediction across six TCGA cancer types plus a CPTAC cohort, and on classification tasks for lung cancer subtyping and prostate cancer grading. EAGLE-Net consistently beat or matched the state-of-the-art methods like CLAM and TransMIL.

Tom: The numbers are solid — up to three percent higher classification accuracy, and top concordance indices in six out of seven cancer types for survival. But honestly, the more impressive part to me is the interpretability. The attention maps they generate are smooth and biologically coherent — they light up tumor regions, invasive fronts, necrosis, immune infiltration — the stuff pathologists actually look at.

Jane: That's the part that gets me excited. We're not just getting better predictions; we're getting models that show their work in a way that aligns with clinical reasoning. That's huge for trust and for eventual clinical adoption.

Tom: And we haven't even talked about how they handle patients with multiple slides, or how they validated across different foundation models. That's coming up.

Improvements: Tom: So Jane, we've covered the basics of EAGLE-Net, but the paper has some clever engineering details that really set it apart. Let's talk about the improvements they propose over existing methods.

Jane: One thing I love is their patient-level tissue packing. Most MIL methods process each slide independently, but a single patient often has multiple slides from different parts of the tumor. EAGLE-Net packs those slides together into one coherent canvas — rotating them to minimize footprint and concatenating them side by side.

Tom: That's clever because it gives the model a holistic view of the patient's disease. One slide might show the invasive front, another might show a lymphovascular nest. If you analyze them together, the attention mechanism can connect those dots and see the full histological spectrum. That's something no other method in their benchmark does.

Jane: And it's not just about performance — it's about capturing inter-slide heterogeneity. The paper makes the point that different tissue samples capture different facets of the tumor, and analyzing them in isolation misses that.

Tom: Another improvement is how they handle positional encoding. Standard vision transformers use fixed positional encodings designed for square images of fixed size. But whole slide images are huge and variable — you can't just resize them. Some methods try to reshape the patch grid, but that distorts the true spatial relationships. EAGLE-Net's MASE module learns positional encodings dynamically from the actual tissue structure.

Jane: And they use convolutions for that, which is smart because convolutions have this built-in spatial inductive bias — they naturally understand locality and hierarchy. The two-stage approach — first extracting multi-scale features with different kernel sizes, then combining them with a one times one convolution — lets the model capture both fine-grained cellular detail and broader tissue architecture.

Tom: They also share weights between the classification layer and the attention mechanism. That's a trick that forces the model to align what it's paying attention to with what it's predicting. It's not just a post-hoc explanation; the attention is directly tied to the decision-making process.

Jane: And the neighborhood-aware loss is another big improvement. Instead of just looking at the top-attended patches in isolation, they look at the local neighborhood around those patches. The paper shows that this helps the model focus on cohesive tumor microenvironments rather than scattered high-scoring tiles.

Tom: Which makes biological sense. In pathology, the interesting stuff happens at boundaries — tumor-stroma interfaces, immune infiltration zones. A single patch might look ambiguous, but if you see it's surrounded by immune cells and necrosis, that tells you a lot.

Jane: And they validate all of this across three different foundation models — REMEDIES, Uni-V1, and Uni2-h. That's important because it shows their improvements aren't tied to one specific feature extractor. The gains hold up regardless of the backbone.

Tom: The ablation studies are thorough too — they test removing the MASE module, varying the neighborhood size, adjusting the spatial radius, and removing the regularization losses. Each component contributes something. It's not just one trick carrying the whole thing.

Jane: And the interpretability analysis is really impressive. They had pathologists annotate three hundred lung cancer slides with seven different tissue types — tumor, stroma, immune, vessels, bronchi, necrosis, and lung. Then they compared where the model's attention went against those annotations.

Tom: EAGLE-Net allocated seventy-four point six percent of its attention to tumor regions — higher than CLAM or Gated-AbMIL — and it also paid more attention to necrosis and immune infiltration, which are clinically important. Meanwhile, it spent less attention on normal lung and vascular structures. That's exactly the kind of focus you'd want from a clinical tool.

Jane: They even used frequency-domain metrics — analyzing the Fourier transform of the tumor masks — to show that EAGLE-Net captures more biologically plausible tumor shapes. That's a level of rigor you don't usually see in pathology AI papers.

Tom: So the improvements aren't just about accuracy numbers. They're about building a model that thinks more like a pathologist and can explain itself in a way that clinicians trust.

Jane: And that's the bridge to real clinical impact. Let's wrap this up and think about what it means for the field.

Conclusion: Tom: Jane, we've spent this whole episode on "The Next Layer: Augmenting Foundation Models with Structure-Preserving and Attention-Guided Learning for Local Patches to Global Context Awareness in Computational Pathology" — and I think we've only scratched the surface.

Jane: We really have. This paper from MD Anderson is a big step forward for computational pathology. It takes the power of foundation models — those massive pre-trained feature extractors — and adds the spatial awareness and local context that's been missing.

Tom: The key innovations are worth repeating: the MASE module that preserves tissue structure through multi-scale convolutions, the neighborhood-aware loss that focuses on cohesive microenvironments, and the background suppression that keeps the model honest. Plus that patient-level tissue packing that gives a holistic view across multiple slides.

Jane: And the results speak for themselves — better survival prediction across six of seven cancer types, higher classification accuracy, and attention maps that align with expert pathologist annotations. The model highlights tumor regions, invasive fronts, necrosis, and immune infiltration — the things that actually matter clinically.

Tom: What excites me most is the generalizability. They tested it across three different foundation models and it held up. That means this isn't a one-trick pony — it's a framework that can ride on whatever foundation model comes next.

Jane: And the interpretability is a game-changer. When a model can show its work in a way that matches clinical reasoning, that builds trust. And trust is what we need for AI to actually get adopted in hospitals.

Tom: There are limitations, of course. The paper acknowledges the computational overhead of the spatial encoding, and the gains were less dramatic on biopsy datasets where spatial context is limited. But those are solvable problems.

Jane: The bigger picture here is that we're moving toward AI systems that don't just predict — they understand. They see the tumor as a complex ecosystem, not just a collection of pixels. That's the kind of thinking that could lead to better biomarker discovery, better patient stratification, and ultimately better treatment decisions.

Tom: And that's why this paper matters. It's not just another incremental improvement in accuracy. It's a step toward AI that can actually partner with pathologists in the clinic.

Jane: So we say goodbye to EAGLE-Net and this paper — but I have a feeling we'll be hearing more about this approach. The code is going to be available, and I suspect a lot of researchers are going to build on this.

Tom: Absolutely. Thanks for joining us, everyone. We'll be back next time with another paper that's pushing the boundaries of what AI can do in medicine.

Jane: Until then, keep questioning, keep exploring, and keep looking at the whole picture — not just the patches.

Muhammad Waqas, Rukhmini Bandyopadhyay, Eman Showkatian, Amgad Muneer, Anas Zafar, Frank Rojas Alvarez, Maricel Corredor Marin, Wentao Li, David Jaffray, Cara Haymaker, John Heymach, Natalie I Vokes, Luisa Maren Solis Soto, Jianjun Zhang, Jia Wu

MD Anderson Cancer Center · Northwestern University · The University of Texas MD Anderson Cancer Center · The University of Texas MD Anderson Cancer Center · The University of Texas MD Anderson Cancer Center · The University of Texas MD Anderson Cancer Center

q-bio.QM, cs.AI, stat.ML

Submitted: 2025-08-27

Updated: 2026-08-17

Comments: 43 pages, 7 main Figures, 8 Extended Data Figures

Code: https://github.com/WuLabMDA/EAGLE-Net

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 78/100

Key concepts

Foundation Models
These are powerful pre-trained feature extractors used in AI. The paper tests EAGLE-Net across three different foundation models, showing that the proposed method's improvements are generalizable and not tied to a single specific feature extractor.
EAGLE-Net
This is a multiple instance learning framework designed to process whole slide images by combining features from thousands of small patches. It uses three key ingredients: Multi-scale Absolute Spatial Encoding (MASE) for structure, neighborhood-aware loss for local context, and background suppression loss.
Structure-Preserving
This refers to the method's goal to keep the spatial layout of tissue intact while the model learns. It uses Multi-scale Absolute Spatial Encoding (MASE) which builds a three-dimensional matrix based on patch coordinates, allowing the model to understand where each patch sits in relation to others.
Interpretability
The paper's attention maps are smooth and biologically coherent, lighting up tumor regions, invasive fronts, necrosis, and immune infiltration. This allows the model to show its reasoning in a way that aligns with what pathologists actually look for clinically.

Terminology

Summary

Summary

This paper introduces EAGLE-Net, a multiple instance learning (MIL) framework designed to augment foundation models in computational pathology by integrating global tissue structure and local microenvironment context. The authors state: We present EAGLE-Net, a structure-preserving, attention-guided MIL architecture designed to augment prediction and interpretability.

The framework comprises four key components: "(i) Tiling and feature extraction, in which tissue patches are extracted and embedded through the pretrained foundation model; (ii) Multi-scale Absolute Positional Encoding (MASE) block; (iii) Attention pooling; and (iv) Neighborhood-aware and background-suppression loss terms. The MASE module aims to simultaneously learns patch-level information and global tissue structure using absolute positional encoding, employing a two-stage convolutional approach to learn absolute positional encodings of tissue patches for every slide. This uses convolution kernels of sizes 1×1, 3×3, 5×5, and 7×7 to capture spatial context at multiple scales, where Small kernels encode fine spatial details, such as cellular structures, while larger kernels incorporate broader spatial context, like surrounding blood vessels and tissues."

The authors propose a top-K neighborhood-aware loss that incorporates top attended instances and their connecting local neighborhood to bag-level loss function, allowing the model to self-guide on clinically relevant niches. They also introduce a background suppression loss term that penalizes the attention weights of the background patches. The total loss is computed as Loss = L1 + lambdaL2 + betaL3, where L1 is the task-specific bag-level loss, L2 is the neighborhood-aware loss, and L3 is the background suppression term.

For patient-level analysis, the authors implement a tissue packing step that packs tissue samples from multiple slides into one coherent canvas, enabling a holistic view across multiple tissue samples of same patient, capturing inter-slide heterogeneity that would otherwise be neglected.

The model was evaluated on seven prognostic and three diagnostic tasks of totally 14,432 WSIs from multiple institutions and scanners in 7 distinct cancer types. For survival prediction, they used six datasets sourced from The Cancer Genome Atlas (TCGA) and the Clinical Proteomic Tumor Analysis Consortium (CPTAC) including LUSC, LUAD, STAD, UCEC, THCA, KIRC, and CPTAC-LUAD, totaling 4,172 slides from 2,956 patients. For classification, they used TCGA and CPTAC lung cancer subtyping and the PANDA dataset for ISUP prostate cancer grading.

Results show that EAGLE-Net consistently achieved improved or comparable prognostic performance compared to benchmarked algorithms. Specifically, in TCGA-KIRC, EAGLE-Net achieved a C-index of 0.708 ± 0.018, surpassing CLAM (0.668 ± 0.026), Gated-AbMIL (0.693 ± 0.008), and AbMIL (0.687 ± 0.010). For TCGA-LUAD, it scored 0.672 ± 0.018, ahead of CLAM (0.662 ± 0.012) and TransMIL (0.650 ± 0.016). On CPTAC-LUAD, EAGLE-Net achieved 0.723 ± 0.052, outperforming all other models.

For classification, EAGLE-Net achieved Accuracy of 0.980 on TCGA and 0.916 on CPTAC for NSCLC subtyping, compared to CLAM's 0.980 and 0.889, and TransMIL's 0.972 and 0.892. On the PANDA dataset, EAGLE-Net achieved a Cohan's kappa of 0.984 on the Radboud cohort and 0.985 on the Karolinska cohort.

The authors tested generalizability across three foundation models—REMEDIES, Uni-V1, and Uni2-h—finding that EAGLE-Net consistently achieved the highest C-index across all three backbones and cancer types. Performance gains ranged from 0.1–3.9% over the next-best model with REMEDIES, 0.2–1.7% with Uni-V1, and 0.3–1.1% with Uni2-h.

For interpretability, the authors compared attention heatmaps against pathologist annotations on 300 TCGA-LUAD slides covering seven distinct biological regions—Tumor, Stroma, Immune, Vessel, Bronchi, Necrosis, and Lung. EAGLE-Net achieved the highest Dice score (0.56) and the lowest FPR (0.101) compared to other attention-based methods. In frequency-domain analysis, EAGLE-Net achieved a lower radial difference of 0.147 and a lower AED Difference of 0.316. Statistical testing with two-sided Wilcoxon signed-rank tests with Bonferroni correction showed statistically significant improvement in Dice score over CLAM and in FDR over both CLAM and Gated-AbMIL.

Attention allocation analysis showed that EAGLE-Net allocates 74.6% of its attention to tumor regions, exceeding CLAM (73.5%) and Gated-AbMIL (72.8%), while reducing attention to normal lung regions (6.10% vs. 8.20% for CLAM) and vascular structures (2.93% vs. 3.39% for CLAM).

The authors conclude that EAGLE-Net provides a generalizable, interpretable framework that complements foundation models, enabling improved biomarker discovery, prognostic modeling, and clinical decision support. They note limitations including increased computational overhead during training from the MASE module and less pronounced performance gains in biopsy datasets like PANDA.

Improvements for AI systems

Based on the EAGLE-Net paper, here are the specific improvements I can implement in an AI system and what the improved system can do:


Implementation: Add a two-stage convolutional positional encoding layer that transforms 2D patch feature matrices into spatially enriched representations using kernels of sizes 1×1, 3×3, 5×5, and 7×7, followed by a 1×1 convolution to fuse multi-scale spatial context.

What the improved system can do:

  • Preserve global tissue architecture without distortion, unlike fixed sinusoidal encodings or pyramid-based methods that break spatial relationships

  • Capture both fine-grained cellular details (small kernels) and broader tissue context (large kernels) simultaneously

  • Handle variable-length WSI inputs without sequence-length constraints

Implementation: Add a loss term that identifies the top-K highest-attention patches and computes cross-entropy loss on their surrounding neighborhood patches (within a configurable radius), weighted by their attention scores.

Implementation: Add a regularization term that penalizes attention weights assigned to non-tissue patches, using an indicator vector that marks tissue vs. background regions.

Implementation: Implement a greedy rotation-and-concatenation algorithm that packs multiple WSIs from the same patient into a single coherent canvas, preserving within-slide micro-architecture while enabling cross-slide analysis.

Implementation: Share weights between the classification layer and attention computation, so that the same learned weights determine both slide-level prediction and patch-level saliency.

  • 3% higher classification accuracy on lung cancer subtyping (TCGA: 98.0%, CPTAC: 91.6%)

  • Top concordance indices in 6 of 7 cancer types for survival prediction (C-index range: 0.685–0.743)

  • Cohen's kappa of 0.984–0.985 on multiclass prostate cancer grading (PANDA dataset)

  • Smooth, biologically coherent attention maps that align with expert annotations

  • 74.6% attention allocated to tumor regions (vs. 73.5% for CLAM, 72.8% for Gated-AbMIL)

  • Lower false-positive rate (0.101) in tumor boundary detection

  • Improved frequency-domain metrics (radial difference: 0.147, angular energy dispersion difference: 0.316), indicating better capture of invasive/irregular tumor contours

  • Consistent performance across three foundation models (REMEDIES, Uni-V1, Uni2-h), with gains of 0.1–3.9% over next-best baselines

  • Works on both surgical specimens and biopsies, though gains are more pronounced on large resections where spatial context is richer

  • Compatible with any existing histology foundation model without architectural changes

  • Identifies invasive fronts, necrosis, and immune infiltration without explicit ROI annotations

  • Reduces attention to benign tissue (normal lung: 6.10% vs. 8.20% for CLAM), sharpening focus on diagnostically relevant compartments

  • Provides explainable predictions that pathologists can visually verify, supporting regulatory adoption


def MASE(features 2d, kernel sizes=[1,3,5,7]):

latent = Linear(feature dim, d latent)(features 2d) # reduce dim

outputs = []

for k in kernel sizes:

pad = (k-1)//2

conv out = Conv2d(k, padding=pad)(latent) # [rows, cols, d latent]

outputs.append(conv out)

stacked = Concat(outputs, dim=-1) # [rows, cols, 4*d latent]

positional = Conv2d(1x1)(stacked) # [rows, cols, d latent]

enriched = latent + positional # residual connection

return enriched.flatten(0,1) # [n patches, d latent]

def neighborhood loss(attention scores, features 2d, top k=10, radius=2):

top indices = topk(attention scores, k=top k)

loss = 0

for idx in top indices:

row, col = unravel index(idx, features 2d.shape[:2])

neighbors = get neighbors(row, col, radius) # square window

for n in neighbors:

patch logits = classification head(features 2d[n])

loss += cross entropy(patch logits, bag label) * attention scores[idx]

return loss / (top k * len(neighbors))

def background loss(attention scores, tissue mask):

return sum(attention scores * (1 - tissue mask)) # penalize non-tissue attention

The improved system is a drop-in replacement for existing MIL frameworks in computational pathology, requiring only patch coordinates and foundation model embeddings as input. It delivers state-of-the-art performance while producing attention maps that are biologically meaningful and clinically actionable.

Sources

Related papers