Anatomy Contextualized Adaption of CT Foundation Models

arXiv:2607.27154 · cs.CV, cs.AI · Submitted 2026-08-19 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Anatomy Contextualized Adaption of CT Foundation Models".

Jane: The paper was written by R. Kenia et al. from.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Summary: Tom: Now we're looking at how they actually built this system—the methodology behind Anatomy Contextualized Adaptation of CT Foundation Models. They aren't just applying a simple filter; they’ve designed a multi-layered approach to enforce structural reasoning.

Jane: The core idea is that the general model provides raw power, but then these specific adaptation modules refine that output based on anatomical location or type of finding. It’s like giving the system a specialized lens for every part of the body.

Lu: What I found most compelling is how they are integrating spatial reasoning into this AI framework; it's not just looking at pixels, but modeling the relationships between those pixels and the known physical structure of the body as a coherent whole.

Meng: I was looking closely at their experimental setup, and it seems like they use a multi-stage training pipeline that forces the model to learn both local features—like a calcification spot—and global context simultaneously, which is technically quite demanding.

Lalam: And thinking about the impact on healthcare culture, this rigorous architectural approach suggests that AI won't just be an overlay; it will become an integrated, structurally knowledgeable partner in the medical team, improving both trust and efficacy globally.

Jane: It gives us confidence because they aren't just claiming "this works"; they are showing *how* it works by defining these precise adaptation points within the system. This makes the whole complex model feel much more robust and trustworthy to clinicians who rely on structural integrity.

Tom: Absolutely, robustness is paramount when patient care is involved. So, if we understand this mechanism, what specific improvements do they propose for the future of this technology? That’s where things get really exciting for us in the next segment.

Improvements Suggested: Jane: We're moving into the future state of this technology now! The paper doesn't just present a model; it suggests several ways to take this adaptive concept even further, which is incredibly helpful for industry and research.

Tom: One major improvement they point out relates to handling multi-modality data, meaning integrating CT scans with MRIs or ultrasounds through the same contextualized framework. That’s a massive increase in complexity and power.

Lu: And I love that they are pushing for greater interpretability; it's not enough for the AI to give an answer; it needs to show *why* it thinks that answer is correct, pointing back to specific anatomical features and contextual rules.

Meng: Interpretability isn't just a nice feature; it’s a mandatory requirement for clinical adoption. If I can’t trace the AI’s reasoning back to the anatomy, I can’t trust it in an emergency setting where every second counts. The suggested improvements directly address this lack of transparency.

Lalam: From a broader cultural perspective, enhancing interpretability means that AI doesn't just dictate; it educates the human doctor by solidifying best practices and knowledge transfer across different medical institutions worldwide.

Jane: So, these suggested improvements are essentially about moving from a powerful but somewhat opaque tool to a transparent, highly collaborative diagnostic assistant that can handle diverse types of scans and information.

Tom: It sounds like they're laying out a roadmap for the next generation of medical AI tools, which is really valuable information for everyone listening who works in this field.

Lu: And I think connecting it to multi-modality data is where the true wild potential lies—combining all that anatomical knowledge from different sources opens up entirely new diagnostic possibilities we haven't even seen yet.

Meng: Practically speaking, tackling that multi-modality integration means building standardized APIs and pipelines, which is a massive undertaking, but one they are clearly mapping out for us in the future.

Lalam: If we can build these integrated systems based on the principles of Anatomy Contextualized Adaptation, the ultimate implication is that diagnosis becomes holistic—seeing the patient not as a collection of separate scans, but as one interconnected biological system.

Conclusion: Tom: Wow, we’ve covered so much ground today in Anatomy Contextualized Adaptation of CT Foundation Models. We’ve discussed the theoretical need for context, the technical architecture they are using, and the exciting future improvements. Jane, how do you feel about the overall impact this has on our field?

Jane: I feel incredibly optimistic about this work; it represents a thoughtful maturation of AI in medicine—moving from brute-force pattern matching to achieving true anatomical comprehension.

Lu: The fact that they are systematically addressing the context and adaptability issues while using frozen base models is impressive. It shows a deep understanding of how to make large, complex systems work together.

Meng: I’m glad they focused on efficiency, too; it seems like their approach is practical for deployment without needing massive computational resources for full retraining.

Lalam: Ultimately, the goal of Anatomy Contextualized Adaptation of CT Foundation Models is to empower clinicians with a level of AI that respects human anatomy, ensuring the technology serves our goal of delivering better care.

Tom: It really does feel like they're providing a robust solution to a fundamental challenge in medical AI. We’ve seen how this adaptation works across different datasets and how it promises to improve diagnostic accuracy.

Jane: It’s clear that by combining fine-grained alignment with global contextualization, this paper is setting a new standard for the way we approach CT vision-language modeling.

Lu: I hope the researchers continue exploring these avenues of inter-anatomy context, as that is where the most complex relationships in medicine truly live.

Meng: I think it's vital that ensuring its practical applicability and adapting this framework to make it run efficiently is the key to moving forward.

Lalam: We are excited to see how much more this technology can improve our global culture of health outcomes when using Anatomy Contextualized Adaptation of CT Foundation Models.

Tom: Well, we've covered a lot of ground, and I think that brings us to the end of our discussion for today. Thank you all for sharing your insights on this groundbreaking work!

Conclusion: Tom: We’ve spent a lot of time digging into this research today and it really seems like we are seeing a major step forward in how medical AI can function. The researchers are showing us that by using this concept of adaptation instead of retraining, we can achieve much higher accuracy in finding classification.

Jane: I think the biggest win here is that the model isn't just looking at random patterns; it’s actually gaining a deep understanding of human anatomy, which makes its decision-making process feel much more reliable to clinicians.

Lu: It’s truly fascinating how they are integrating this spatial reasoning into the AI framework; it’s not just seeing pixels but modeling the actual physical relationships between organs. That complexity is where the real potential lies.

Meng: I was looking at the practical side of things, and it seems like their approach is incredibly efficient, meaning that we can deploy this kind of structured knowledge in hospitals without needing massive computational resources for full model retraining.

Lalam: And when you consider the global scale of healthcare needs, allowing AI to be so structurally knowledgeable means its insights will be more trustworthy and less prone to errors across different medical systems worldwide.

Tom: It really does feel like they are addressing a foundational challenge in medical AI, though. But how do they actually execute this complex adaptation at once that's the real question we left unanswered?

Jane: Exactly, but I think understanding the mechanism is key to understanding why this is such a breakthrough for better patient outcomes.

Lu: The way they are teaching the model *how* parts relate to each other is definitely a revolutionary concept that deserves more attention.

Meng: I agree, because practical adoption demands that knowing how it works is just as important as the performance metrics.

Lalam: It’s about helping us move toward a holistic view of the patient, seeing them not as separate parts but as one interconnected biological system.

Tom: All these points—the efficiency, the structural knowledge, and the practical application—are all encapsulated in this work called Anatomy Contextualized Adaption of CT Foundation Models. It’s a huge piece of progress for a field that needs it.

Jane: I hope we get to discuss how this is being implemented in real clinical settings next time around.

R. Kenia et al.

cs.CV, cs.AI

Submitted: 2026-08-19

Updated: 2026-08-21

Code: https://github.com/lotterlab/ACA

Importance score: 84/100

The gist: The following is a detailed summary of the scientific paper, quoting relevant parts of the text: * Anatomy Contextualized Adaptation (ACA) for CT Foundation Models The research addresses limitations

Key concepts

Anatomy Contextualized Adaptation
A methodology where a general AI model is refined using specific modules. This allows the system to adjust its output based on the anatomical location or type of finding, giving it structural knowledge beyond simple pattern matching.
Structural Reasoning
The process by which the AI models not just individual pixels, but the physical relationships between those pixels and the known physical structures of the body as a coherent whole. This makes the AI's analysis robust and anatomically informed.
Multi-modality Data
The integration of different types of medical scans, such as CT scans with MRIs or ultrasounds, through a single contextualized framework. This significantly increases the complexity and potential power of the diagnostic tool.

Terminology

Summary

The following is a detailed summary of the scientific paper, quoting relevant parts of the text:


Anatomy Contextualized Adaptation (ACA) for CT Foundation Models

The research addresses limitations in current CT vision-language foundation models, which are typically trained with whole-volume representations that dilute fine-grained anatomical signals. While fine-grained vision-language pre-training (FVLP) has emerged as a strategy to align anatomy-specific visual embeddings with corresponding textual descriptions, these methods suffer from discarding the global context that whole-volume models provide. Furthermore, existing FVLP approaches have relied on training the entire model from scratch, which is computationally expensive and can be difficult to scale.

The authors introduce Anatomy Contextualized Adaptation (ACA), a framework designed to adapt frozen CT foundation model representations for anatomy-level vision-language alignment while enhancing global contextualization. ACA aims to "combine the strengths of both whole-volume foundation models and FVLP: the broad, frozen representations of pretrained CT foundation models are adapted using lightweight trainable modules to produce anatomy-level embeddings, which are then contextualized across anatomies using transformer-based attention."

Methodology

ACA operates through three core components:

  1. Embedding Construction:
  • Anatomy-level visual embeddings are extracted by applying TotalSegmentator [30] to segment each CT volume into 44 anatomical structures. The scan is passed through a frozen foundation model backbone (Merlin or CT-CLIP). To generate anatomy-level features, the authors apply a non-overlapping max-pool over each binary organ mask using a kernel matched to the backbone’s effective patch size... yielding a discrete patch-presence grid at the feature map resolution.

  • The resulting embeddings are 2-normalized. Furthermore, inter-anatomy spatial position encodings are computed to capture geometric relationships between anatomical structures.

  • Text Embedding Extraction: Anatomy-specific findings are extracted from radiology reports using Qwen3-4B-Instruct [38] (LLM), and these findings are encoded using the corresponding frozen text encoder to produce anatomy-level text embeddings. A report-level embedding is also generated by passing the full report through the text encoder.

  1. Inter-Anatomy Transformer:

To enable contextual reasoning, all present anatomy embeddings are passed jointly through a transformer encoder. Each embedding is augmented with two learned tokens: a positional encoding vector (p a) and an anatomy type embedding (t a). The final transformation is defined as h a = a W vis + MLP(p a) + t a. The sequence h a a in P is then passed through an L-layer transformer encoder with prelayer normalization, producing a contextualized embedding for each anatomy.

  1. Combined Loss Function:

The model is supervised by two complementary loss terms:

  • Anatomy-Level Contrastive Loss (L anatomy): This aligns each anatomy’s contextualized embedding to its corresponding text description. The loss uses a soft target matrix T, which accounts for both matching image-text pairs and instances of the same structure from different patients.

  • Scan-Level Report Loss (L scan): This provides a global supervision signal by aligning the mean-pooled scan-level image embedding (f scan img) to the text embedding of the full radiology report (f scan txt).

The combined objective is L total = L anatomy + lambda L scan.

Results and Discussion

The evaluation on the Merlin and CT-RATE datasets shows that ACA consistently outperforms both the frozen foundation model baselines and existing fine-grained methods in zero-shot finding classification.

Key findings include:

  • Performance: ACA achieves significant gains, with its in-distribution average AUROC on Merlin rising from 0.7729 to 0.8213, and on CT-RATE rising from 0.7082 to 0.7311 (Table 1).

  • Efficiency: ACA is a lightweight framework, requiring less than one hour of training once embeddings are cached.

  • Contextual Reasoning: The model excels at findings that require cross-anatomy context, such as organomegaly. The authors note that ACA’s inter-anatomy transformer, by attending jointly over all organ tokens in the same forward pass, can represent the relative size of each structure with respect to the others, which is crucial for judging relative pathology.

  • Learned Associations: Visualization shows that both models learn anatomically plausible cross-anatomy associations that emerge purely from contrastive training.

The authors conclude that ACA provides a practical path toward improving representation learning in a compute-efficient manner by combining fine-grained alignment with global contextualization.

Improvements for AI systems

(Warning: The following recommendations are based on a rigorous analysis of performance degradation across domain shifts and the need for clinical reliability in life-critical diagnostic systems.)

  • Improvement: Integrate a dedicated Domain Adversarial Module that explicitly trains the core feature extractors (e.g., CT-CLIP and fVLM) to be invariant to dataset shift, rather than merely optimizing for high AUROC on a single test set. This requires augmenting the training objective with an adversarial loss function that penalizes feature representations highly correlated with the source domain label, forcing generalization across diverse institutional data distributions (e.g., simultaneously training on CT-RATE and Merlin data streams).

  • Improved Capability: The system will significantly mitigate the Domain Shift Gap. Instead of performance metrics degrading when moving from an in-distribution set to an out-of-distribution set (as observed between Table 12 and Table 14), the model will maintain a robust, high AUROC across multiple, distinct hospital or scanner datasets. This dramatically increases deployment safety and reliability.

  • Improvement: Replace simple parallel concatenation of modality outputs (e.g., CLIP + MLP + fVLM) with a structured, weighted attention mechanism that operates hierarchically. The HMAFN will first assess the diagnostic certainty of each input modality for a given finding (e.g., Is the visible structure clearer in the MLP's spatial context or the fVLM's semantic context?). It will then dynamically assign weights to fuse only the most reliable feature vectors, allowing one modality to compensate for the failure of another (e.g., if image quality is low, it prioritizes semantic relationships).

  • Improved Capability: The system achieves Adaptive Diagnostic Reliability. It moves beyond merely averaging scores and instead provides a weighted synthesis of evidence. This ensures that the final diagnosis is based on the strongest available evidence, even if one input stream (e.g., ViSD-Boost) performs poorly due to local noise or artifacting.

  • Improvement: Incorporate a Bayesian Deep Learning layer at the final classification stage of all models. Instead of outputting a single point estimate (a probability, P), the system must output a full predictive distribution, characterized by both the mean prediction (mu) and an associated variance (sigma squared). This requires training the network to model parameter uncertainty alongside prediction uncertainty.

  • Improved Capability: The system provides Clinician-Grade Risk Assessment. When a diagnosis is made (e.g., Cardiomegaly), it simultaneously outputs a confidence interval (e.g., We predict Cardiomegaly with mu = 0.92 and sigma = 0.03 ). This allows the treating physician to immediately triage cases: if the variance (sigma) is high, it signals that the model itself is unsure, prompting mandatory human review regardless of a high mean prediction (mu).

  • Improvement: Augment standard heatmaps (like those used in CLIP) with a causal inference layer. This layer must not just identify where the model is looking (spatial attention), but why it is focusing there by linking local image features back to known pathological causality. For instance, instead of just highlighting a Calcification, the system must highlight and report the specific anatomical relationship that confirms its likelihood (e.g., Calcification located adjacent to the aortic root, consistent with Atherosclerosis).

  • Improved Capability: The system provides Actionable Interpretability. It transforms a black-box prediction into a fully explainable, evidence-based diagnostic report. This is critical for medical adoption, as it allows the clinician to validate the AI's reasoning pathway against their own expert knowledge, fundamentally building trust and reducing diagnostic error due to overlooked contextual details.

Sources

Related papers