SAMRI-2: A Memory-based Model for Cartilage and Meniscus Segmentation in 3D MRIs of the Knee Joint

arXiv:2502.10559 · eess.IV, cs.AI, cs.CV · Submitted 2025-02-14 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "SAMRI-2: A Memory-based Model for Cartilage and Meniscus Segmentation in 3D MRIs of the Knee Joint".

Jane: The paper was written by Danielle L. Ferreira, Bruno A. A. Nunes, Xuzhe Zhang, Laura Carretero Gomez, Maggie Fung et al. from GE HealthCare and Columbia University and LAIMBIO, Rey Juan Carlos University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Summary: Tom: So, building on that concept, the paper gives us a detailed summary of their methodology in "SAMRI-two: A Memory-based Model for Cartilage and Meniscus Segmentation in three dee MRIs of the Knee Joint."

Jane: They tested four different AI models—a standard CNN called three dee-VNet, two transformer-based models, and that interactive one we're calling SAMRI-two.

Lu: It's fascinating to see them comparing traditional convolutional networks directly against these more advanced transformer architectures.

Meng: The key takeaway from the summary is that they ran these tests on two major datasets, D7 and D50, which are external holdout sets, meaning they haven' a never seen the data before.

Lalam: I want to emphasize that the results showed SAMRI-two outperformed all other models across both datasets.

Tom: That performance is measured using metrics like Dice Score (DSC) and Intersection over Union (IoU), which are basically standard ways to see how much overlap there is between predicted segmentation and the actual expert reference standard.

Jane: And the summary makes it clear that this model didn't just achieve high accuracy; it also managed to reduce the amount of effort required from human annotators.

Lu: That’s a huge step toward automation, making complex tasks much more accessible for any specialist who might be looking at these images.

Meng: The fact that the summary is focused on performance across different scanners and populations shows they are building a model that is robust, not just one specific to one lab.

Lalam: By seeing how it performs in this comprehensive manner, the future of AI in musculoskeletal imaging looks incredibly bright indeed.

Improvements: Tom: We’ve seen the summary, but now let's talk about what makes this model so effective—the actual improvements mentioned in "SAMRI-two: A Memory-based Model for Cartilage and Meniscus Segmentation in three dee MRIs of the Knee Joint."

Jane: The most significant contribution seems to be this thing they call the Hybrid Shuffling Strategy, or HSS.

Lu: It sounds like a clever way to enforce spatial consistency that goes beyond just looking at individual slices.

Meng: I think HSS is what allows the AI to understand that if you see a certain shape in slice five, it' should probably flow into the slice six and seven.

Lalam: That’s exactly how we move from seeing isolated pixels to seeing a coherent, continuous anatomical structure.

Tom: The paper explains that HSS shuffles data at the level of "chunks" or sub-volumes instead of just individual slices.

Jane: That chunk-based processing helps the AI capture those spatial dependencies across the slices far more effectively than standard methods allow it to.

Lu: It’s like teaching a student to think about a whole paragraph, not just one word at a time, when they are reading text.

Meng: And because HSS improves how the memory encoder processes three dee data, it significantly boosts the model's ability to maintain spatial awareness during training.

Lalam: This structural improvement is what allows the AI to finally match human consistency across a whole volume, which is a massive leap for reliability.

Tom: It’s definitely a huge factor in ensuring this is not just another incremental improvement but a fundamental change in how the segmentation process works.

Results: Tom: We've talked about the theory and now we need to look at the actual results presented in "SAMRI-two: A Memory-based Model for Cartilage and Meniscus Segmentation in three dee MRIs of the Knee Joint."

Jane: The data shows that, on average, SAMRI-two achieved a significant improvement of five points in Dice Score compared to all other models.

Lu: And the peak improvements for specific areas, like the tibial cartilage, were even higher than that.

Meng: That’s impressive because it means the model is not just generally better; it’s excelling at specific, difficult anatomical regions.

Lalam: I think we should mention how low its average error was in measuring cartilage thickness—it reduced discrepancies by up to three times compared to other models.

Tom: That speaks directly to the accuracy of its morphometric assessment, which is so crucial for clinical tracking of osteoarthritis progression.

Jane: The paper shows that the model maintains high performance even with minimal human input, requiring as few as three user clicks per volume.

Lu: That’s a huge win for efficiency and consistency in a practical clinical setting.

Meng: It provides a level of precision that is hard to achieve manually, especially when you factor in inter-observer variability among radiologists.

Lalam: This means the technology is ready to assist clinicians in a way that significantly improves the quality and speed of their diagnostic process.

Conclusion: Tom: We are wrapping up our discussion on "SAMRI-two: A Memory-based Model for Cartilage and Meniscus Segmentation in three dee MRIs of the Knee Joint," and I think we've covered a lot of ground.

Jane: To summarize, we' have seen how this interactive, memory-based model combines advanced training strategies like HSS with the power of VFMs to solve some very difficult segmentation problems.

Lu: It’s clear that by combining these elements, the researchers have created something that generalizes exceptionally well across different types of MRI scans.

Meng: My main takeaway is that this is a practical, robust tool that can significantly reduce the workload on specialists while providing highly accurate results.

Lalam: I believe the future where AI assists in every aspect of diagnostic imaging has arrived with this model, improving patient care globally.

Tom: Before we go, I want to hear one final thought from each of you.

Lu: I’m just excited about the possibilities; how many other complex structures can we apply this memory-based approach to next time in the research?

Meng: We need more clinical validation studies, but based on these results, it looks like a viable product is ready to be tested in real-world settings.

Lalam: I see the model as a bridge; it connects complex data with human reliability, making sure that's where we always aim to improve our culture of care.

Tom: Thank you all for this great discussion on "SAMRI-two: A Memory-based Model for Cartilage and Meniscus Segmentation in three dee MRIs of the Knee Joint."

Jane: We hope you enjoyed hearing about this groundbreaking work, and we're excited to share the next paper with you soon.

Danielle L. Ferreira, Bruno A. A. Nunes, Xuzhe Zhang, Laura Carretero Gomez, Maggie Fung, Ravi Soni

GE HealthCare · Columbia University · LAIMBIO, Rey Juan Carlos University

eess.IV, cs.AI, cs.CV

Submitted: 2025-02-14

Updated: 2026-08-25

Journal ref: Scientific Reports 16, 1825 (2026)

DOI: 10.1038/s41598-025-31503-2

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

Importance score: 85/100

The gist: The paper details the development and evaluation of a memory-based model, SAMRI-2, for segmenting cartilage and meniscus in 3D MRIs of the knee joint.

Key concepts

SAMRI-2
A memory-based AI model designed for segmenting cartilage and meniscus in 3D MRIs of the knee joint. It was tested against other AI models and showed superior performance across multiple datasets.
Hybrid Shuffling Strategy (HSS)
A significant improvement in the model's architecture that shuffles data at the level of 'chunks' or sub-volumes, rather than individual slices. This helps the AI maintain spatial consistency and understand anatomical structures across a whole volume.
Dice Score (DSC) and IoU
Standard metrics used to measure how much overlap there is between a predicted segmentation (by the AI) and the actual expert reference standard. Higher scores indicate better accuracy of the model's predictions.

Terminology

Summary

The paper details the development and evaluation of a memory-based model, SAMRI-2, for segmenting cartilage and meniscus in 3D MRIs of the knee joint. The methodology incorporates advanced training strategies and evaluates various mask propagation techniques across different datasets.

Training Strategy:

During training, a Hybrid Shuffling Strategy (HSS) is employed. This strategy dictates that at the start of each epoch, the Baseline model is fed RadA batches... [and] at the end of each epoch, the RadA chunks are shuffled, and the process repeats in the next epoch.

Evaluation of Mask Propagation:

The study evaluates mask propagation for SAMRI-2 and SAM2 (without fine-tuning) using different slice selection strategies. These strategies include:

  1. Placing click prompts on every sagittal slice of the volume (SAMRI-2 ALL).

  2. Placing a click prompt on the first slice where the structure appears, then propagating it across subsequent slices at intervals of 10, 20, or 50 slices (SAMRI-2 every 10, SAMRI-2 every 20, or SAMRI-2 every 50).

Segmentation Performance and Results:

The performance is measured using Dice Scores for four segmentation classes: femoral cartilage, tibial cartilage, patellar cartilage, and the meniscus. The results demonstrate the quantitative comparison of these methods across multiple datasets (D7 Rad A/B, D50) and slice selection strategies. For instance, when evaluating the D7 Rad A dataset:

  • The SAMRI-2 ALL strategy achieved a Dice Score of 0.722 for femoral cartilage.

  • The SAMRI-2 every 10 strategy achieved a score of 0.826 for femoral cartilage.

  • The SAMRI-2 every 50 strategy achieved a score of 0.7 for femoral cartilage.

Similar comparative results are presented across the tibial, patellar, and meniscus segments, showing varying scores depending on the dataset and slice selection method (e.g., for the meniscus in D7 Rad A, SAMRI-2 ALL scored 0.852).

Visualization of Agreement:

Visual analysis confirms model performance through 3D renderings. Figure 2 shows that strong agreement is observed on sagittal slices through the center of the structures. However, discrepancies are noted in peripheral areas: in peripheral slices... models begin to disagree, particularly for tibial cartilage and the meniscus, though these discrepancies are limited only to a few slices.

Furthermore, Figure 4 provides a visual comparison of segmentation outputs by showing two cases: the best (Case 6) and the worst agreement (Case 2) between the two radiologists Rad A and Rad B, in the D7 holdout testset, illustrating real-world variability.

Improvements for AI systems

As a researcher focused on high-stakes AI implementation, I have thoroughly reviewed the architecture and findings of the SAMRI-2 paper. The system is highly effective, but its reliance on specific training methodologies and limitations in real-world interaction present critical avenues for advancement.

The improvements detailed below focus not only on optimizing the existing successful components (HSS and Memory) but also on addressing the noted limitations to create a robust, clinically viable, next-generation system.


1. Generalization of the Hybrid Shuffling Strategy (HSS)

  • The Improvement: Implement HSS not just for contiguous slices along the Z-axis, but also adapt it to handle non-uniform or anisotropic data structures where spatial correlation is critical (e.g., segmenting complex vessels in CT angiography, or spinal cord segmentation). This involves defining chunks based on structural coherence rather than purely geometric continuity.

  • What the Improved System Can Do: It will significantly enhance the model's ability to maintain internal spatial consistency during training, allowing it to learn complex anatomical dependencies even when data is shuffled, ensuring that the model achieves high accuracy in scenarios where slice-by-slice correlation is paramount.

2. Dynamic Adaptive Prompt Generation (Addressing the Predefined Limitation)

  • The Improvement: Replace the current practice of deriving prompt placement from predefined reference standard clicks with a Predictive Uncertainty Mapping (PUM) mechanism. This module will analyze the model’s internal confidence maps and identify regions where its segmentation probability drops sharply or where inter-model disagreement is highest (i.e., areas of high uncertainty). It then automatically generates optimal click prompts within those high-uncertainty zones.

  • What the Improved System Can Do: It transforms the system from a semi-automatic tool requiring human input to a highly efficient, autonomous workflow. The AI will proactively guide the radiologist to areas needing attention, drastically reducing annotation time while ensuring that even subtle or ambiguous anatomical features are correctly identified without human bias.

3. Meta-Learning for Cross-Scanner Generalization

  • The Improvement: Implement meta-learning techniques during fine-tuning (using the CUBE dataset). Instead of simply training on the data, the model will be trained to learn how to adapt its weights quickly given a new input distribution. This involves training on a suite of diverse meta-scenarios (simulated variations in FOV, contrast, and noise).

  • What the Improved System Can Do: The model will achieve true generalization across different MRI vendors, acquisition protocols, and patient populations. It will perform with consistent high performance regardless of whether the input image came from a Siemens scanner or a GE scanner, making it reliable for multi-center clinical deployment.

4. Integrated Morphometric Risk Assessment (From Segmentation to Prediction)

  • The Improvement: Integrate the precise volume and thickness measurements derived from the SAMRI-2 segmentation directly into a predictive modeling pipeline (e.g, a regression network). This pipeline will correlate changes in cartilage volume/thickness over time with clinical outcomes.

  • What the Improved System Can Do: It elevates the system from a pure segmentation tool to a Clinical Decision Support System. It can automatically quantify disease progression (e.g, Cartilage loss is 15% compared to baseline) and provide automated risk stratification for patients, enabling earlier intervention in osteoarthritis management.

5. Multi-Modal Consensus Aggregation

  • The Improvement: Develop a framework that allows the SAMRI-2 output to be combined with other data sources (e.g., patient demographics, lab results). The model will learn how to adjust its segmentation confidence based on these external clinical variables.

  • What the Improved System Can Do: It increases the robustness of the segmentation in complex cases where anatomical features are ambiguous. For instance, if a patient's age or specific clinical markers suggest early OA, the system can be more conservative and refine its prediction to match known pathological patterns, ensuring higher accuracy than a purely image-driven model.


Summary of Core Capability:

The improved system will operate as an Autonomous Clinical Segmentation Engine. It will not only achieve superior segmentation accuracy (exceeding 0.85 DSC) while minimizing human effort through predictive prompting but will also provide actionable, quantitative morphometric data that allows clinicians to track disease progression and make evidence-based treatment decisions in a highly automated, generalizable manner.

Abstract

Accurate morphometric assessment of cartilage-such as thickness/volume-via MRI is essential for monitoring knee osteoarthritis. Segmenting cartilage remains challenging and dependent on extensive expert-annotated datasets, which are heavily subjected to inter-reader variability. Recent advancements in Visual Foundational Models (VFM), especially memory-based approaches, offer opportunities for improving generalizability and robustness. This study introduces a deep learning (DL) method for cartilage and meniscus segmentation from 3D MRIs using interactive, memory-based VFMs. To improve spatial awareness and convergence, we incorporated a Hybrid Shuffling Strategy (HSS) during training and applied a segmentation mask propagation technique to enhance annotation efficiency. We trained four AI models-a CNN-based 3D-VNet, two automatic transformer-based models (SaMRI2D and SaMRI3D), and a transformer-based promptable memory-based VFM (SAMRI-2)-on 3D knee MRIs from 270 patients using public and internal datasets and evaluated on 57 external cases, including multi-radiologist annotations and different data acquisitions. Model performance was assessed against reference standards using Dice Score (DSC) and Intersection over Union (IoU), with additional morphometric evaluations to further quantify segmentation accuracy. SAMRI-2 model, trained with HSS, outperformed all other models, achieving an average DSC improvement of 5 points, with a peak improvement of 12 points for tibial cartilage. It also demonstrated the lowest cartilage thickness errors, reducing discrepancies by up to threefold. Notably, SAMRI-2 maintained high performance with as few as three user clicks per volume, reducing annotation effort while ensuring anatomical precision. This memory-based VFM with spatial awareness offers a novel approach for reliable AI-assisted knee MRI segmentation, advancing DL in musculoskeletal imaging.

Sources

Related papers