Automated Lesion Segmentation of Stroke MRI Using nnU-Net: A Comprehensive External Validation Across Acute and Chronic Lesions

arXiv:2601.08701 · q-bio.QM, cs.CV · Submitted 2026-01-13 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Genomics Radio. Generated commentary on the latest computational biology and genomics papers.

Ines: Today's paper: "Automated Lesion Segmentation of Stroke MRI Using nnU-Net".

Marcus: Accurate and generalisable segmentation of stroke lesions from magnetic resonance imaging (MRI) is essential for advancing clinical research, prognostic modelling, and personalised interventions.

Ines: First, who's behind it and why it matters.

Paper summary: Ines: So, looking at the conclusion of "Automated Lesion Segmentation of Stroke MRI Using nnU-Net: A Comprehensive External Validation Across Acute and Chronic Lesions," it seems they've established that automated lesion segmentation can indeed match the performance levels seen in human experts when evaluated on independent datasets.

Marcus: And I think the authors are really emphasizing that the practical guidance they offer is focused on dataset curation and model development, specifically highlighting lesion volume, training data quality, and dataset diversity as critical factors for accuracy.

Yuki: From a broader perspective, this suggests that if we want to use these tools to study stroke effects across populations, we really need to be careful about the characteristics of the training data we choose.

Ines: I feel that this paper gives us a solid foundation for moving automated segmentation into actual clinical research and prognostic modeling, which is where the abstract said it was essential.

Marcus: And I think the implication is that we can start building more robust, transparent tools by focusing on those specific factors they highlighted, rather than just training models on whatever data is easiest to get.

Yuki: It really shows that the history of stroke research has led us to a point where we can start leveraging these computational methods more seriously, provided we respect the nuances they described regarding volume and quality.

Conclusion: Ines: So, we're wrapping up our look at this paper that tackles automated lesion segmentation using nnU-Net across both acute and chronic stroke data sets. Marcus, what are your initial thoughts on the title and who did the work?

Marcus: I think it’s a very clear title because it immediately tells us exactly what they did: comprehensive external validation across two major stroke phases using an automated framework. The authors used nnU-Net, which is impressive for its ability to adapt to different datasets, but we need to keep an eye on the specific cohorts they used for that validation.

Yuki: From a population genetics standpoint, it’s interesting because it shows a method that works across different anatomical locations and disease chronicity without needing massive pre-trained models specific to one very narrow group. That level of generalizability is what matters when we think about broader human health trends.

Ines: I agree with Yuki on the generalizability aspect; I'm curious what the authors actually recover about stroke biology from these segmentation results. Are they just drawing boxes around tissue, or are there subtle differences in lesion shape or intensity that matter biologically?

Marcus: The paper highlights that lesion volume is a key determinant of accuracy across chronic studies, which suggests that the physical size of the infarct is a powerful statistical feature for predicting outcome. This aligns with what we see in genomics data where trait size often dictates phenotypic expression.

Yuki: Exactly, and it’s not just about size; it’s how those lesions are distributed spatially within the brain structure, which speaks to underlying vascular or genetic vulnerabilities that might be present across different patient populations.

Ines: That connects to my point about temporal dynamics; since they tested both acute DWI and chronic T1w, we can start thinking about how these models might help us track the progression of injury over time in real-world clinical settings.

Marcus: And that’s where the impact could be huge; if we can reliably segment these lesions across diverse patient groups using this framework, it opens up possibilities for large-scale epidemiological studies that currently struggle with manual annotation costs and inter-rater variability.

Yuki: It moves us closer to having standardized quantitative measures of brain injury that can be used in population studies, which is a big step forward in understanding the genetic risk factors associated with stroke susceptibility.

Ines: So, the big picture here is using this technology to translate raw imaging data into quantifiable biological metrics that are robust enough for serious research and clinical modeling. Where should we go next?

Tammar. Truzman, Matthew A. Lambon Ralph, Ajay D. Halai

MRC Cognition and Brain Sciences Unit, University of Cambridge

q-bio.QM, cs.CV

Submitted: 2026-01-13

Updated: 2026-09-28

Comments: 32 pages, 7 figures. Submitted to Brain. Code and trained models available

Code: https://github.com/AjayHalai/Lesion-Segmentation-nnUNet

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 90/100

The gist: Accurate and generalisable segmentation of stroke lesions from magnetic resonance imaging (MRI) is essential for advancing clinical research, prognostic modelling, and personalised interventions.

Key concepts

nnU-Net
A fully automated framework that automatically configures the neural network architecture and training settings based on the specific MRI dataset it is analyzing. It handles preprocessing, training, and inference with minimal manual intervention, making it flexible for different stroke studies.
Dice Coefficient
A standard medical image segmentation metric used to measure how well a model's predicted lesion area overlaps with the actual human-annotated lesion area. A higher Dice score indicates better spatial accuracy in defining the boundaries of the stroke damage.
Lesion Volume
The physical size or total amount of tissue affected by a stroke, measured in cubic units. The study found that larger lesions were segmented more accurately, suggesting that volume is a key feature for models to learn and generalize across different patients.

Terminology

Summary

Accurate and generalisable segmentation of stroke lesions from magnetic resonance imaging (MRI) is essential for advancing clinical research, prognostic modelling, and personalised interventions. The gist: Models trained on DWI consistently outperformed those trained on FLAIR in acute stroke lesion segmentation, while lesion volume emerged as a key determinant of accuracy across chronic stroke studies.

Methodology and Framework

The study systematically evaluated the nnU-Net framework across five independent stroke MRI datasets spanning acute and chronic stages. The authors utilized nnU-Net, a fully automated, self-configuring framework that adapts network architecture and training parameters to the dataset at hand, employing its standard workflow from automated preprocessing (nnUNetv2 plan and preprocess) through training and inference. For acute stroke datasets (SOOP and ISLES), models were trained using paired DWI and ADC maps as a two-channel input, while for chronic stroke datasets (ATLAS v2.0, ARC, CCNRP), models were trained using single-channel T1w images. Initial model training involved five-fold cross-validation with a loss function combining the Dice coefficient and cross-entropy loss, followed by identifying the optimal model using nnUNetv2 find best configuration based on mean cross-validation Dice scores.

Evaluation Metrics and Performance

Performance was rigorously evaluated using standard medical image segmentation metrics, specifically the Dice coefficient to quantify spatial overlap, and the 95th percentile Hausdorff Distance (HD95) to measure boundary distance. The results demonstrated that models achieved robust generalisation, with segmentation accuracy approaching reported inter-rater reliability for manual annotations, ranging from Dice ≈.72–.84. For acute stroke, the SOOP-trained DWI+FLAIR model achieved the best segmentation accuracy (median Dice =.777), showing a statistically significant improvement over the SOOP-trained DWI-only model. For chronic stroke, the ATLASv2 model achieved a median Dice of.81 on external test sets.

Key Determinants of Generalisability

The research identified several critical factors governing model generalisability across different conditions. The performance varied systematically based on:

  1. Training Data Composition: Increasing training set size led to systematic performance gains for chronic stroke, with the largest improvements observed when moving from small to medium datasets, indicating diminishing returns beyond several hundred training cases.

  2. Lesion Characteristics: Lesion volume emerged as a key determinant of accuracy, with larger lesions were segmented more accurately, and models trained on broad volume distributions showed more symmetric cross-dataset performance. Conversely, very small lesions remained challenging to segment, and models trained on restricted volume ranges generalized poorly.

  3. Image Quality: Differences in MRI quality contributed to variability; ARC exhibited systematically lower image quality across multiple dimensions (SNR, tissue boundary definition), yet models trained on higher-quality datasets showed robustness at inference.

Modality and Laterality Effects

The study provided specific insights into modality and anatomical factors. For acute stroke, models trained on DWI consistently outperformed those trained on FLAIR, suggesting that DWI provides more reliable features for the acute phase due to the temporal evolution of lesions. For chronic stroke, lesion laterality did not impact segmentation accuracy or generalisability; models trained on left, right, or bilateral lesions achieved comparable performance. Furthermore, volume-matched analysis showed that apparent generalisation asymmetries largely disappeared once lesion distributions were aligned by matching training and test set volumes.

Conclusion and Implications

The findings demonstrate that automated lesion segmentation can achieve performance comparable to human experts when evaluated on independent datasets. The study provides practical guidance for dataset curation and model development, highlighting key factors such as lesion volume, training data quality and dataset diversity as key determinants of accuracy. Future work should focus on capturing temporal dynamics and extending automated segmentation to modalities like CT. All trained models and code are openly shared to promote transparency and community benchmarking. The conclusion is that these findings offer practical guidance for dataset curation and model development toward more robust, generalisable, and clinically relevant lesion segmentation tools.


The gist

Models trained on DWI consistently outperformed those trained on FLAIR in acute stroke lesion segmentation, while lesion volume emerged as a key determinant of accuracy across chronic stroke studies.

How it works

  1. The study utilized the nnU-Net framework, a fully automated, self-configuring framework that adapts network architecture and training parameters to the dataset at hand, following its standard workflow from automated preprocessing (nnUNetv2 plan and preprocess) through training and inference.

  2. For acute stroke datasets (SOOP and ISLES), models were trained using paired DWI and ADC maps as a two-channel input, while for chronic stroke datasets (ATLAS v2.0, ARC, CCNRP), models were trained using single-channel T1w images.

Improvements for AI systems

Here are the specific improvements that can be made to AI systems, derived from the findings of this study, and what those improved systems could achieve:


  1. Improve segmentation performance across stroke stages by prioritizing Diffusion Weighted Imaging (DWI) inputs for acute lesions.

  2. Develop modality-invariant or multimodal models that effectively integrate DWI and FLAIR information to overcome the limitations of using only one sequence, especially in the subacute phase where lesion appearance evolves differently between modalities.

  3. Implement dynamic training strategies for chronic stroke segmentation by incorporating lesion volume distribution as a key feature, specifically training models on broad volume distributions to improve cross-dataset generalizability.

  4. Design adaptive training pipelines where model configuration and parameters are automatically optimized based on the specific input dataset's characteristics (e.g., using the nnU-Net framework effectively) rather than relying on fixed settings.

  5. Improve robustness to image quality variations by incorporating image quality metrics (like those derived from MRIQC, particularly SNR, CNR, and residual partial volume effect) directly into the model training or inference process to mitigate performance degradation on lower-quality scans.

  6. Create automated pre-processing steps that normalize intensity based on dataset-specific characteristics (Zscore normalization) to ensure better feature learning across heterogeneous MRI datasets.

  7. Enhance the reliability of models by incorporating uncertainty estimation or leveraging consensus mechanisms in the ground truth annotation process to account for inherent variability in human labeling, particularly for challenging lesion boundaries and small lesions (< 10 cm3).

These improved AI systems can achieve the following:

  1. Accurately segment acute stroke lesions with performance approaching human inter-rater reliability (Dice ≈.78), specifically by leveraging DWI data, leading to more reliable prognostic biomarkers for acute stroke severity.

  2. Provide a single, robust segmentation tool capable of performing reliably across both acute and chronic stroke stages using the nnU-Net framework, offering consistent performance on independent datasets from diverse clinical sites.

  3. Enable personalized stroke intervention by accurately quantifying lesion volume and shape in chronic lesions, allowing clinicians to stratify patients based on tissue damage severity without needing manual tracing expertise.

  4. Develop smart models that can perform well even when deployed in real-world clinical settings where image quality is variable (e.g., lower SNR scans), ensuring a consistent level of diagnostic accuracy regardless of scanner quality.

  5. Improve the efficiency of research by enabling researchers to quickly train and test models on new stroke datasets, accelerating the development cycle for novel segmentation algorithms.

  6. Create more transparent and reproducible AI tools by providing open-source models and detailed documentation that allow other researchers to understand exactly how performance was achieved across different training set sizes, volumes, and modalities.

Related papers