Unified-protocol voxel-level pulmonary embolism annotations for three public CT angiography datasets
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Unified-protocol voxel-level pulmonary embolism annotations for three public CT angiography datasets".
Jane: This study quantifies how evaluation annotations influence measured pulmonary embolism (PE) segmentation performance relative to model training changes, establishing a human-referenced framework for future studies.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So, looking at the title again, "Unified-protocol voxel-level pulmonary embolism annotations for three public CT angiography datasets," it really hammers home that they are focusing on standardizing the labeling process itself. It’s not just about training a better model; it's about making sure everyone is talking about the same thing when they talk about PE segmentation.
Jane: Exactly, Tom; and the authors listed show a good mix of expertise, with people from places like TUM and various radiology departments, which gives them a solid foundation for this kind of rigorous validation work. It tells us this isn't just theoretical work; it’s grounded in clinical application.
Lu: The implication here is huge because when you have multiple datasets that all use different annotation conventions, comparing model performance becomes almost meaningless unless you normalize the ground truth first. This paper provides that normalization tool using the FairPE method, which is really clever.
Meng: From an engineering standpoint, if we adopt this unified protocol mentioned in the title, it means our data ingestion pipeline could be much more robust against variations in how different teams label things manually or semi-automatically. That consistency is key for deploying reliable tools.
Lalam: I see a potential cultural impact here; by establishing this common reference standard, it helps move the conversation away from just chasing the latest model architecture and toward ensuring that the data quality itself is consistent across different research groups.
The paper's summary: Tom: Now, let's talk about what they actually found in their summary of this paper; they quantified exactly how much the way annotations are done influences segmentation performance compared to just changing the model’s training setup. It shows that the difference between good and bad labels is a real, measurable factor in how well an AI segments pulmonary emboli.
Jane: That's a crucial distinction, Tom; it confirms that we can’t just blame the model architecture when results vary; we have to account for the label effect. They measured both the "label effect" and the "model effect" on one hundred forty-nine cases from CADPE, FUMPE, and READ.
Lu: The summary highlights a key finding: that annotation quality actually has a significant impact on performance, which they measure by how much segmentation metrics change when only the annotations are refined. They found that for models like nnU-Net-A and nnU-Net-B, changing just the annotations increased mean DSC by about zero point one four three or zero point one eight eight, depending on the model.
Meng: That quantitative measurement is what really matters for us in practice; seeing that a small change in annotation quality translates to a measurable improvement in segmentation metrics tells us exactly where our efforts should be focused if we are working with labeled data.
Lalam: It's exciting because it gives researchers a clear roadmap: prioritize improving the annotation protocol before tweaking every single model configuration for every new dataset. This shifts the focus of development quite a bit.
The paper's improvements: Tom: The paper suggests some really important improvements to this whole process, especially around how we assess those annotations. They look at three criteria: physical consistency like within-mask CT attenuation, human agreement on small subsets, and even testing models that haven't seen the public annotations at all.
Jane: That assessment of quality is a major step up because it moves beyond just comparing final segmentation masks; they are checking the underlying physical reality and inter-rater reliability. They even found that within-mask attenuation standard deviation fell in all three datasets after re-annotation, which is pretty telling.
Lu: The paper also points out some specific systematic errors in the original annotations for each dataset; for CADPE, it was dominated by over-segmentation onto things like normal pulmonary arteries and lung parenchyma. For FUMPE, it showed significant under-segmentation with missed emboli in twenty out of thirty-three cases.
Meng: Those specific error patterns are incredibly valuable because they give us concrete targets for data curation; instead of just saying "the labels are bad," we know exactly *what* the labels are wrong about, which helps engineers fix the input pipeline directly.
Lalam: This detailed error analysis is something I think will have a big impact on how we design data validation tools in our AI development process; knowing where the known pitfalls lie allows for targeted quality control.
Conclusion: Tom: So, to wrap up, the main conclusion of this paper is that the annotation quality can influence measured pulmonary embolism segmentation performance as strongly as changes made to the model configuration itself. They established a human-referenced evaluation framework, which provides a common reference standard for everyone moving forward in this field.
Jane: That’s a big statement, Tom; it means researchers can now attribute performance variations much more accurately to the model's capability rather than just assuming everything is about the labeling convention used. They released the refined dataset FairPE and the evaluation toolkit publicly, which is really helpful for others building on this work.
Lu: The implication for future AI development is that we need to integrate this framework—the annotation protocol, uncertainty ranges, and multi-dimensional metrics—directly into our validation pipelines so we can make these kinds of assessments systematically.
Meng: From a practical standpoint, having a standard like FairPE means that when we compare two different models or datasets down the road, we won't be guessing about whether the performance gap is due to better architecture or just better labeling; we have a way to test that hypothesis rigorously.
Lalam: I think this whole study on "Unified-protocol voxel-level pulmonary embolism annotations for three public CT angiography datasets" sets a really important precedent for how we structure validation in complex medical imaging problems, and it gives us a solid foundation to build future tools upon.
Qihang Suna, Zhongxiao Liuc, Bailiang Jiana, *Shenman Qiuc, Jingyuan Wangd, Lei Zhange, Lixiang Xiec, Jiazhen Pana
Technical University of Munich (TUM) · Munich Center for Machine Learning (MCML) · Department of Radiology, The Affiliated Hospital of Xuzhou Medical University · Department of Radiology, The Third Affiliated Hospital of Soochow University · Department of Radiology, The Affiliated Taizhou People’s Hospital of Nanjing Medical University
eess.IV, cs.CV
Submitted: 2026-08-25
Updated: 2026-09-29
Comments: 18 pages, 5 figures, 1 table
Code: https://github.com/iblueer777/PEbench
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 79/100
The gist: This study quantifies how evaluation annotations influence measured pulmonary embolism (PE) segmentation performance relative to model training changes, establishing a human-referenced framework for
Key concepts
- Unified Protocol
- A standardized labeling process for pulmonary embolism annotations across different datasets. This aims to ensure everyone uses the same method when labeling medical images, making comparisons between models and datasets meaningful.
- Label Effect
- The measurable influence that the quality of annotations has on how well an AI model segments pulmonary emboli. The study found that refining labels can measurably improve segmentation metrics even when keeping the model architecture the same.
- FairPE Method
- A tool developed in this paper used to provide normalization for ground truth data across different annotation conventions. It helps standardize evaluations so that performance comparisons are not skewed by differences in how data was initially labeled.
Terminology
Summary
This study quantifies how evaluation annotations influence measured pulmonary embolism (PE) segmentation performance relative to model training changes, establishing a human-referenced framework for future studies. It addresses the critical question of whether differences in segmentation metrics stem from annotation quality or model configuration by measuring both the label effect
and the model effect
on a unified set of 149 cases derived from three public datasets (CADPE, FUMPE, and READ). This work is significant because it provides a common reference standard—the FairPE framework—allowing researchers to attribute performance variations more accurately to model capability rather than annotation fidelity.
Materials and Methods
The retrospective study screened 166 voxel-annotated CT pulmonary angiography cases from CADPE (n = 91), FUMPE (n = 35), and READ (n = 40), including a total of 149 included cases. The study employed a multi-rater annotation protocol, where a primary rater created Annotation 1, followed by revision by a senior thoracic radiologist to produce Annotation 2. Three additional raters independently annotated a subset of cases to produce Annotations 3 and 4.
Key methodological steps included:
-
The label effect was measured by evaluating two pretrained nnU-Net models (nnU-Net-A and nnU-Net-B) against original and refined annotations, taking the
per-case difference in each metric between these two evaluations
as the label effect. -
The model effect was measured by comparing the same architecture trained on different dataset combinations with annotations fixed, using a benchmark model (nnPE) trained with leave-one-dataset-out and pooled five-fold cross-validation.
-
All three public datasets were annotated under a unified protocol to construct FairPE, which serves as the refined reference standard.
Annotation Quality Assessment
The researchers systematically assessed the quality of the original annotations against three independent criteria:
-
Within-mask CT attenuation (physical). The study found that "Within-mask attenuation SD fell in all three datasets after re-annotation (all P <.001)."
-
Agreement among four annotation sets produced under the same protocol on a 15-case subset (human).
-
Performance of two externally trained models that had not used either set of public test annotations (model).
The analysis revealed systematic errors in the original annotations:
** CADPE was dominated by over-segmentation onto normal pulmonary arteries (74%), lung parenchyma (32%), bronchi (30%), and pulmonary veins (16%).**
FUMPE was dominated by under-segmentation, with missed emboli in 20/33 cases (61%).
Model Performance Comparison
The study compared the performance of two pretrained nnU-Net models against the original and refined annotations. The results showed that both models performed significantly better against the re-annotation than against the original annotations.
For instance, for nnU-Net-A on CADPE, DSC improved from 0.48 ± 0.25 to 0.63 ± 0.25 (∆% +40.9%).
The performance of the benchmark model (nnPE) was also assessed across dataset combinations:
Within each dataset, the pair of training-set combinations with the largest mean DSC difference was selected, and the per-case difference between that same pair was then computed for all metrics.
Label Effect vs. Model Effect
The principal finding indicates that annotation quality significantly impacts performance. The label effect was measured by changing only the annotations while holding model predictions fixed:
-
Changing only the annotation increased mean DSC by 0.143 (0.122–0.166) for nnU-Net-A and 0.188 (0.163–0.213) for nnU-Net-B (both P <.001).
-
The label effect exceeded the model effect on CADPE and FUMPE, but was 0.045 on READ, where the Control effect was not significant (Figure 6B).
-
This ordering held for boundary, volumetric, and lesion-level metrics across all three dimensions (Figure 6C, D).
Conclusion and Framework
The principal conclusion is that the annotation can influence measured PE segmentation performance as strongly as the model configuration.
The study established a human-referenced evaluation framework comprising an annotation protocol, multi-rater uncertainty ranges, a multi-dimensional metric set, and a baseline trained on the refined annotations (nnPE), providing a common reference standard for future studies.
This framework allows researchers to attribute performance differences to the model rather than solely to labeling convention. The refined dataset FairPE and the evaluation toolkit are publicly released.
Improvements for AI systems
Here are specific improvements for AI systems derived from this research, focusing on leveraging the findings regarding annotation quality versus model architecture:
)1. Robust Evaluation Framework for Segmentation Models:
The paper establishes a human-referenced benchmark (FairPE) and quantifies the label effect
versus the model effect.
AI development should move beyond simple comparison of architectures (e.g., nnU-Net vs. nnU-Net) or fixed training sets by integrating these metrics into the validation pipeline.
- Dynamic Annotation Quality Assessment:
Improve models by implementing a pre-processing step that assesses annotation quality using criteria derived from the study (e.g., within-mask attenuation SD, inter-rater agreement scores). If an input dataset exhibits high noise or low agreement (as seen in FUMPE's under-segmentation), the system should flag this uncertainty and potentially adjust its confidence scores or rely more heavily on model performance against a human-referenced standard rather than the raw annotation.
- Domain-Specific Performance Tuning:
Since the label effect varies by dataset (label effect > 0.045 in CADPE/FUMPE vs. < 0.045 in READ), AI training strategies should be tailored based on the expected input data characteristics (e.g., contrast opacification patterns). For datasets like READ, where label effects are minimal, model improvements should focus more intensely on architectural changes or training set composition rather than solely relying on annotation refinement.
- Uncertainty-Aware Segmentation:
Integrate the findings regarding residual gaps between the model (nnPE) and human raters. AI systems should be designed to output not just a segmentation mask, but also a quantifiable uncertainty map (similar to how DSC fell for small-volume cases). This allows clinicians to distinguish between genuine model deficit
(where performance drops) and measurement artifact/boundary disagreement
(where performance is stable but boundary accuracy suffers), improving clinical trust in the AI output.
- Multi-Scale and Threshold-Aware Detection:
The paper provides detailed metrics for lesion detection across three thresholds (1px, 10%, 20% overlap). AI models should be trained to predict segmentation metrics at multiple scales simultaneously. This allows the system to perform coarse
detection (e.g., at 20% overlap) while maintaining high precision for fine-grained
lesions (e.g., at 1px), optimizing performance across different clinical needs without requiring separate model training for each threshold.
- Automated Error Pattern Identification:
Use the annotation error patterns identified by raters (e.g., over-segmentation into pulmonary artery wall in CADPE) to develop specific regularization techniques or loss functions that penalize these known failure modes during training, leading to more anatomically plausible segmentations directly from the model.
Sources
Related papers
- Revisiting Integration of Image and Metadata for DICOM Series Classification: Cross-Attention and Dictionary Learning
- VesselSDF: Distance Field Priors for Vascular Network Reconstruction
- cSVR: Convolutional Slice-to-Volume Reconstruction
- NAIMA: Semantics Aware RGB Guided Depth Super-Resolution
- AneumoBench: A Source-Linked Benchmark for Synthetic-Geometry Transfer in Aneurysm CFD
- RETO: A Rotary-Enhanced Transformer Operator for High-Fidelity Prediction of Automotive Aerodynamics