ImageCAS-X: a dataset and benchmark for coronary artery segmentation and centerline extraction in coronary CT angiography
Kit M. Bransby, Esther Øksnebjerg, Kristoffer Kjær, Jacob Kirkeby, Yasmin El Youssef, Aïda Jiménez, Philip R. Pedersson, Martina C. de Knegt, Klaus F. Kofoed
Technical University of Denmark, DTU Compute, Kongens Lyngby, Denmark · Copenhagen University Hospital – Rigshospitalet, Copenhagen, Denmark
cs.CV, cs.AI
Submitted: 2026-08-31
Updated: 2026-08-31
Comments: Pre-print (under review)
Code: https://github.com/kitbransby/ImageCAS-X
Project page: https://kitbransby.github.io/ImageCAS-X
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 94/100
The gist: The paper details a comprehensive shared training and inference framework designed for coronary artery segmentation and centerline extraction, ensuring that "all methods are trained and evaluated
Terminology
Summary
The paper details a comprehensive shared training and inference framework designed for coronary artery segmentation and centerline extraction, ensuring that all methods are trained and evaluated through a single shared framework that fixes everything which must remain identical for a fair comparison.
Data Input and Preprocessing:
The input data involves processing the cMPR volume. Specifically, Random crops from cMPR volume of size N × H × W were used as input,
where N was set at 10 frames with 0.4mm spacing between frames, and H and W were set at 64, representing a 16 × 16mm cross section using 0.25 mm pixel spacing.
For the shared framework (Section E), every scan is first processed by resampling to isotropic 0.5 mm spacing,
clipping intensities to a [−200, 1000] HU window and linearly mapped to [0, 1].
The framework supports three input types: "whole volume where the full preprocessed volume is passed to the network; random crop where a fixed-size patch is cropped at a random location from the whole volume; and center crop where patches are extracted at precomputed coordinates. For patch-based methods,
foreground (fg) sampling is used to ensure at least 50% of every batch contains patches with lumen class."
Anatomical Segmentation Classes:
The study utilizes a defined set of coronary segments, which are mapped to specific class indices:
Branch Segment Class ID
:---:---:---
Left main LM (Left main) 1
Left anterior descending artery LAD (LAD) 2-3 (Implied range/separate classes) [Note: The table lists LAD as 2]
Left circumflex artery LCx (LCx) 3-4 [Note: The table lists LCx as 3]
First diagonal branch D1 (D1) 4-5 [Note: The table lists D1 as 4]
Second diagonal branch D2 (D2) 5-6 [Note: The table lists D2 as 5]
First obtuse marginal branch OM1 (OM1) 6-7 [Note: The table lists OM1 as 6]
Second obtuse marginal branch OM2 (OM2) 7-8 [Note: The table lists OM2 as 7]
Intermediate marginal (ramus intermedius) IM (IM) 8-9 [Note: The table lists IM as 8]
Right coronary artery RCA (RCA) 9-10 [Note: The table lists RCA as 9]
Right posterior descending artery R-PDA (R-PDA) 10-11 [Note: The table lists R-PDA as 10]
Right posterior lateral artery R-PLA (R-PLA) 11-[Note: The table lists R-PLA as 12]
Left posterior descending artery L-PDA (L-PDA) 12-[Note: The table lists L-PDA as 13]
Left posterior lateral artery L-PLA (L-PLA) 13-[Note: The table lists L-PLA as 14]
Other - -
Training and Augmentation Pipeline:
Augmentation is applied during training based on the nnU-Net’s default pipeline, including: "additive Gaussian noise (p=0.1, sigma in [0, 0.01]), Gaussian blur (p=0.2, sigma in [0.5, 1.0]), multiplicative brightness (p=0.15, [0.75, 1.25]), contrast (p=0.15, [0.75, 1.25]), simulated low resolution (p=0.25, downsample factor [0.5, 1.0]), inverted gamma (p=0.1, gamma in [0.7, 1.5]) and gamma (p=0.3, gamma in [0.7, 1.5]). Notably,
Axis flipping over coronal and sagittal planes are included, while rotations and axial flipping are omitted because
coronary anatomy has a genuine superior–inferior orientation."
The training schedule defines an epoch as 250 training and 50 validation batches drawn with replacement.
Most methods use Adam (with weight decay 10-8) and a polynomial learning-rate decay (power 0.9), while nnU-Net used SGD.
Inference Pipeline:
For inference, the process varies by method type: whole-volume methods used a single forward pass, while patch- and crop-based methods are tiled with a sliding window of the method’s own training patch size at 50% overlap.
Overlapping predictions are combined using a Gaussian-weighted average with per-axis sigma = 0.125 times patch size.
A critical step is Axis mirroring test-time augmentation,
where predictions are averaged over all four flip combinations of the two in-plane axes, matching the flip axes used during training.
Finally, post-processing is applied at full original resolution and remains identical for all methods: a threshold at 0.5, followed by removal of connected components smaller than 100 voxels.
Method-Specific Implementation Notes:
The framework required
Improvements for AI systems
Based on my rigorous analysis of the ImageCAS-X paper, here are the specific improvements for AI systems and what these systems can achieve:
A. Targeted Anatomical Segmentation (Segment-Level Training)
-
Improvement: The system is trained not merely on a binary lumen mask, but is explicitly segmented by the defined 18-segment model. The input data structure allows for loss functions or weighted sampling that targets specific anatomical regions (e.g., Left Circumflex vs. Right Coronary Artery).
-
What the System Can Do: The AI can be trained to understand and maintain connectivity within distinct segments, reducing the risk of
phantom
connections between different arterial branches, leading to more reliable local quantification of plaque burden in specific territories.
B. Topology-Aware Training (Prioritization of Structural Integrity)
-
Improvement: The system is optimized using metrics like ** beta err (Betti number error)** and clDice, which penalize topological errors such as broken vessels or unintended loops, rather than just volumetric overlap (DSC).
-
What the System Can Do: The AI can reliably detect subtle vessel breaks or spurious connections that traditional Dice-based models ignore. This is critical for accurate subsequent computational modeling (e.g., CT-FFR), ensuring the model's predicted geometry accurately reflects physical reality, not just statistical mean overlap.
C. Contextual/Stratified Training
-
Improvement: The system can be trained to recognize and adapt its performance based on clinical context descriptors (coronary dominance, disease presence, vessel diameter).
-
What the System Can Do: Instead of providing a single aggregate score, the AI can provide a segmented risk assessment. For instance, it can be trained to flag that
segmentation accuracy drops significantly in small side branches (e.g., D1/D2) compared to main trunks,
allowing clinicians to prioritize areas where automated performance is inherently limited.
A. Robust Performance Benchmarking
-
Improvement: The system is benchmarked against inter-observer variability derived from the ImageCAS-X dataset, rather than just against a single
ground truth
model (as was common in previous studies). -
What the System Can Do: The AI' can be validated to ensure its performance falls within the statistically acceptable range of human expertise. This provides a rigorous, clinically relevant measure of reliability for deployment in clinical trials and ensures that the automated solution is not significantly worse than a highly experienced human expert.
B. Localized Performance Analysis
-
Improvement: The system is evaluated using local DSC sampling across different points along the vessel (e.g., every fifth vertex).
-
What the System Can Do: The AI can provide dynamic performance metrics, indicating how its accuracy changes as it moves from the aortic ostium to the distal end of a vessel. It can flag areas where contrast concentration diminishes or diameter reduces, allowing researchers to understand why the model struggles in specific physiological regions.
A. Automated Plaque and Adipose Quantification
-
Improvement: The high-quality, precise segmentation allows for the reliable subsequent segmentation of perivascular adipose tissue (PVAT) and heterogeneous plaque composition (calcified vs. non-calcified).
-
What the System Can Do: The AI can generate accurate 3D models of plaque burden and measure the volume/attenuation of surrounding fat, enabling precise quantification that was previously hampered by poor segmentation accuracy in older datasets.
B. High-Fidelity Computational Modeling
-
Improvement: The generated mesh surfaces and centerlines provide a geometrically exact representation of the coronary tree.
-
What the System Can Do: The AI enables highly accurate simulation of hemodynamic factors (e.g., CT-FFR) and provides precise geometric data for automated stent planning algorithms, ensuring that interventions are planned based on the true vessel anatomy rather than approximated contours.
Abstract
Accurate segmentation of the coronary vessel lumen is a prerequisite for quantitative assessment of atherosclerotic plaque and perivascular adipose tissue in coronary computed tomography angiography (CCTA). Cardiologists rely on semi-automated methods for this task because manual vessel tracing and segmentation are labour-intensive. Although many automated methods have been proposed, their validation remains limited by the lack of large, high-quality publicly available datasets. We provide a new dataset of voxel-wise annotations of the vessel lumen and coronary segments, alongside centerlines, and mesh surfaces for 800 scans from the publicly available ImageCAS dataset. Using this dataset, we benchmark established lumen segmentation methods against inter-observer variability, stratifying performance by disease, image quality, coronary dominance, coronary segment, vessel diameter, and lumen attenuation. These labels allow segmentation accuracy to be described in anatomical and clinical context rather than reported as a single aggregate score. The dataset supports the development and validation of methods for lumen segmentation, plaque and perivascular quantification, and haemodynamic modelling.
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models