A Neighborhood Attention Transformer Network for Enhanced 3D Segmentation of the Left Anterior Descending Artery

arXiv:2608.12274 · cs.CV, cs.AI · Submitted 2026-08-12 · Read on arXiv

Rafi Ibn Sultan, Chengyin Li, Yiannos Demetriou, Ahmed I. Ghanem, Joshua P. Kim, Justine Cunningham, Hassan Bagher-Ebadian, Dongxiao Zhu, Kundan S. Thind

Wayne State University · Henry Ford Health · Alexandria University · Michigan State University · Oakland University

cs.CV, cs.AI

Submitted: 2026-08-12

Updated: 2026-08-13

Comments: Acceteed by Medical Physics 2026

Code: https://github.com/rafiibnsultan/NA_UNETR

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

Terminology

Summary

Summary

This paper introduces NA-UNETR, a 3D transformer-based segmentation framework designed to improve the segmentation of the Left Anterior Descending (LAD) artery in free-breathing, non-contrast CT images for cardiac dose sparing in thoracic radiotherapy. The LAD is described as extremely small, exhibits poor soft-tissue contrast, and varies substantially across patients, making manual delineation challenging, with manual coronary artery Dice scores ranging from 0.10 to 0.53 on non-contrast CT.

The proposed architecture integrates Neighborhood Attention (NA) and Dilated Neighborhood Attention (DiNA) blocks within a UNETR-style backbone. NA restricts token interactions to local k×k×k windows, introducing a spatial inductive bias that promotes geometrically coherent responses among adjacent voxels, while DiNA expands the receptive field by sampling dilated neighborhoods, enabling integration of longer vessel segments across slices without resorting to fully global interactions. The encoder is divided into four stages with NAT blocks (counts set to (3, 4, 6, 18, 5) and kernel sizes (7, 7, 7, 3, 3)), each preceded by a residual convolution layer. The decoder follows a symmetric U-shaped design with skip connections, deconvolution layers, and residual blocks.

To address the limited availability of annotated LAD data, the model is pretrained on 1,000 CTA volumes from the ImageCAS dataset, then fine-tuned on 20 institutional free-breathing CT scans (LAD-SEG dataset) using Low-Rank Adaptation (LoRA) with rank r=8, where only θdec, A, and B were updated, while all other pretrained weights remained fixed. The LAD-SEG dataset has an average foreground ratio of 1.7 × 10−5 and an average of 540 foreground voxels per case.

The loss function combines Dice-Focal loss with a Hausdorff boundary loss, balanced dynamically via homoscedastic uncertainty weighting. The total loss is computed as Ltotal = (1/2σ12)LDice-Focal + (1/2σ22)L̃Hausdorff + log σ1 + log σ2, where σ12 and σ22 are learnable variance parameters. Preprocessing includes intensity clipping to [−200, 400] HU, contrast enhancement with gamma adjustment between 1.6 and 1.8, Savitzky–Golay filtering, and artery-centric patch sampling with a 1:1 foreground-to-background ratio. Postprocessing retains the largest connected component, removes components smaller than 64 voxels, and fills holes.

On the LAD-SEG dataset, NA-UNETR achieved 45.64% Dice, 38.16 mm HD95, and 10.01 mm ASD, improving Dice by 3.10 percentage points over nnU-Net and reducing HD95 by 2.96 mm relative to Swin UNETR. It also achieved the highest clDice at 44.39%. On ImageCAS, it achieved 79.49% Dice, 8.89 mm HD95, and 1.02 mm ASD, with statistically significant improvements (p < 0.05) confirmed by Mann–Whitney U testing. On LAD-SEG, differences were not statistically significant (p > 0.05) due to the small sample size (n=20).

Ablation studies confirmed that residual blocks, variable kernels, and uncertainty-weighted loss each contributed to performance. Removing residual blocks decreased DSC from 45.64% to 43.01%, and using a fixed 3×3×3 kernel reduced DSC by roughly 2.5%. LoRA rank sensitivity showed r=8 as optimal, with performance degrading at r=16. Training on LAD-SEG alone reduced DSC from 45.64% to 36.39%, confirming the value of pretraining. Computational analysis showed NA-UNETR has 19.6M trainable parameters, 314.1B FLOPs, and 1.33 s inference time, comparable to Swin UNETR but more memory-efficient (4.17 GB peak VRAM).

The authors conclude that NA-UNETR effectively balances local precision and global context, enabling accurate segmentation of thin, low-contrast LAD structures, and that its integration of Neighborhood Attention, uncertainty-weighted loss, and LoRA-based fine-tuning demonstrates a robust and computationally efficient framework for substructure-level cardiac segmentation in radiotherapy planning. They note the approach is not yet suitable for clinical deployment and that further refinement, larger annotated datasets, and validation across multiple institutions will be essential.

Improvements for AI systems

Improvements to AI Systems:

  1. Adaptive Multi-Scale Attention with Local Inductive Bias – Integrate Neighborhood Attention (NA) and Dilated Neighborhood Attention (DiNA) blocks into any vision transformer backbone to enforce spatial coherence for small, low-contrast structures. The system can now segment thin tubular anatomies (e.g., arteries, vessels, ducts) with higher precision by restricting token interactions to local windows while using dilation to capture long-range continuity without global attention overhead.

  2. Uncertainty-Weighted Multi-Component Loss – Replace fixed-weight loss combinations with homoscedastic uncertainty weighting (learnable σ1, σ2) that dynamically balances Dice-Focal and Hausdorff boundary losses. The improved system automatically adapts to datasets with extreme class imbalance (foreground ratio 1.7×10−5) and boundary noise, reducing false positives and improving boundary fidelity without manual hyperparameter tuning.

  3. LoRA-Based Parameter-Efficient Transfer Learning for Sparse Medical Data – Apply Low-Rank Adaptation (rank r=8) to fine-tune only decoder and LoRA matrices (A, B) on a pretrained 3D transformer, keeping encoder weights frozen. This enables robust segmentation with as few as 20 annotated volumes, reducing overfitting and training time while retaining pretrained anatomical knowledge. The system can generalize to new substructures with minimal labeled data.

  4. Artery-Centric Patch Sampling with Contrast Enhancement – Implement preprocessing that clips intensity to [−200, 400] HU, applies gamma adjustment (1.6–1.8), Savitzky–Golay filtering, and samples patches with a 1:1 foreground-to-background ratio. The improved system now focuses computational effort on vessel-rich regions, boosting detection of tiny structures (average 540 voxels per case) and improving Dice by 3.1 points over nnU-Net without extra model capacity.

  5. Post-Processing with Connected Component and Hole-Filling – Add automatic retention of the largest connected component, removal of components <64 voxels, and morphological hole-filling. The system now eliminates spurious small false positives and ensures topological continuity, increasing clDice (centerline Dice) to 44.39% for thin vessels, which is critical for radiotherapy dose sparing.

  6. Hybrid Local-Global Encoder Design with Variable Kernel Sizes – Use a four-stage encoder with NA kernel sizes (7,7,7,3,3) and residual convolution layers before each stage. The improved system balances local detail (small kernels in deep stages) and global context (large kernels in shallow stages), reducing HD95 by 2.96 mm vs. Swin UNETR while using 4.17 GB peak VRAM—enabling deployment on standard clinical GPUs.

  7. Pretraining on Contrast-Enhanced CT for Non-Contrast Fine-Tuning – Pretrain on large CTA datasets (e.g., ImageCAS with 1,000 volumes) then fine-tune on non-contrast CT with LoRA. The system now transfers vascular appearance knowledge across modalities, improving Dice from 36.39% (training from scratch) to 45.64% on free-breathing non-contrast scans, enabling use in routine radiotherapy planning without contrast agents.

  8. Statistical Validation and Uncertainty-Aware Inference – Incorporate Mann–Whitney U testing for performance comparison and report confidence intervals. The improved system can flag low-confidence segmentations (e.g., when Dice 40 mm) for manual review, addressing the current limitation of non-significant results on small datasets (n=20) and guiding safe clinical adoption.

Abstract

Background: Accurate segmentation of the Left Anterior Descending (LAD) artery in 3D free-breathing, non-contrast CT is critical for cardiac dose sparing in thoracic radiotherapy. The LAD is extremely small, has poor soft-tissue contrast, and varies substantially across patients; even manual contours show limited inter-observer agreement, underscoring the ambiguity of the vessel boundaries. Purpose: To develop a transformer-based framework that improves LAD delineation in low-contrast, imbalanced CT through local-global context modeling and uncertainty-guided optimization. Methods: We propose NA-UNETR, a 3D transformer-based segmentation model whose Neighborhood Attention (NA) and Dilated NA (DiNA) blocks jointly capture fine structural detail and long-range context. Given the scarcity of annotated LAD data, the model is pretrained on 1,000 CTA volumes of general coronary anatomy and fine-tuned with LoRA-based parameter-efficient adaptation on 20 free-breathing institutional CT scans. A composite Dice-Focal and Hausdorff loss, dynamically balanced via homoscedastic uncertainty, improves overlap and boundary accuracy. Results: NA-UNETR reached 45.64% Dice, 38.16 mm HD95, and 10.01 mm ASD, improving Dice by 3.10 percentage points over nnU-Net and reducing HD95 by 2.96 mm relative to Swin UNETR, with the strongest boundary accuracy among all models and improved centerline stability. On ImageCAS it achieved 79.49% Dice, 8.89 mm HD95, and 1.02 mm ASD. Ablations confirmed that residual blocks, variable kernels, and uncertainty-weighted loss each contributed. Conclusions: NA-UNETR balances local precision and global context for thin, low-contrast LAD structures, offering a computationally efficient framework for substructure-level cardiac segmentation in radiotherapy planning.

Sources

Related papers