Uncertainty-Aware and Explainable Ensemble Deep Learning Framework for Multi-Class Skin Lesion Classification

arXiv:2608.11280 · eess.IV, cs.AI, cs.CV, cs.LG · Submitted 2026-08-11 · Read on arXiv

Rofiqul Islam, Lilatul Ferdouse

Wilfrid Laurier University

eess.IV, cs.AI, cs.CV, cs.LG

Submitted: 2026-08-11

Updated: 2026-08-13

Comments: 5 pages, 3 figures, IEEE AIBThings 2026

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: This paper proposes an uncertainty-aware and explainable deep learning framework for multi-class skin lesion classification.

Terminology

Summary

This paper proposes an uncertainty-aware and explainable deep learning framework for multi-class skin lesion classification. The framework combines a vision transformer model (MaxViT-Tiny) with CNN-based models (ConvNeXt-Tiny and EfficientNetV2-B0) through deep ensemble learning. Monte Carlo (MC) Dropout estimates predictive uncertainty and identifies unreliable predictions, while Grad-CAM++, an explainable AI (XAI) technique, provides visual explanations by highlighting lesion regions that influence model decisions. Evaluated on the HAM10000 dataset, the framework achieves 96% accuracy and 99% ROC-AUC under uncertainty-aware filtering (entropy < 1.0, confidence ≥ 0.7), with macro-average precision, recall, and F1-score of 94%, 95%, and 95%, respectively, and 96% weighted-average scores across all three metrics. The results demonstrate accurate, interpretable, and uncertainty-aware skin lesion classification for trustworthy computer-aided diagnosis.

The main contributions of the paper are: (1) proposing an uncertainty-aware deep ensemble framework combining MaxViT-Tiny with ConvNeXt-Tiny and EfficientNetV2-B0 for robust multi-class skin lesion classification; (2) integrating deep ensemble learning with Monte Carlo Dropout to estimate predictive uncertainty, enabling reliable and confidence-aware clinical decision support; (3) employing Grad-CAM++ to generate visual explanations, improving interpretability and transparency; (4) validating the framework on the HAM10000 dataset, achieving 96% accuracy and 99% ROC-AUC under the uncertainty-aware filtering criterion (entropy < 1.0 and confidence ≥ 0.7); and (5) supporting clinically relevant operating scenarios by enabling high-coverage large-scale skin lesion assessment and high-confidence point-of-care diagnosis.

The framework is evaluated on the HAM10000 dataset, which consists of 10,015 RGB dermoscopic images covering seven diagnostic classes: benign keratosis-like lesions (bkl), melanocytic nevi (nv), dermatofibroma (df), melanoma (mel), vascular lesions (vasc), basal cell carcinoma (bcc), and actinic keratoses and intraepithelial carcinoma (akiec). The dataset is divided into training (80%) and testing (20%) subsets using stratified sampling. All images are resized to 224 × 224 pixels and normalized using ImageNet normalization parameters. Data augmentation techniques including random horizontal and vertical flipping, random rotation up to 25°, color jittering, and random affine transformations are applied during training. A weighted sampling strategy is employed to mitigate class imbalance.

For classification, a deep ensemble of three state-of-the-art architectures is constructed: MaxViT-Tiny, ConvNeXt-Tiny, and EfficientNetV2-B0. Each pretrained model is initialized using ImageNet weights and fine-tuned for seven-class skin lesion classification. Focal loss is employed to mitigate class imbalance. For uncertainty estimation, MC Dropout is combined with deep ensemble learning. During inference, each ensemble model performs T = 20 stochastic forward passes, and the predictive probability of each model is obtained by averaging the stochastic outputs. The final predictive distribution is obtained by averaging the outputs of all M = 3 ensemble models. Predictive entropy is used to quantify uncertainty. A selective prediction strategy is adopted where a prediction is considered reliable when entropy < 1.0 and confidence ≥ 0.7. Samples not satisfying these criteria are considered uncertain and can be referred to dermatologists for further assessment.

For explainability, Grad-CAM++ is employed to visualize discriminative image regions influencing classification decisions, using the final feature extraction stage of the MaxViT-Tiny backbone as the target layer. Representative visualizations include correctly classified high-confidence cases for nv (confidence: 0.988, entropy: 0.081), mel (confidence: 0.941, entropy: 0.254), and akiec (confidence: 0.930, entropy: 0.319), as well as an uncertain but correctly classified bkl case (confidence: 0.553, entropy: 1.039).

The experimental results on the uncertainty-filtered test set show class-wise performance with F1-scores ranging from 0.86 to 0.98 across categories. Specifically, bkl, nv, df, vasc, and bcc exhibit F1-scores ranging from 0.95 to 0.98. The mel category remains the most challenging due to high visual similarity to other pigmented lesions, achieving a recall of 0.86. The framework obtains macro-average precision, recall, and F1-score of 0.94, 0.95, and 0.95, respectively, and weighted-average scores of 0.96 across all three metrics.

The confusion matrix analysis shows that most predictions are concentrated along the main diagonal, with the largest confusion occurring between mel and nv (10 mel samples predicted as nv and 30 nv samples predicted as mel), reflecting similar dermoscopic characteristics. Minor misclassifications are also observed among bkl, bcc, and akiec, while df and vasc show very few classification errors.

The paper also discusses two clinical deployment scenarios: large-scale skin lesion assessment with high prediction coverage using a relatively relaxed uncertainty criterion, and point-of-care decision support with high-confidence predictions using stricter uncertainty filtering (entropy < 1.0 and confidence ≥ 0.7). The conclusion states that further validation on larger and more diverse multi-center datasets is required to assess generalizability, and future work will focus on improving computational efficiency, exploring multimodal learning approaches, and developing more robust uncertainty estimation strategies.

Improvements for AI systems

Improvements to AI Systems:

  1. Uncertainty-Aware Selective Prediction: Integrate MC Dropout with deep ensembles to compute predictive entropy and confidence scores. The AI system can abstain from making predictions when entropy ≥ 1.0 or confidence < 0.7, flagging these cases for human review. This reduces false diagnoses in high-risk settings (e.g., melanoma detection) and enables dynamic threshold tuning for different deployment contexts (high-coverage screening vs. high-precision point-of-care).

  2. Hybrid Vision Transformer + CNN Ensemble Architecture: Combine a vision transformer (e.g., MaxViT-Tiny) with multiple CNN backbones (ConvNeXt-Tiny, EfficientNetV2-B0) to capture both global contextual features and local texture patterns. The improved system achieves higher robustness and generalizability across diverse lesion morphologies, outperforming single-model baselines by leveraging complementary inductive biases.

  3. Focal Loss with Weighted Sampling for Class Imbalance: Use focal loss (γ > 2) alongside stratified weighted sampling to prioritize hard-to-classify minority classes (e.g., dermatofibroma, actinic keratoses). The AI system can improve recall on rare but critical conditions (e.g., melanoma) without sacrificing precision on majority classes, as demonstrated by the 0.86 recall for melanoma while maintaining 0.95+ F1 on other classes.

  4. Explainable Grad-CAM++ Visualizations for Clinical Trust: Generate class-discriminative heatmaps from the transformer’s final feature layer. The improved system can provide visual justifications for each prediction, highlighting lesion regions (e.g., pigmented networks, atypical vessels) that drive decisions. This enables dermatologists to verify model reasoning, detect spurious correlations, and build trust in AI-assisted diagnosis.

  5. Uncertainty-Calibrated Confidence Metrics: Combine ensemble variance with MC dropout stochasticity to produce well-calibrated confidence scores. The improved system can distinguish between aleatoric uncertainty (inherent lesion ambiguity) and epistemic uncertainty (model knowledge gaps), allowing clinicians to prioritize cases needing biopsy or additional imaging.

  6. Stratified Multi-Class Confusion Analysis for Error Correction: Leverage the confusion matrix insights (e.g., mel↔nv confusion) to implement targeted data augmentation or contrastive learning for visually similar classes. The AI system can reduce cross-class misclassification by learning finer-grained discriminative features, particularly between melanoma and benign nevi.

What the Improved AI System Can Do:

  • Autonomous Triage: In a dermatology clinic, the system can automatically classify skin lesions with 96% accuracy and 99% ROC-AUC, while flagging 10-15% of uncertain cases for immediate specialist review, reducing workload and diagnostic delays.

  • Explainable Second Opinion: For each high-confidence prediction, it provides a Grad-CAM++ heatmap overlaid on the dermoscopic image, enabling a physician to visually confirm whether the model focuses on clinically relevant structures (e.g., irregular borders, pigment globules) before accepting the recommendation.

  • Adaptive Deployment: In resource-limited settings, it can operate in high-coverage mode (relaxed uncertainty threshold) to screen large populations, and switch to high-precision mode (strict threshold) for confirming high-risk cases before surgery or biopsy.

  • Continuous Learning: By logging uncertain predictions and their outcomes (e.g., biopsy results), the system can be fine-tuned on new data, improving its uncertainty estimates and reducing future abstention rates while maintaining safety.

  • Cross-Dataset Generalization: The ensemble architecture’s robustness allows it to be fine-tuned on smaller, multi-center datasets with fewer than 1,000 images per class, achieving reliable performance where single models fail due to overfitting.

Sources

Related papers