Soft-Attention Improves Skin Cancer Classification Performance

arXiv:2105.03358 · eess.IV, cs.CV, cs.LG · Submitted 2026-08-12 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Soft-Attention Improves Skin Cancer Classification Performance".

Jane: The paper was written by Soumyya Kanti Datta, Seyed Mohammad Abuzar Hashemi, Sargur N Srihari and Mingchen Gao from State University of New York, Buffalo.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Welcome back to the arXiv Review, everyone. I’m Tom, and as always, I’m here with my co-host, Jane. Today we’ve got a paper that’s really close to home for a lot of us — it’s called “Soft-Attention Improves Skin Cancer Classification Performance.”

Jane: And Tom, I have to say, the title is pretty much the whole thesis, right? It’s not hiding anything. The authors — Soumyya Kanti Datta, Seyed Mohammad Abuzar Hashemi, Sargur N. Srihari, and Mingchen Gao from SUNY Buffalo — they’re basically saying, hey, if you add this one mechanism to your neural network, skin cancer classification gets better.

Tom: Exactly. And that mechanism, soft attention, is something we should break down because it’s not as scary as it sounds. In plain terms, it’s like telling the network, “Hey, don’t look at the whole picture. Focus on the part that actually matters.”

Jane: Right. Think of it like a radiologist looking at a scan. They don’t stare at the entire image equally — they zoom in on the suspicious spot. Soft attention does that for a neural network. It learns which pixels are important and boosts their signal while suppressing the noise, like hair or veins in a skin lesion image.

Tom: And that’s huge because skin lesions are tricky. Malignant and benign ones look almost identical to the untrained eye. The authors mention that low inter-class variation is a big problem. So having a model that can zoom in on the right features is a game-changer.

Jane: Yeah, and they tested this on two big datasets — HAM10000 and ISIC-two thousand seventeen. They took well-known architectures like VGG, ResNet, Inception ResNet v2, and DenseNet, and they added this soft attention block to each one. The results were pretty consistent — attention helped almost every model.

Tom: The biggest win was with Inception ResNet v2. They got precision up to ninety-three point seven percent on HAM10000, which beat the baseline by four point seven percent. That’s not a tiny bump, Jane. That’s a real improvement in a medical context where every percentage point matters.

Jane: And it’s not just about accuracy. The authors also show that soft attention makes the model more transparent. You can see where the network is looking when it makes a decision. That’s huge for building trust with dermatologists who might use this as a tool.

Tom: So we’ve got a paper that’s not only improving performance but also making the black box a little less black. I’m excited to dig into the actual method in the next segment, but first, let’s just sit with that title — it’s simple, it’s honest, and it delivers.

Jane: It does. And honestly, that’s refreshing. Sometimes papers overpromise. This one just says, we added attention, things got better, here’s the proof.

Tom: Alright, stick around. Next up, we’re going to look at how they actually built this soft attention block and why it works so well.

Summary: Tom: So we’re back, and we’re still on “Soft-Attention Improves Skin Cancer Classification Performance.” Jane, let’s get into the meat of it — how does this soft attention thing actually work?

Jane: Okay, so the authors describe it pretty clearly. You have a feature tensor, which is basically the output of a convolutional layer — a stack of feature maps that represent different patterns the network has detected. They pass that through a three dee convolution layer with a bunch of filters, then apply a softmax function to turn those outputs into attention maps.

Tom: And those attention maps are like heat maps, right? They show which areas of the image the network thinks are important.

Jane: Exactly. They generate sixteen of these attention maps, then aggregate them into one unified map. That unified map acts as a weighting function. They multiply it with the original feature tensor, and that scales up the important features and scales down the irrelevant ones.

Tom: And there’s a learnable scalar called gamma, right? I remember reading that they initialize it from zero point zero one so the network slowly learns how much attention it needs.

Jane: Right. That’s a nice touch because it means the network isn’t forced to use attention right away. It eases into it during training. And then they concatenate the attention-scaled features with the original features as a residual branch. So the network can still use the original information if it needs to, but it also has this enhanced version.

Tom: That’s clever. It’s like giving the network a magnifying glass but also letting it look at the full picture if it wants. And the results speak for themselves. On HAM10000, the Inception ResNet v2 with soft attention hit an AUC of ninety-eight point four percent. That’s really high.

Jane: And on ISIC-two thousand seventeen they got sensitivity up to ninety-one point six percent, which is a three point eight percent improvement over the baseline. Sensitivity is crucial in cancer detection because it measures how well the model catches actual positive cases. You don’t want to miss a melanoma.

Tom: Right, because missing a cancer is way worse than a false alarm. And they actually discuss that trade-off. Their model has slightly lower specificity compared to some baselines, but they argue that sensitivity is more important in this context. I think that’s a defensible position.

Jane: Definitely. And they also compared their attention maps to Grad-CAM, which is a common technique for visualizing what a network is looking at. They found that soft attention maps were more focused on the actual lesion area, while Grad-CAM sometimes spread out onto healthy skin.

Tom: So it’s not just about better numbers — it’s about the model looking at the right things. That’s the kind of evidence that makes me believe this could actually help in a clinical setting.

Jane: And that’s what we’ll dig into next — what this means for real-world applications and how it might change the way dermatologists work.

Tom: Stay with us.

Improvements: Tom: Alright, we’re back on “Soft-Attention Improves Skin Cancer Classification Performance,” and I want to talk about what this actually improves in practice. Jane, you mentioned clinical settings — let’s get into that.

Jane: Yeah, so the big improvement here is twofold. First, you’ve got the raw performance boost — higher precision, higher sensitivity, better AUC scores across multiple architectures. But second, and maybe more importantly, you’ve got the transparency. The soft attention mechanism gives you a visual map of where the network is focusing.

Tom: And that matters because a dermatologist isn’t going to trust a black box that just says “this is malignant.” They want to see why. With soft attention, you can show them the heat map and say, “Look, the model is focusing on the lesion, not the hair or the veins around it.”

Jane: Exactly. And the authors make a point about this — they compare their attention maps to Grad-CAM, and the soft attention maps are tighter, more focused on the actual lesion. That’s a big deal for building trust.

Tom: Let’s bring in our senior researcher, Lu, from Tsinghua. Lu, what do you think about the broader implications here?

Lu: Thanks, Tom. I think the most exciting part is that this isn’t a brand-new architecture. They took existing, well-known models and added a relatively simple block. That means it’s easy to retrofit into systems that are already deployed. You don’t need to retrain from scratch — you just add this attention module and fine-tune.

Jane: That’s a great point. And it’s not just skin cancer. The authors mention that this could be applied to other medical imaging tasks. If you’re looking at X-rays or MRIs, you also have the problem of irrelevant features dominating the image.

Tom: Let’s get Meng’s take. Meng, you’re the engineer — how practical is this to implement?

Meng: Honestly, it’s very practical. The soft attention block is just a few layers — a three dee convolution, a softmax, a multiplication, and a concatenation. Any engineer who’s worked with Keras or PyTorch can implement this in an afternoon. The training time is the real cost, but even that isn’t crazy — they trained for one hundred fifty epochs with early stopping.

Tom: So it’s not a research toy. This is something that could actually ship.

Meng: Absolutely. And the fact that they tested it on multiple architectures — VGG, ResNet, DenseNet, Inception — means you have a good sense of how it generalizes. It’s not a fluke that only works on one model.

Jane: And let’s not forget the data side. They dealt with class imbalance by over-sampling and under-sampling, which is a real-world problem. Anyone working with medical data knows that you rarely have equal numbers of each condition.

Tom: So we’ve got a method that’s simple, effective, and transparent. What’s not to like?

Lu: Well, one thing to note is that it’s not magic. The improvement varies by architecture. For DenseNet, the gain was smaller — around zero point five percent in precision. So it’s not a universal silver bullet, but it consistently helps.

Tom: Good to have that nuance. Alright, we’re going to wrap up with our final thoughts and hear from Lalam about the bigger picture.

Conclusion: Tom: And we’re back for the final stretch on “Soft-Attention Improves Skin Cancer Classification Performance.” Jane, let’s wrap this up.

Jane: Sure. So the takeaway is pretty clear — adding soft attention to existing deep learning models improves skin cancer classification, both in terms of accuracy and interpretability. The best result was with Inception ResNet v2, hitting ninety-three point seven percent precision on HAM10000 and ninety-one point six percent sensitivity on ISIC-two thousand seventeen.

Tom: And it’s not just the numbers. It’s the fact that the model shows its work. You can see where it’s looking, and it’s looking at the right places. That’s a big step toward getting these tools into actual clinics.

Jane: Let’s bring in Lalam for a final thought on the cultural and societal impact.

Lalam: Thank you, Tom and Jane. I think the most profound impact here is democratization. Skin cancer is a global problem, and dermatologists aren’t equally available everywhere. A model like this, running on a smartphone with a dermoscopy attachment, could provide a first-line screening in underserved areas. The attention maps also empower patients — they can see what the model sees and have a more informed conversation with their doctor.

Tom: That’s a beautiful way to put it. It’s not just about the algorithm; it’s about access.

Meng: And from an engineering standpoint, the fact that it’s a drop-in module means it can be integrated into existing telemedicine platforms without a massive overhaul. That lowers the barrier to adoption.

Lu: I’d add that this also opens the door for more research into attention mechanisms in other medical domains. If it works for skin lesions, it’s worth trying for retinal scans, pathology slides, even CT scans.

Jane: So we’re saying goodbye to this paper, but the ideas in it are going to stick around. Soft attention is a tool we’ll be seeing more of.

Tom: Agreed. And with that, we’re wrapping up “Soft-Attention Improves Skin Cancer Classification Performance.” Thanks to the authors for their work, and thanks to all of you for listening. We’ll be back soon with another paper. Until then, keep learning, keep questioning, and take care of your skin.

Jane: See you next time, everyone.

Soumyya Kanti Datta, Seyed Mohammad Abuzar Hashemi, Sargur N Srihari, Mingchen Gao

State University of New York, Buffalo

eess.IV, cs.CV, cs.LG

Submitted: 2026-08-12

Comments: 8 pages, 9 figures, 4 tables

Journal ref: Publication date: 2021/9/21; Conference: Interpretability of Machine Intelligence in Medical Image Computing; Issue: 978-3-030-87444-5; Pages: 13-23; Publisher: Springer, Cham

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 61/100

Key concepts

Soft Attention
This mechanism helps a neural network focus on important parts of an image instead of looking at the whole picture. It works by generating attention maps that act as weighting functions to scale up important features and suppress irrelevant noise, like hair or veins in skin lesions.
Sensitivity
In cancer detection, sensitivity measures how well a model catches actual positive cases. The authors found that soft attention improved sensitivity on ISIC-two thousand seventeen by 3.8 percent, which is crucial because missing a cancer is worse than a false alarm.
Interpretability
Soft attention makes the model more transparent by providing visual heat maps showing where the network is focusing when making a decision. This allows users, like dermatologists, to see why a diagnosis was made, helping build trust in the AI tool.

Terminology

Summary

Summary

This paper investigates the effectiveness of the Soft-Attention mechanism in deep neural architectures for the classification of skin lesions. The central aim of Soft-Attention is to boost the value of important features and suppress the noise-inducing features. The authors compare the performance of VGG, ResNet, Inception ResNet v2, and DenseNet architectures with and without the Soft-Attention mechanism.

The experiments are performed on two datasets: the HAM10000 dataset, which consists of 10015 dermatoscopic images of size 450 × 600 with 7 diagnostic categories (Melanoma, Melanocytic Nevi, Basal Cell Carcinoma, Actinic Keratosis, and Intra-Epithelial Carcinoma, Benign Keratosis, Dermatofibroma, Vascular lesions), and the ISIC 2017 dataset, which consists of 2600 images of size 767 x 1022, with the training dataset containing 2000 images of 3 categories: benign nevi, seborrheic keratosis, and melanoma. The data in both datasets is cleaned to remove class imbalances through over-sampling and under-sampling, and images are normalized by dividing each pixel by 255.

The Soft Attention module takes the feature tensor (t) flowing down the deep neural network as input. The feature tensor t ∈ Rh×w×d is input to a 3D convolution layer with weights Wk ∈ Rh×w×d×K, where K is the number of 3D weights. The output is normalized using softmax to generate K = 16 attention maps. These attention maps are aggregated to produce a unified attention map that acts as a weighting function α. This α is multiplied with t to attentively scale salient feature values, further scaled by γ, a learnable scalar. Finally, the attentively scaled features are concatenated with the original feature t in the form of a residual branch. During training, γ is initialized from 0.01 so the network can slowly learn to regulate the amount of attention required.

For the model setup, the Adam optimizer with a 0.01 learning rate and 0.1 epsilon is used, with batch normalization added after each layer. For the HAM10000 dataset, an output layer with 7 hidden units is implemented, followed by a softmax activation unit. The soft attention layer is integrated differently across architectures: in Inception ResNet v2, it is added to the Inception Resnet C block where the feature size is 8 x 8, followed by a maxpool layer, concatenated with the filter concatenate layer, then a relu activation and a 0.5 dropout layer. In DenseNet201, the soft attention layer is added to the 4th dense block where the feature map size is 7 x 7. In ResNet34, it is added after the 3rd convolution block (feature map size 28 x 28), and in ResNet50 after the 5th convolution block (feature map size 7 x 7), with the same integration procedure as Inception ResNet v2. In VGG16, the soft attention layer is added after conv layer 4 where the feature map size is 28 x 28. The networks are trained for 150 epochs with early stopping patience of 30 (Inception ResNet v2), 35 (DenseNet201), and 300 epochs with early stopping patience of 65 (VGG16).

The loss function used is categorical cross entropy, and the models are evaluated using Precision, Sensitivity, Accuracy, Specificity, and AUC scores.

From the ablation analysis on the HAM10000 dataset, the authors found that Inception ResNet v2 when coupled with Soft Attention (IRv2+SA) shows significant improvements, with a precision and AUC score of 93.7% and 98.4% respectively, which are the highest scores amongst all models. Soft Attention boosts the performance of IRv2 by 3.2% in terms of precision compared to the original IRv2 model. Soft Attention also boosts the precision of DenseNet201, ResNet34, ResNet50, and VGG16 by 0.5%, 0.8%, 1.2%, and 2% respectively. The authors selected an 85/15% training/testing split as the model with 85% training data outperforms the model with 80% and 70% training data by 2.2% and 2.6% respectively.

The proposed approach outperforms the baseline by 4.7% in terms of precision on the HAM10000 dataset, achieving a precision of 93.7%. In terms of AUC scores, it outperforms state-of-the-art models by 0.5% to 4.3%.

On the ISIC-2017 dataset, the authors tested two configurations: IRv25x5 +SA and IRv212x12 +SA, where the attention layer was added when the feature map size is 5x5 and 12x12 respectively. The model IRv25x5 +SA outperforms IRv212x12 +SA in terms of AUC scores, Accuracy, and Specificity by 2.4%, 0.6%, and 12.2% respectively, whereas IRv212x12 +SA outperforms IRv25x5 +SA in terms of Sensitivity by 2.9%. When IRv25x5 +SA is compared with the ARL-CNN50 baseline model, it performs on par in terms of AUC score but outperforms it in accuracy and Sensitivity by 3.6% and 3.8% respectively, achieving 91.6% sensitivity. However, ARL-CNN50 takes the upper hand in Specificity by 3.4%.

Qualitatively, the Soft Attention heat maps are compared with Grad-CAM heatmaps. The authors observe that the SA map focuses on the main part of the lesion area whereas the Grad-cam heatmap is slightly shifted towards top left and is also spread out on the uninfected area of skin, concluding that the Soft Attention maps are focused more on the relevant locations of the image compared to Grad-CAM heatmaps.

The authors conclude that the Soft Attention mechanism eliminates the need for external mechanisms like GradCAM, internally provides the location of where the model focuses while categorizing a disease, boosts the performance of the main network, and naturally deals with image noise internally. The model can be implemented in dermoscopy systems to assist dermatologists and can be easily implemented to classify data from other medical databases.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in an AI system, along with what the improved system can do:

Improvements to the AI System:

  1. Integrate a Soft-Attention Module into the Feature Extraction Pipeline:
  • Insert a 3D convolution layer (with K=16 filters) after a designated convolutional block (e.g., after the 5th block in ResNet50, or the Inception-ResNet-C block in Inception ResNet v2) where the feature map size is 7x7 or 8x8.

  • Apply a softmax function to the 3D convolution output to generate K attention maps, then aggregate them into a single unified attention map (α).

  • Multiply the original feature tensor (t) by α, then scale by a learnable scalar γ (initialized at 0.01) to create an attentively scaled feature map.

  • Concatenate this scaled map with the original feature tensor as a residual branch, followed by a ReLU activation and a 0.5 dropout layer for regularization.

  1. Modify the Network Architecture for Optimal Attention Placement:
  • For Inception ResNet v2: Add the Soft-Attention block after the Inception-ResNet-C block (feature map 8x8), followed by a 2x2 max-pool, then concatenate with the filter concatenation layer.

  • For ResNet34/50: Add the block after the 3rd (28x28) or 5th (7x7) convolution block respectively, followed by a 2x2 max-pool and concatenation with the standard max-pool output.

  • For DenseNet201: Add the block after the 4th dense block (7x7 feature map).

  • For VGG16: Add the block after conv layer 4 (28x28 feature map).

  1. Implement Class-Balanced Training with Data Augmentation:
  • Apply over-sampling and under-sampling to ensure equal image counts per class in the training set (addressing the class imbalance in HAM10000 and ISIC-2017).

  • Normalize images by dividing pixel values by 255 to keep them in [0,1].

  • Use a train-test split of 85/15 (as the paper shows this outperforms 80/20 and 70/30 splits by 2.2% and 2.6% respectively).

  1. Optimize Training Hyperparameters:
  • Use the Adam optimizer with a learning rate of 0.01 and epsilon of 0.1.

  • Add batch normalization after each layer.

  • Train for 150 epochs (300 for VGG16) with early stopping patience of 30 (65 for VGG16).

  • Use categorical cross-entropy loss for multi-class classification.

  1. Add an Internal Explainability Mechanism:
  • Instead of relying on external Grad-CAM, use the Soft-Attention maps directly to visualize which parts of the image the model focuses on, making the classification process transparent to medical personnel.

What the Improved AI System Can Do:

  1. Achieve Higher Classification Accuracy on Skin Lesion Images:
  • On HAM10000: Reach a precision of 93.7% and an average AUC of 98.4%, outperforming the baseline (Rezvantalab et al.) by 4.7% in precision and 0.5-4.3% in AUC.

  • On ISIC-2017: Achieve an accuracy of 90.4%, sensitivity of 91.6%, and AUC of 95.9%, improving sensitivity by 3.8% over the ARL-CNN50 baseline.

  1. Focus on Clinically Relevant Image Regions:
  • The system will automatically suppress noise-inducing features (e.g., hair, veins, contrast variations) and highlight the lesion area, as demonstrated by the attention maps in Figure 7 and 9, which are more focused than Grad-CAM heatmaps.
  1. Provide Class-Specific Performance Improvements:
  • For AKIEC, BCC, DF, and NV classes, precision improves by 17%, 3%, 33%, and 4% respectively compared to the original Inception ResNet v2.

  • For BKL and MEL, AUC improves by 1.2% and 0.9% respectively.

  1. Generalize to Other Medical Imaging Tasks:
  • The Soft-Attention mechanism is architecture-agnostic and can be applied to any CNN (VGG, ResNet, DenseNet, Inception) for other medical image classification problems (e.g., histopathology, radiology), as the paper notes it can be easily implemented to classify data from other medical databases.
  1. Operate Without External Visualization Tools:
  • The system internally generates attention maps that show the exact location of the lesion, eliminating the need for post-hoc explainability methods like Grad-CAM, while simultaneously boosting performance.

Abstract

In clinical applications, neural networks must focus on and highlight the most important parts of an input image. Soft-Attention mechanism enables a neural network toachieve this goal. This paper investigates the effectiveness of Soft-Attention in deep neural architectures. The central aim of Soft-Attention is to boost the value of important features and suppress the noise-inducing features. We compare the performance of VGG, ResNet, InceptionResNetv2 and DenseNet architectures with and without the Soft-Attention mechanism, while classifying skin lesions. The original network when coupled with Soft-Attention outperforms the baseline[16] by 4.7% while achieving a precision of 93.7% on HAM10000 dataset [25]. Additionally, Soft-Attention coupling improves the sensitivity score by 3.8% compared to baseline[31] and achieves 91.6% on ISIC-2017 dataset [2]. The code is publicly available at github.

Sources

Related papers