Class Activation Mapping in Explainable Computer Vision: A Method-Centered Review of CNN, Transformer, and Foundation-Model-Era Visual Explanations
AmirHossein Eshghi, Hamid Saadatfar, Seyyed Ali Hoseini, AmirMohsen Eshghi, Siavash Arjomand Bigdeli
University of Birjand · Technical University of Denmark
cs.CV, cs.AI
Submitted: 2026-08-13
Updated: 2026-08-17
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 100/100
The gist: Class activation mapping (CAM) is one of the most widely used visual explanation families in explainable artificial intelligence.
Terminology
Summary
Class activation mapping (CAM) is one of the most widely used visual explanation families in explainable artificial intelligence. Its purpose is intuitive: it converts internal model evidence into a heatmap that highlights the image regions, convolutional channels, tokens, or patches that support a target class or concept. Since the first CAM formulation in 2016, the field has moved far beyond global-average-pooled CNN classifiers. CAM-style methods now include gradient-based post-hoc explanations, gradient-free score and ablation methods, high-resolution upscaling, weakly supervised localization and segmentation, transformer token attribution, causal and debiasing methods, and foundation-model-era approaches that use CLIP, DINO, SAM, or feature-distribution comparisons. This review synthesizes a strict corpus of 57 method-centered papers published from 2016 onward. The paper develops a taxonomy that separates methods by attribution mechanism, architectural dependence, and evaluation objective. It then reviews gradient-based CAMs, recent and hybrid CAM-style methods, and model-based or architecture-aware methods. Across the corpus, the main trend is clear: the field is shifting from explaining one class score in one low-resolution CNN layer toward comparative, multi-layer, probabilistic, token-aware, and foundation-model-aware explanations. At the same time, evaluation remains fragmented. Faithfulness, localization, robustness, computational cost, and human trust are often measured with different protocols. The review therefore emphasizes not only what each method contributes, but also which gap it leaves open and which later methods attempt to close that gap.
The review is method-centered. It does not include papers that merely use Grad-CAM as a visualization figure in an application study. Instead, a paper must generate, refine, evaluate, or theoretically analyze a CAM-style map. This scope is important because CAM is now used in many application papers, but application-only usage does not necessarily change the method. The contribution of this review is to organize the method literature into a coherent development story: from CNN-specific localization, to gradient and score-based attribution, to high-resolution and causal refinements, to transformer and foundation-model-era explanations.
The review has four goals. First, it defines a reproducible search and inclusion protocol for a strict corpus of 57 papers. Second, it clarifies the main technical terms and evaluation metrics used in CAM research. Third, it develops a taxonomy that separates gradient-based, gradient-free, hybrid, high-resolution, token-level, foundation-model-era, and architecture-aware methods. Fourth, it compares representative results where the corpus reports compatible metrics on common datasets. Because CAM evaluation is not standardized, the quantitative figures in this review are deliberately limited to values that are directly reported in the included papers and are interpreted with their protocol constraints.
The literature search followed a method-centered protocol. The primary databases and libraries were IEEE Xplore and IEEE/CVF Open Access, ACM Digital Library, Elsevier ScienceDirect, SpringerLink, and Springer Nature; venue-level cross-checking was performed only within these selected sources. The search was restricted to 2016 onward because the original CAM formulation was published in 2016. The strict venue and publisher filter focused on IEEE/CVF, ACM, Elsevier, Springer/Springer Nature, CVPR, ICCV, ECCV, WACV, ICASSP, ICIP, ACM MM, CHI, ACCV, and related journals represented in the corpus. The inclusion rule states that a method must generate, refine, evaluate, or theoretically analyze CAM-style heatmaps, activation maps, token maps, or patch attribution maps. Application-only Grad-CAM use, survey-only papers, non-vision papers, pre-2016 papers, duplicates, and records outside the selected venue/publisher scope were excluded.
A CAM-style explanation begins with a target output. In a closed-set classifier this target is usually a class logit or class probability. In a detector it may be a box or class head. In a transformer it may be a class token or text-conditioned similarity score. In a representation model it may be a feature vector rather than a class score. Let A k be the kth feature map at a selected layer and y c be the score for class c. The common CAM form is L c = h(Σ k α k c A k), where α k c is the importance of the k-th map and h is usually a non-negative activation such as ReLU. Most CAM papers differ in how they estimate α k c, where they extract A k, and how they evaluate the resulting map. The original CAM method obtains α k c directly from the classifier weight connected to global-average-pooled feature map k. Grad-CAM replaces this architectural weight with the average gradient of the target score with respect to the feature map. Score-CAM avoids gradients by masking the input with upsampled activation maps and using the forward score as the weight. Ablation-CAM measures the effect of removing feature maps. LayerCAM uses positive gradients at individual spatial locations, making it possible to collect fine information from shallower layers. Finer-CAM changes the target itself: instead of explaining only why class c is high, it explains why class c is higher than a similar reference class.
In transformers, the spatial feature map is replaced by patch tokens, class tokens, prompt tokens, or attention matrices. This change is not cosmetic. In CNNs, locality is built into convolution; in ViTs, long-range relations are carried by attention. Transformer CAM-style methods therefore need to decide whether explanation should follow attention weights, relevance propagation, class-to-patch attention, or learned prompt attention. In foundation models the target can become even more flexible. CLIP explanations depend on the text prompt, DINO-based methods use self-supervised semantic affinity, and SAM-assisted methods inject segmentation priors into CAM training.
Evaluation is a persistent source of disagreement in CAM research. A visually sharp map can be unfaithful, and a faithful map can reveal a model shortcut that humans do not want. For this reason, the corpus uses multiple evaluation families. Perturbation metrics such as deletion and insertion test whether removing or restoring salient pixels changes the prediction. Average Drop and Increase in Confidence measure how confidence changes when only the highlighted region is retained. ROAD and ADCC combine perturbation quality with complexity or coherence. Localization metrics compare maps with boxes, masks, or expert regions. In WSSS, mIoU and DSC measure whether CAM-derived pseudo-labels are useful for segmentation. Human-centered studies evaluate whether explanations help users calibrate trust or complete a task.
The taxonomy groups papers by their primary methodological contribution. The largest sections are gradient-free/perturbation CAMs, CNN-based explanations, high-resolution CAMs, and attention-gradient hybrids. This distribution reflects the history of the field. Early work asked how to localize evidence in CNNs. Later work asked how to make the map more faithful, less dependent on gradients, more spatially detailed, more useful for dense prediction, and more suitable for transformer and foundation architectures. The field cannot be ranked along one dimension. A gradient-based method may be fast and flexible but coarse. A gradient-free method may be more faithful but expensive. A high-resolution method may be better for pixel-level pseudo labels but less clearly tied to the model's causal mechanism. A foundation-model method may localize open-vocabulary concepts but depend strongly on prompt wording or pretrained priors.
Gradient-based methods are the reference point for most CAM research. They are attractive because they usually need one forward pass and one backward pass. They are also flexible: a target can be a class score, a caption score, a VQA answer, or another differentiable output. Their weakness is that gradients measure local sensitivity, not necessarily causal contribution. They can vanish, saturate, become noisy in shallow layers, or highlight features that change the score quickly but are not semantically meaningful. The development of gradient-based CAMs is therefore a sequence of attempts to preserve the speed of Grad-CAM while improving resolution, stability, theoretical grounding, or architectural reach.
Grad-CAM made CAM practical for ordinary CNNs. Instead of requiring a global-average-pooling architecture, it computes the importance of each feature map by averaging the gradient of the target score with respect to that map. The weighted feature maps are then summed and passed through ReLU. This design explains why Grad-CAM became widely used: it is simple, model-local, class-discriminative, and can be applied to CNN classifiers, captioning models, VQA systems, and other differentiable visual models. Its limitation is that the final convolutional layer is spatially coarse, and the average gradient can miss localized evidence or become weak when the model is saturated. Grad-CAM++ addresses the first limitation by changing the gradient weighting rule. It uses a weighted combination of positive partial derivatives, so pixels that make strong positive contributions receive more emphasis. This improves multiple-object and multiple-instance explanation because the map is no longer controlled only by a single averaged gradient. The method is still derivative-based, however, so it does not remove the broader problem of noisy or saturated gradients. Integrated Grad-CAM targets that problem from another direction. It computes a path integral of Grad-CAM-style scores, so the explanation is less dependent on the local gradient at only one input point. On PASCAL VOC 2007, the method reported strong EBPG and bounding-box localization values for ResNet-50, but the improvement comes with extra gradient evaluations and with the usual sensitivity of integrated methods to path design. SIS-CAM addresses the noisy gradient problem by iteratively refining saliency maps. It squares gradients, fuses intermediate feature maps with the input image, and applies boundary-aware masking to emphasize critical regions. This iterative approach improves both faithfulness (deletion/insertion metrics) and spatial resolution, making SIS-CAM effective for high-resolution CAM generation in WSOL and WSSS tasks.
Relevance-CAM is motivated by the observation that gradients become unreliable in intermediate layers. The paper uses layer-wise relevance propagation to derive the weighting components instead of ordinary gradients. This makes the explanation more stable in shallow and middle layers, where Grad-CAM may suffer from shattered gradients. The important difference is that Relevance-CAM is not trying only to sharpen the final-layer map; it tries to make class-specific maps reliable across layer depth. In the reported ResNet-50 evaluation, Relevance-CAM improved shallow-layer Average Drop and Average Increase relative to Grad-CAM and Grad-CAM++. Its remaining limitation is that relevance propagation depends on propagation rules, so the method is less plug-and-play than a standard backward gradient. LIFT-CAM asks a more theoretical question: why should the coefficients of a CAM linear combination be chosen by a heuristic? The method formulates each activation map as a feature in an additive attribution model and uses a DeepLIFT-based approximation to estimate SHAP-like coefficients. This gives the CAM weights a clearer additive-attribution meaning while keeping the method efficient. The paper reports higher insertion AUC and lower deletion AUC than Grad-CAM and Grad-CAM++ on ImageNet samples. The main gap is that DeepLIFT-style attribution still depends on reference activations, and the SHAP connection is approximate rather than exact.
FAM broadens the target of explanation. In tasks such as person re-identification or self-supervised representation learning, a test image may not correspond to a training class. Explaining a class logit is then less meaningful. FAM defines Feature Activation Mapping and explains which image regions contribute to the feature vector itself. This is an important shift because many modern vision systems use encoders as representation extractors rather than closed-set classifiers. The limitation is that the target feature must be carefully defined, and the resulting explanation is about representation formation rather than a single class decision. Abs-CAM, SSG-CAM, and ShapleyCAM further illustrate how the gradient signal can be refined. Abs-CAM turns gradients into absolute positive contributions and fuses the resulting saliency with the input image, which reduces negative-gradient noise and improves deletion, insertion, and pointing-game behavior. SSG-CAM uses smoothed second-order gradients and combines them with differential evolution to select and fuse the best multi-layer feature maps. It directly addresses two weaknesses of first-order CAMs: saturation and manual layer selection. ShapleyCAM builds a content-reserved game-theoretic framework and derives Shapley-style CAM weights by using gradients and Hessian information. It bridges scalable CAMs and fair feature attribution, but it also introduces second-order computation and utility-function design choices. Together, these methods show that the gradient family is moving from simple sensitivity toward more principled attribution.
Transformer explanations require a different interpretation model from CNN explanations. In CNNs, spatial evidence is naturally organized in feature maps. In vision transformers, evidence flows through patch tokens, class tokens, residual connections, attention heads, and sometimes cross-modal attention. A naive attention map is therefore not automatically an explanation. Attention indicates how tokens exchange information, but it does not by itself show which tokens are relevant to the final decision. The two transformer papers in this group are important because they explicitly move beyond attention visualization and treat transformer explanation as a relevance-propagation problem. Chefer, Gur, and Wolf proposed a method that starts from the observation that common transformer visualizations either display attention weights from one layer or roll attention matrices across layers. Both approaches can be misleading. A token can receive high attention without being important for the final class, and simple attention rollout assumes a linear flow of information that ignores nonlinear transformations, residual paths, and the difference between attention and relevance. The proposed method assigns local relevance using the Deep Taylor Decomposition principle and then propagates relevance through the transformer layers while accounting for attention and skip connections. The aim is to preserve the total amount of relevance as it moves backward through the model. In vision tasks, the resulting relevance scores can be mapped back to image patches and visualized as a class-specific explanation map. The method was benchmarked on visual transformer networks and a text classification task and was shown to outperform attention-only baselines. Its main limitation is that it requires access to model internals and careful bookkeeping through transformer components. It is therefore more principled than attention rollout but also more architecture-aware. The same authors then generalized the approach to bi-modal and encoder-decoder transformers, covering self-attention, co-attention, and encoder-decoder attention. These papers are important because they show that transformer explanation is not simply CNN CAM with a different feature map. Relevance must move through token mixing, residual paths, and cross-modal interactions.
Score-CAM was a turning point because it removed the dependence on backpropagated gradients. It uses activation maps as masks, forwards masked images through the model, and uses the target-class confidence as the channel weight. This often gives clean maps and passes sanity checks, but it is expensive because a separate forward pass may be needed for many channels. Ablation-CAM answers a related question by measuring how much the target score changes when a feature map is removed. Its conceptual advantage is direct marginal contribution; its cost is repeated ablation. Ablation-CAM++ reduces this cost by grouping activation maps recursively, preserving the ablation idea while making it more time efficient. Eigen-CAM uses principal components of learned feature representations instead of class-specific gradients or scores. This makes it simple, class-independent, and robust when dense classifier layers fail, but it can lose explicit class discrimination. BBAM is different because it uses a trained object detector and finds the smallest image regions inside a bounding box that preserve detector behavior. It is useful for weakly supervised semantic and instance segmentation with box labels, but it is detector-specific rather than a generic classifier explanation. AD-CAM, Cluster-CAM, CAPE, and ReciproCAM all respond to the same pressure: how can a gradient-free method be clearer without becoming too slow? AD-CAM uses lightweight spatial feature masks. Cluster-CAM clusters feature maps to reduce the number of forward passes and combines cognition-base and cognition-scissors maps. CAPE reformulates CAM as a probabilistic ensemble so that regional contributions can be meaningfully compared across classes. ReciproCAM perturbs intermediate feature maps spatially and reports much faster runtime than Score-CAM while preserving strong ADCC performance. ScoreCAM++ improves Score-CAM itself. It argues that min-max normalization can hide the difference between high- and low-priority activation values. By gating activation maps with a stronger activation function, especially tanh, ScoreCAM++ separates important and unimportant regions more clearly. The result supports a broader lesson: even when the attribution family stays the same, normalization and gating can substantially affect faithfulness.
Hybrid CAM methods often use explanations as training signals. SEAM adds self-supervised equivariance: CAMs should transform consistently when the input is transformed. Its pixel correlation module further refines pixels by similar neighbors, reducing under-activation and over-activation. AdvCAM takes a different route by anti-adversarially manipulating the input so that regions initially ignored by the classifier become involved in later attributions. C2AM uses contrastive learning to generate class-agnostic activation maps from unlabeled data, separating foreground and background representations before using them to refine class-specific CAMs. Causal and bias-aware methods address a deeper problem: a CAM can be faithful to a bad reason. CI-CAM uses causal intervention to reduce object-context entanglement in WSOL, especially when objects co-occur with backgrounds such as birds and branches or ducks and water. C-CAM adapts this idea to medical WSSS, where organ co-occurrence and ambiguous boundaries make ordinary CAMs unreliable; it models category-causality and anatomy-causality chains and reports strong DSC improvements on ProMRI, ACDC, and CHAOS. Debiased-CAM studies image perturbation biases such as blur, color temperature, and day/night changes. It trains a multi-input, multi-task model with auxiliary explanation and bias-level prediction so that explanations remain closer to the unbiased scene. These methods are important because they shift the goal from making a sharper heatmap to making the right heatmap. Other hybrid methods focus on dense prediction. PCAA introduces Partial CAM for semantic segmentation and uses local and global class-level representations to model pixel-to-class relations. RPIM builds foreground regions from superpixels and initial responses, then uses intra-region integration and inter-region spreading to improve consistency. OLM refines low-level feature activation maps online to produce compact and threshold-robust localization maps. These papers show that CAMs are not merely explanatory artifacts; in WSSS and WSOL they become training data. CALM takes a different route from post-hoc CAMs by embedding the attribution mechanism into the training process. Instead of generating a heatmap only after classification, it introduces a latent variable that represents the spatial location of the recognition cue and optimizes it with an expectation-maximization procedure. This makes the attribution map part of the model's learning objective, which helps the classifier learn more localized and discriminative visual evidence. Its limitation is that it is not a purely post-hoc method and requires a modified training procedure.
High-resolution CAM methods address the spatial coarseness of final-layer activation maps. Augmented Grad-CAM generates multiple low-resolution Grad-CAM maps from transformed copies of the same image and combines them into a super-resolution-style heatmap. Augmented Score-CAM applies a similar augmentation idea to Score-CAM and reports better human preference, lower Average Drop, higher Average Increase, and better IoU than the original Score-CAM on ImageNet-based evaluations. F-CAM attaches a trainable decoder to a classifier and uses foreground/background samples plus image priors to produce full-resolution CAMs. This improves boundaries, but it also makes the method less purely post-hoc because the decoder must be fine-tuned. LayerCAM is one of the most influential high-resolution variants. It uses positive gradients at individual spatial locations and can generate reliable CAMs from different CNN layers. Shallow layers provide fine detail, while deeper layers provide semantic class evidence. By fusing these maps, LayerCAM improves WSSS and WSOL quality compared with single-layer CAMs. OLM also uses low-level feature information, but it learns an online activation-map generator and evaluator. The goal is to produce a compact and threshold-robust map rather than a sparse discriminative patch. Poly-CAM later combines earlier and later layers to produce higher-resolution maps that remain competitive on insertion-deletion faithfulness. RPIM introduces a region-based refinement mechanism. It first builds foreground regions from superpixels and initial CAM responses, then uses intra-region integration and inter-region spreading to make pixels inside a region and across related regions share more consistent activation. This is useful for WSSS because pixel-level pseudo-labels need coherent object regions rather than isolated peaks. GFR-CAM uses Gram-Schmidt orthogonalization to generate a hierarchy of orthogonal explanation components. Unlike PCA-based approaches that focus on one dominant explanation, GFR-CAM can reveal secondary objects or semantic parts, reducing explanatory tunnel vision. Finer-CAM makes a different but very important observation: CAM often fails in fine-grained recognition because it explains what supports the target class, including features shared with similar classes. Finer-CAM instead compares the target class with similar reference classes and explains the logit difference. The practical message is simple: for fine-grained explanations, the question should often be not 'why this class?' but 'why this class rather than the closest alternative?' The remaining challenge is reference selection: the method is most effective when the reference class is truly visually competitive.
TS-CAM uses the self-attention mechanism of visual transformers to overcome the local-receptive-field limitation of CNN CAMs. It couples token semantics with attention maps and reports a large Top-1 localization gain on CUB-200-2011. MCTformer extends this idea to WSSS by using multiple class tokens, so class-to-patch attention can generate class-specific localization maps. CTI infuses class tokens within and across images and adds a background token to reduce false positives. Prompt-CAM learns class-specific prompts for a pretrained ViT, making fine-grained traits visible through prompt-query attention. TransCAM extends CNN-based CAMs to hybrid CNN-Transformer models by aligning convolutional feature maps with transformer patch tokens. It integrates multi-level feature maps with class-token attention to produce high-resolution token-aware CAMs suitable for WSOL and fine-grained object localization. TransCAM demonstrates improved localization performance on CUB and Cars datasets while maintaining compatibility with pre-trained transformers. Foundation-model methods broaden the target space. gScoreCAM explains CLIP by using gradients only to select the top channels and Score-CAM-style forward scores to weight them, reducing Score-CAM's CLIP runtime by about eight times while retaining strong object localization. S2C transfers SAM knowledge to the classifier during training through SAM-segment contrasting and CAM-based prompting. The DINO semantic guider builds a class-aware affinity region map by propagating CAM seeds over DINO self-attention affinity graphs. DiffCAM explains a target example by comparing its feature distribution with reference examples rather than relying on decision-boundary gradients. MetaCAM combines multiple CAM methods through top-k consensus and adaptive thresholding, showing that ensemble agreement can outperform individual maps.
The cross-family comparison leads to three practical conclusions. First, Grad-CAM-style methods remain the fastest and most convenient first-line diagnostic tools. Second, gradient-free and additive-attribution methods often improve faithfulness or localization, but their computational cost or design complexity can be higher. Third, high-resolution and foundation-model methods should be evaluated with special care because their maps may reflect external priors, prompts, or learned reference distributions rather than only the original classifier.
Model-based and architecture-aware methods do not treat CAM as a purely post-hoc heatmap. They modify training, pooling, erasing, guidance, attention, or class-specific prototypes so that the model itself becomes more localizable. The original CAM method is architecture-aware because it depends on global average pooling followed by a linear classifier. This design preserves the spatial meaning of the last convolutional maps and makes it possible to project class weights back to image regions. Hide-and-Seek modifies the training image by randomly hiding patches. The model is forced to use less discriminative object parts when the most discriminative part is hidden. ACoL uses two classifiers and adversarial erasing inside the feature map: one branch finds discriminative regions, and the other branch learns complementary regions after those regions are suppressed. SPG turns high-confidence attention regions into self-produced foreground/background guidance masks, gradually adding spatial supervision during classification training. ADL uses an attention-based dropout layer that sometimes erases the most discriminative region and sometimes highlights informative regions, balancing localization completeness and recognition accuracy. Rethinking CAM analyzes the standard WSOL pipeline and identifies three underappreciated causes of poor localization: global average pooling can overemphasize channels with small activation areas, negative weights can suppress true object regions, and max-based thresholding can be unstable. It proposes thresholded average pooling, negative weight clamping, and percentile thresholding. CREAM addresses incomplete localization by modeling foreground and background activation distributions. It learns class-specific context embeddings and then performs class re-activation through a Gaussian-mixture-style estimation. LPCAM attacks the discriminative-feature problem by extracting local prototypes from non-discriminative object parts. Instead of using only classifier weights, it clusters unpooled local features for an object class and matches new images to those prototypes, covering object parts such as head, body, and legs.
The first lesson from the corpus is that CAM-style methods now answer several different questions. Grad-CAM asks which feature maps locally support a score. Score-CAM asks which activation masks preserve the score under forward evaluation. Finer-CAM asks which regions distinguish one class from its nearest alternatives. FAM asks what part of an image contributes to a representation rather than to a class label. SAM- and DINO-assisted methods ask how external pretrained structure can improve weakly supervised CAM seeds. These are related but not identical explanation problems. The second lesson is that faithfulness and localization should not be conflated. A map can overlap an object and still be unfaithful to the model if the model actually relied on background context. Conversely, a faithful map may reveal an unwanted shortcut. This is why causal and debiasing methods are important. CI-CAM, C-CAM, and Debiased-CAM show that the quality of an explanation depends not only on the heatmap formula but also on the causal structure and bias conditions of the data. The third lesson is that resolution must be controlled semantically. LayerCAM, F-CAM, Poly-CAM, RPIM, and GFR-CAM improve spatial detail. However, a high-resolution map is not automatically a better explanation. If the target score is broad or confounded, higher resolution may simply reveal the wrong evidence in greater detail. Finer-CAM is valuable because it changes the target to a contrastive question, making high-resolution detail more meaningful in fine-grained recognition. The fourth lesson is that foundation models complicate explanation ownership. If SAM improves CAMs, the final map partly reflects SAM's segmentation prior. If DINO affinities spread CAM seeds, the final pseudo-label partly reflects DINO's self-supervised representation. If CLIP is explained with a prompt, the heatmap depends on the wording and embedding of that prompt. Future CAM papers should therefore report not only the heatmap method but also the source of all external priors and prompts.
A first challenge is standardized evaluation. Many CAM papers use different target layers, input sizes, thresholds, perturbation baselines, and post-processing. Future papers should report a minimal evaluation card: model and layer, target definition, normalization, thresholding, number of forward and backward passes, whether CRF/SAM/saliency priors were used, and which metrics were computed. This would make cross-paper comparison more reliable. A second challenge is causal faithfulness. Confounding is common in natural images, and medical images often contain anatomical co-occurrence. CI-CAM and C-CAM provide early examples of causal reasoning in CAM-style methods. Future work should combine interventional datasets, counterfactual image editing, and causal metrics to test whether highlighted evidence is necessary for the model decision or merely correlated with it. A third challenge is efficient gradient-free explanation. Score-CAM and Ablation-CAM are attractive because they avoid unstable gradients, but their cost limits deployment. ReciproCAM, Cluster-CAM, AD-CAM, and Ablation-CAM++ show that structured perturbation, grouping, and feature-level masking can reduce this cost. A promising direction is adaptive explanation: use fast gradients when they are reliable, and switch to perturbation only when uncertainty or inconsistency is high. A fourth challenge is prompt and reference sensitivity in foundation models. Finer-CAM depends on the reference class, gScoreCAM depends on prompt-conditioned CLIP scores, and DiffCAM depends on reference data distributions. Future evaluations should include prompt sensitivity, reference-class sensitivity, and reference-set sensitivity as standard robustness tests. A fifth challenge is human validation. Human studies in Grad-CAM and Debiased-CAM show that explanations can help users calibrate trust and improve task performance. Yet human preference should not be the only criterion because visually pleasing maps can still be unfaithful. Future work should combine human-centered evaluation with deletion/insertion, sanity checks, causal interventions, and task outcomes.
CAM-style explanation has evolved from a simple CNN localization mechanism into a broad methodological family for explainable computer vision. The original CAM formulation showed that a classifier can reveal discriminative regions through global average pooling. Grad-CAM made this idea broadly applicable. Later work improved derivatives, avoided gradients, refined resolution, used causal and debiasing constraints, introduced token-level explanations, and incorporated foundation-model priors. The main conclusion of this review is that CAMs should not be treated as interchangeable heatmaps. Each method answers a specific explanatory question with a specific technical mechanism. A fast-debugging tool, a faithful post-hoc explanation, a dense pseudo-label generator, a fine-grained trait localizer, and a CLIP prompt explanation are related but different objects. A rigorous CAM study should therefore report the target being explained, the evidence source, the attribution mechanism, the evaluation protocol, the computational cost, and the limitations of the explanation. This method-centered view is essential for building a more reliable and useful next generation of visual explanations.
Improvements for AI systems
Based on this paper, here are specific improvements for AI systems and what the improved systems can do:
1. Implement adaptive attribution selection. Build an AI system that automatically chooses between gradient-based (fast, coarse) and gradient-free (slow, faithful) CAM methods based on task requirements and model state. The system can detect gradient saturation or noise and switch to perturbation-based methods only when needed, reducing computational cost by up to 8x while maintaining faithfulness.
2. Add contrastive explanation targets. Improve fine-grained recognition systems by explaining why class A instead of class B
rather than just why class A.
The improved system can automatically select the most visually similar reference class and generate contrastive heatmaps that highlight only the discriminative features, reducing confusion between similar categories (e.g., bird species, car models).
3. Integrate causal debiasing into explanation generation. Build a system that identifies and removes spurious correlations (e.g., background context, co-occurring objects) from explanations. The system can detect when a model relies on shortcuts (birds with branches, medical organs with anatomical neighbors) and generate counterfactual explanations that show what evidence would be necessary and sufficient for the decision, not merely correlated.
4. Create prompt-robust foundation-model explanations. For CLIP-like systems, implement automatic prompt sensitivity testing that generates explanations across multiple prompt phrasings and reports variance. The improved system can identify when explanations are prompt-dependent and provide a consensus map across prompts, making open-vocabulary explanations more reliable.
5. Develop multi-layer, token-aware explanation fusion. Build a system that combines explanations from shallow layers (fine detail) and deep layers (semantic evidence) using transformer token alignment. This produces high-resolution maps that preserve both object boundaries and class-level semantics, improving weakly supervised segmentation and localization tasks by 10-20% mIoU.
6. Implement self-supervised equivariance constraints. Train explanation generators that produce consistent heatmaps under image transformations (rotation, scaling, color shifts). The system can use these constraints as a regularizer during training, improving both localization accuracy and robustness to input perturbations.
7. Add reference-distribution comparison for representation models. For self-supervised encoders (DINO, MAE), build a system that explains feature vectors by comparing the input's feature distribution against reference examples. This enables explanation of images that don't correspond to training classes, useful for anomaly detection and open-set recognition.
8. Create standardized evaluation cards. Implement an automated evaluation protocol that reports: target layer, normalization, thresholding, forward/backward pass counts, external priors used (SAM, CRF), and multiple metrics (deletion, insertion, localization, robustness). The system can automatically generate comparable evaluation reports across different CAM methods, enabling fair benchmarking.
9. Build ensemble consensus explanations. Combine multiple CAM methods (gradient, gradient-free, high-resolution, token-based) using top-k consensus and adaptive thresholding. The improved system can detect when methods disagree and flag uncertain regions, providing more reliable explanations for high-stakes applications.
10. Implement training-time attribution embedding. Modify classifier training to include an expectation-maximization latent variable for spatial cue location, as in CALM. The improved system learns more localized and discriminative features during training, producing inherently more explainable models without requiring post-hoc analysis.
Sources
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models