MorphoSHAP: Rethinking the Unit of Attribution in Explanation for Deep Visual Models
cs.CV, cs.AI
Submitted: 2026-09-22
Updated: 2026-09-22
Comments: 21 pages
License: http://creativecommons.org/licenses/by/4.0/
The gist: Visual attribution methods typically explain predictions using pixels, superpixels, or regular patches.
Terminology
Abstract
Visual attribution methods typically explain predictions using pixels, superpixels, or regular patches. These representations can localize important regions, but provide limited information about their structure. We introduce MorphoSHAP, a model-agnostic post-hoc method that instead uses morphological shapes as the players of a Shapley attribution game. Using the Tree of Shapes, each shape is described by its scale, geometry, and signed contribution, providing explanations of where the evidence lies, what type of structure carries it, and how strongly it affects the prediction. This shared morphological vocabulary enables spatial, textual, and global class-level explanations beyond image-specific heatmaps. To the best of our knowledge, MorphoSHAP is the first SHAP-based image attribution framework to combine these different forms of explanation. Across five diverse datasets and three architectures, MorphoSHAP achieves strong insertion/deletion performance and outperforms competing attribution methods on several benchmarks. Finally, a user study shows that MorphoSHAP provides explanations that are easy to use and are preferred over standard attribution baselines.
Sources
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- A Proof of the Tree of Shapes in n-D
- Enhancing Concept Localization in CLIP-based Concept Bottleneck Models
- CAMs as Shapley Value-based Explainers
- Morphological Network: How Far Can We Go with Morphological Neurons?
- Distributionally Robust Neural Networks for Group Shifts: On the Importance of Regularization for Worst-Case Generalization
- Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models