Faithful Faithfulness Evaluations: Challenges & Pitfalls Learned from a Breast MRI Case Study

arXiv:2609.25978 · cs.CV, cs.HC, cs.LG, eess.IV · Submitted 2026-09-22 · Read on arXiv

cs.CV, cs.HC, cs.LG, eess.IV

Submitted: 2026-09-22

Updated: 2026-09-22

Comments: Accepted at MICCAI iMIMIC Workshop 2026

Code: https://github.com/friendorus/Faithful-Faithfulness-Evaluations

License: http://creativecommons.org/licenses/by/4.0/

The gist: Saliency maps are widely used to explain deep learning predictions in medical imaging, yet visually plausible explanations do not necessarily reflect a model's true decision process and may therefore

Terminology

Abstract

Saliency maps are widely used to explain deep learning predictions in medical imaging, yet visually plausible explanations do not necessarily reflect a model's true decision process and may therefore mislead clinicians. We investigate this problem using a Vision Transformer-based breast MRI classifier trained on the ODELIA Breast MRI Challenge dataset and evaluate multiple saliency methods, including Last-layer Attention, Attention Rollout, Grad-SAM, Gradient Attention Rollout, GMAR, Grad-CAM, and HiResCAM. Our study highlights two often-overlooked challenges in perturbation-based faithfulness evaluation. First, method rankings depend strongly on the perturbation strategy, varying across intensity-based perturbations and transformer-based attention masking. Second, benchmarking saliency methods requires distinguishing between class-specific and class-agnostic explanations. To enable fair comparisons, we introduce non-class-specific variants of gradient-based methods and evaluate both settings separately. Across protocols, Grad-CAM and Gradient Attention Rollout consistently emerged as the strongest class-specific methods, although their relative ranking depended on the evaluation design. These findings expose important limitations of current saliency-based explainability approaches and highlight the need for more robust and standardized evaluation frameworks for trustworthy clinical AI systems.

Related papers