Are explainable AI (XAI) evaluation strategies aligned? Comparing subjective, objective, and mathematical evaluation measures using saliency maps
cs.HC, cs.AI
Submitted: 2025-04-23
Updated: 2026-09-14
Comments: 31 pages, 8 figures, 6 tables
Journal ref: Kares, Felix, et al. "Are explainable AI (XAI) evaluation strategies aligned? Comparing subjective, objective, and mathematical evaluation measures using saliency maps." Computers in Human Behavior (2026): 109163
Code: https://github.com/jacobgil/pytorch-grad-cam
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: The evaluation of explainable AI (XAI) approaches often relies on three families of methods: subjective measures (e.g., questionnaires on trust or satisfaction), objective measures (e.g., task
Terminology
Abstract
The evaluation of explainable AI (XAI) approaches often relies on three families of methods: subjective measures (e.g., questionnaires on trust or satisfaction), objective measures (e.g., task performance metrics), and mathematical metrics (e.g., for faithfulness). Yet, it remains unclear how these families align or diverge in practice. In a preregistered between-subjects study (N=166), we use three established saliency map techniques (LIME, Grad-CAM, Guided Backpropagation) as a testbed to examine this issue. We find that each family of methods leads to different conclusions: participants reported no differences in trust or satisfaction, Grad-CAM improved user performance, while mathematical metrics favored Guided Backpropagation. At the same time, mathematical metrics were only partially related to user performance, and these relationships were sometimes counterintuitive. Our findings highlight the methodological importance of comparing subjective, objective, and mathematical approaches when evaluating XAI, illustrating both tensions and aspects that are aligned. We discuss implications for XAI evaluation frameworks.
Sources
- One Explanation Does Not Fit All: A Toolkit and Taxonomy of AI Explainability Techniques
- What Makes a Good Explanation?: A Harmonized View of Properties of Explanations
- Towards A Rigorous Science of Interpretable Machine Learning
- Explainable AI for Natural Adversarial Images
- Metrics for Explainable AI: Challenges and Prospects
- Responsibility: An Example-based Explainable AI approach via Training Process Inspection
- On quantitative aspects of model interpretability
- Striving for Simplicity: The All Convolutional Net
Related papers
- EduGage: A Multimodal Dataset and Benchmark for Sensor-Based Momentary Assessment of Engagement in Self-Guided Video Learning
- EvoDesign: Agentic Editable Diagram Creation via Design Expertise Evolution
- HAGI++: Head-Assisted Gaze Imputation and Generation
- Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving
- Review of Explainable Decision Support and Adaptive Human-Machine Interfaces for Automation Transparency in Maritime Autonomous Surface Ships
- Towards Cognitive Process-Aware Proactive Writing Support