CID: Measuring Feature Importance Through Counterfactual Distributions
cs.LG
Submitted: 2025-11-19
Updated: 2026-09-22
Comments: Accepted at Northern Lights Deep Learning (NLDL) 2026 Conference
Code: https://github.com/EddieConti/CID
License: http://creativecommons.org/licenses/by/4.0/
The gist: Assessing the importance of individual features in Machine Learning is critical to understand the model's decision-making process.
Terminology
Abstract
Assessing the importance of individual features in Machine Learning is critical to understand the model's decision-making process. While numerous methods exist, the lack of a definitive ground truth for comparison highlights the need for alternative, well-founded measures. This paper introduces a novel post-hoc local feature importance method called Counterfactual Importance Distribution (CID). We generate two sets of positive and negative counterfactuals, model their distributions using Kernel Density Estimation, and rank features based on a distributional dissimilarity measure. This measure, grounded in a rigorous mathematical framework, satisfies key properties required to function as a valid metric. We showcase the effectiveness of our method by comparing with well-established local feature importance explainers. Our method not only offers complementary perspectives to existing approaches, but also improves performance on faithfulness metrics (both for comprehensiveness and sufficiency), resulting in more faithful explanations of the system. These results highlight its potential as a valuable tool for model analysis. Link to repository: https://github.com/EddieConti/CID
Sources
- A Survey on the Robustness of Feature Importance and Counterfactual Explanations
- Explainable AI needs formalization
- Inherent Inconsistencies of Feature Importance
- Calculating and Visualizing Counterfactual Feature Importance Values
- On the Robustness of Interpretability Methods
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks