Null importance: Disentangling relevance for interpretable machine learning
stat.ML, cs.LG
Submitted: 2026-09-16
Updated: 2026-09-16
Comments: 29 pages, 9 files. Submitted to Statistical Science
Code: https://github.com/krisrs1128/null_importance
License: http://creativecommons.org/licenses/by/4.0/
The gist: Feature importance is central to interpretable machine learning, but the term "importance" encompasses several fundamentally different notions of relevance.
Terminology
Abstract
Feature importance is central to interpretable machine learning, but the term "importance" encompasses several fundamentally different notions of relevance. We develop a unified perspective based on null importance: a population-level characterization of when a feature is irrelevant under a specified notion of relevance. We consider standard notions of null importance arising from marginal and conditional statistical relevance, predictive risk, functional invariance, and causal effects, and show how these notions answer different scientific questions. We illustrate the framework in two applications in which the distinction is particularly consequential: algorithmic fairness, where common fairness criteria correspond to different notions of null importance, and genomic perturbation modeling, where different notions of relevance lead to different conclusions about what a prediction model has learned. The framework connects three aspects of feature analysis: the scientific question defining relevance, the data and model assumptions that shape how different null notions relate, and the methods used to assess importance. We establish sufficient conditions under which null notions coincide and give counterexamples showing how they diverge when those conditions fail. We then characterize which nulls different method families target and when their zero-importance statistics identify those targets. Finally, simulations spanning feature dependence, redundancy, nonlinearity, hidden features and other standard phenomena, along with case studies on image and multiomics data, provide empirical evidence for these theoretical distinctions and their practical consequences. Taken together, these results provide a common statistical language for relating scientific questions, data-generating assumptions, and algorithms, and clarify the conclusions that feature-importance analyses can support.
Sources
- Sparse Autoencoders Find Highly Interpretable Features in Language Models
- Proxy Non-Discrimination in Data-Driven Systems
- Towards A Rigorous Science of Interpretable Machine Learning
- Inherent Trade-Offs in the Fair Determination of Risk Scores
- CausalGAN: Learning Causal Implicit Generative Models with Adversarial Training
- Effects of Distance Metrics and Scaling on the Perturbation Discrimination Score
- Consistent Individualized Feature Attribution for Tree Ensembles
- SmoothGrad: removing noise by adding noise
Related papers
- Behavior of prediction performance metrics with rare events
- Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs
- A Posterior-Dynamics Framework for Imaging Inverse Problems with Pretrained Diffusion Priors
- One Permutation Is All You Need: Fast, Deterministic Feature Importance and Model Stress-Testing
- Online Conformal Prediction for Non-Exchangeable Panel Data
- Deep Time-Series Forecasting in 10 Years: A Survey