Instance-wise Linearization of Neural Network for Model Interpretation
cs.LG, cs.AI, cs.CV
Submitted: 2023-10-25
Updated: 2026-09-05
License: http://creativecommons.org/licenses/by/4.0/
The gist: Neural network have achieved remarkable successes in many scientific fields.
Terminology
Abstract
Neural network have achieved remarkable successes in many scientific fields. However, the interpretability of the neural network model is still a major bottlenecks to deploy such technique into our daily life. The challenge can dive into the non-linear behavior of the neural network, which rises a critical question that how a model use input feature to make a decision. The classical approach to address this challenge is feature attribution, which assigns an important score to each input feature and reveal its importance of current prediction. However, current feature attribution approaches often indicate the importance of each input feature without detail of how they are actually processed by a model internally. These attribution approaches often raise a concern that whether they highlight correct features for a model prediction. For a neural network model, the non-linear behavior is often caused by non-linear activation units of a model. However, the computation behavior of a prediction from a neural network model is locally linear, because one prediction has only one activation pattern. Base on the observation, we propose an instance-wise linearization approach to reformulates the forward computation process of a neural network prediction. This approach reformulates different layers of convolution neural networks into linear matrix multiplication. Aggregating all layers' computation, a prediction complex convolution neural network operations can be described as a linear matrix multiplication F(x) = W times x + b. This equation can not only provides a feature attribution map that highlights the important of the input features but also tells how each input feature contributes to a prediction exactly. Furthermore, we discuss the application of this technique in both supervise classification and unsupervised neural network learning parametric t-SNE dimension reduction.
Sources
- Towards Robust, Locally Linear Deep Networks
- Sensitivity and Generalization in Neural Networks: an Empirical Study
- RISE: Randomized Input Sampling for Explanation of Black-box Models
- Dissecting Deep Neural Networks
- Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- SmoothGrad: removing noise by adding noise
- Object Detectors Emerge in Deep Scene CNNs
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks