A Comparison of Malware Image Transformations Using Grad-CAM and Hybrid Learning Models

arXiv:2608.12077 · cs.CR · Submitted 2026-08-12 · Read on arXiv

Vibha Bhavikatti, Mark Stamp

San Jose State University

cs.CR

Submitted: 2026-08-12

Updated: 2026-08-13

Comments: To appear as a chapter in the book "Artificial Intelligence for Cyber Defense in Emerging Threats", to be published by Springer by early 2027

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: This paper investigates the use of Gradient-weighted Class Activation Maps (Grad-CAM) as an eXplainable AI (XAI) tool for analyzing eight distinct image types derived from malware samples.

Terminology

Summary

This paper investigates the use of Gradient-weighted Class Activation Maps (Grad-CAM) as an eXplainable AI (XAI) tool for analyzing eight distinct image types derived from malware samples. The research builds on prior work by Agrawal et al. [2], which evaluated eight exe-to-image transformation techniques on a 17-family malware dataset derived from the RawMalTF dataset, with 1,000 samples per family (17,000 total). The eight transformations are Grayscale, Entropy Hilbert, Byteclass Hilbert, HIT, Bigram Cartesian, Bigram Polar, Spiral, and Byteclass.

The paper's key contributions include generating Grad-CAM heatmaps for each malware sample across all eight transformations, treating these heatmaps as additional malware images for feature extraction, comparing CNN-based and feature-based approaches, generating HiResCAM heatmaps for comparison, quantitatively evaluating Grad-CAM explanations using faithfulness and stability metrics, comparing classification performance between models trained on original images versus Grad-CAM overlay images, and training various models on feature combinations derived from MobileNetV2 embeddings.

For the CNN experiments, MobileNetV2 is used as the backbone with a classification head consisting of dropout (rate 0.3), a dense layer of size 256 with ReLU activation, and a softmax output layer. Training is performed in two stages: first with the backbone frozen, then fine-tuning the upper 30% of layers. The dataset is split 70:15:15 for training, validation, and testing using stratified splits.

The CNN test accuracy results on the eight image transformations are: Grayscale 0.592, Entropy Hilbert 0.688, Byteclass Hilbert 0.433, HIT 0.532, Bigram Cartesian 0.601, Bigram Polar 0.609, Spiral 0.489, and Byteclass 0.536. The entropy-based Hilbert curve images yield the strongest results, while Byteclass Hilbert and Spiral give the weakest.

For classical machine learning models trained on Grad-CAM feature vectors (50 dimensions consisting of global intensity statistics, shape features, LBP statistics, and GLCM statistics), the results are: Linear SVC 0.075, SVM (RBF) 0.216, XGBoost 0.275, and Random Forest 0.294. These all exceed random accuracy of 0.059, indicating the heatmaps capture meaningful family-level information.

CNNs trained directly on Grad-CAM overlay images achieve: Grayscale 0.566, Entropy Hilbert 0.604, Byteclass Hilbert 0.433, HIT 0.532, Bigram Cartesian 0.555, Bigram Polar 0.553, Spiral 0.489, and Byteclass 0.536. Entropy Hilbert and Grayscale show the best performance.

A hybrid CNN-HOG-XGBoost pipeline, which concatenates 256-dimensional CNN embeddings with HOG features and trains an XGBoost classifier, achieves: Grayscale 0.717, Entropy Hilbert 0.711, Byteclass Hilbert 0.718, HIT 0.731, Bigram Cartesian 0.658, Bigram Polar 0.654, Spiral 0.670, and Byteclass 0.721. This hybrid approach consistently outperforms CNN-only models by double-digit percentages.

The faithfulness metrics (deletion AUC, insertion AUC, faithfulness score, concentration p80, drop top20pct) are evaluated on 200 samples per transformation. Bigram Polar produces the most faithful Grad-CAM explanations (faithfulness 0.229) despite being a mid-tier performer on CNN accuracy. Entropy Hilbert ranks second on faithfulness (0.121) while having the highest CNN accuracy, making it the most balanced transformation overall.

Stability metrics (Spearman correlation, MSE, SSIM, top-K overlap) are evaluated on 50 samples per transformation with Gaussian noise (σ=5, n=5 draws). Byteclass and Grayscale produce the most stable CAMs under noise, while Bigram Polar, which leads on faithfulness, ranks seventh on stability. Spiral is consistently the weakest transformation across both faithfulness and stability.

A key finding is that no single transformation dominates both faithfulness and stability; transformations that produce highly faithful explanations tend to have lower stability, and vice versa. This reflects a fundamental tension between explanation correctness and explanation robustness.

Comparing CNN accuracy on original images versus Grad-CAM overlays, for five out of eight transformations, overlay images perform equal to or better than original images. The largest improvements are for Grayscale (+6.0%) and Byteclass Hilbert (+5.3%). Bigram-based transformations are the only cases where original images outperform overlays.

Progressive feature fusion experiments show that Random Forest consistently outperforms SVM and XGBoost across all transformations. The 512-dimensional combined features (original + overlay embeddings) improve over the 256-dimensional overlay-only features in every transformation. The final experiment combines all 16 model embeddings (eight transformations × original and overlay) into a 4096-dimensional feature vector, achieving a test accuracy of 0.777 with a Random Forest classifier. This exceeds the previous benchmark of 0.751 from prior work [2] on the same dataset.

Per-family analysis reveals that Makoob is the most reliably classified family (0.993 accuracy), while Noon (0.560) and Androm (0.573) remain the most difficult. Noon, Androm, Agensla, and Injuke appear in the bottom five families by accuracy across virtually every experiment.

The comparison between Grad-CAM and HiResCAM shows that the two methods produce effectively equivalent explanations at the final convolutional layer of MobileNetV2, with mean pixel-level difference of 4×10−8 and maximum difference of 3×10−7. This is attributed to the small spatial dimensions of MobileNetV2's final feature map (7×7).

The paper concludes that transformation choice strongly influences classification performance, Grad-CAM overlays preserve meaningful family-discriminative structure, and combining complementary deep feature representations across multiple transformation types is a useful strategy for malware classification. The authors note limitations including the exclusive use of MobileNetV2 as a backbone and the computational cost of generating full Grad-CAM datasets. Future work directions include adding HOG features to the 512-dimensional and 4096-dimensional feature vectors, applying Grad-CAM and HiResCAM at intermediate layers, exploring other XAI techniques, and studying adversarial robustness of XAI explanations.

Improvements for AI systems

Improvements to AI Systems:

  1. Multi-Transformation Ensemble Classifier: Build an AI system that simultaneously processes malware samples through all eight image transformations (Grayscale, Entropy Hilbert, Byteclass Hilbert, HIT, Bigram Cartesian, Bigram Polar, Spiral, Byteclass), extracts deep embeddings from each, and fuses them into a 4096-dimensional feature vector for classification. This system achieves 77.7% test accuracy on 17-family malware classification, outperforming single-transformation CNNs (best: 68.8%) and the prior benchmark (75.1%).

  2. Hybrid CNN-HOG-XGBoost Pipeline: Implement a classifier that concatenates 256-dimensional CNN embeddings (from MobileNetV2) with HOG features and trains an XGBoost model. This improves accuracy by 10–20 percentage points over CNN-only models across all transformations (e.g., Grayscale: 59.2% → 71.7%, HIT: 53.2% → 73.1%). The system can be deployed for high-precision malware family identification where single-model performance is insufficient.

  3. Faithfulness-Stability Balanced XAI Selector: Create an AI module that automatically selects the optimal image transformation for explainability based on a trade-off score combining faithfulness (deletion AUC, insertion AUC) and stability (Spearman correlation, SSIM). For example, Entropy Hilbert offers the best balance (high CNN accuracy 68.8% and second-highest faithfulness 0.121), while Bigram Polar is chosen when faithfulness is prioritized (0.229) despite lower stability. This system can dynamically choose the most trustworthy explanation method per deployment context.

  4. Grad-CAM Overlay Augmentation for Weak Transformations: For transformations with low baseline accuracy (e.g., Byteclass Hilbert at 43.3%, Spiral at 48.9%), train CNNs on Grad-CAM overlay images as an augmentation strategy. This yields improvements of up to +5.3% (Byteclass Hilbert) and +6.0% (Grayscale) over original-image training, enabling better classification for underperforming representations without new data collection.

  5. Progressive Feature Fusion with Random Forest: Implement a feature-fusion AI that progressively combines original-image embeddings, Grad-CAM overlay embeddings, and HOG features, using Random Forest as the final classifier (which outperforms SVM and XGBoost in all tested configurations). The system can scale from 256-dimensional (overlay-only) to 512-dimensional (original+overlay) to 4096-dimensional (all transformations) feature vectors, with accuracy improving at each stage—useful for resource-constrained vs. high-accuracy deployment scenarios.

  6. Per-Family Difficulty-Aware Rejection System: Use the per-family accuracy analysis to build a classifier that flags predictions for known difficult families (Noon, Androm, Agensla, Injuke) as low-confidence, triggering manual review or additional verification. This improves overall system reliability by reducing silent misclassifications on the hardest 4 of 17 families, where accuracy drops to 56–57% compared to 99.3% for the easiest family (Makoob).

  7. XAI-Integrated Malware Triage System: Combine the Grad-CAM heatmaps with the hybrid classifier to provide both a classification decision and a visual explanation (heatmap) for security analysts. The system can highlight which regions of the transformed image drove the decision, with the knowledge that Entropy Hilbert and Grayscale overlays are most stable under noise (for reliable explanations) and Bigram Polar is most faithful (for accurate attribution). This enables faster human verification of AI decisions in security operations centers.

  8. Noise-Robust Malware Detector: Leverage the stability findings (Byteclass and Grayscale produce the most stable CAMs under Gaussian noise) to design a malware classifier that remains reliable under input perturbations. This is critical for adversarial settings where attackers may add noise to evade detection; the system can prioritize these transformations for deployment in hostile environments.

Abstract

Recent studies have shown that binary-to-image representations can enable effective machine learning-based results for malware detection and classification. However, performance can vary significantly, depending on the technique used to convert binaries to images. Furthermore, the explainability and interpretability of image-based models is largely unexplored within the malware domain. In this research, we employ Gradient-weighted Class Activation Maps (Grad-CAM) as an eXplainable AI (XAI) tool, which we use to analyze eight distinct image types derived from malware samples. We provide quantitative faithfulness and stability metrics for Grad-CAM heatmaps and we compare these heatmaps to High-Resolution Class Activation Mappings (HiResCAM). We also show that Grad-CAM heatmaps can provide useful information for malware classification. Specifically, we show that a Random Forest model trained on features extracted from Grad-CAM images via a MobileNetV2 Convolutional Neural Network (CNN) model achieves a test accuracy of 0.777 across 17 malware families, exceeding a previous benchmark of 0.750 for this same dataset. A key finding of this research is that for the malware image transformations considered, accuracy and explanation faithfulness do not coincide, e.g., image transformation techniques that produce the most faithful explanations yield only mid-tier accuracy.

Sources

Related papers