CAM-Guided Saliency Cutout and Image-Based Malware Classification
Yasaman Ebrahimi, Martin Jurecek, Mark Stamp
San Jose State University · Czech Technical University in Prague
cs.CV, cs.LG
Submitted: 2026-08-12
Updated: 2026-08-13
Comments: To appear as a chapter in the book "Artificial Intelligence for Cyber Defense in Emerging Threats", to be published by Springer by early 2027
Code: https://github.com/yasamanebrahimi-byte/CAMRegularization
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: This study evaluated saliency-guided cutout regularization for image-based malware classification.
Terminology
Summary
This study evaluated saliency-guided cutout regularization for image-based malware classification. The controlled experiments compared no cutout, standard random cutout, low-saliency cutout, and high-saliency cutout using ResNet18. For each cutout method, the training set contains the original images plus M augmented cutout copies, with M = 4 and M = 8 tested. The main malware experiment used grayscale RawMal-TF images and every condition was evaluated with seeds 42, 43, and 44 and cutout areas of 5%, 10%, 20%, and 30%, with 100 training epochs per run.
The RawMal-TF results show that saliency-guided cutout did not improve malware classification under this setup. The no-cutout baseline attained 72.83%±0.16% mean best validation accuracy, while the best cutout condition, random cutout with M = 4 and 30% area, reached 71.55% ± 0.45%. Low-saliency cutout was sometimes slightly better and sometimes slightly worse than matched random cutout, and no low-saliency condition exceeded the no-cutout baseline. High-saliency cutout was generally harmful.
CIFAR-100 gave a different result. Low-saliency cutout with M = 4 and 10% area reached 63.51% ± 0.36%, compared with 62.65% ± 0.57% for no cutout. Low-saliency placement also substantially outperformed random placement for large masks, while high-saliency cutout was consistently harmful. This contrast shows that the implementation can produce a positive saliency-guided result, but that the benefit does not transfer to the grayscale malware-image setting.
Overall, our results suggest that saliency-guided cutout is domain dependent. The method may be useful for natural images, but it does not currently improve grayscale RawMal-TF malware classification. This result is useful because it clarifies that malware image saliency should not be treated as equivalent to object saliency in natural images. For malware images, future masking strategies should incorporate executable structure, such as binary regions or PE-section boundaries, rather than relying only on image saliency and square masks.
Improvements for AI systems
Improvements to AI Systems:
-
Domain-Adaptive Augmentation Selection: Implement an AI system that automatically detects the input domain (e.g., natural images vs. grayscale malware binaries) and switches between saliency-guided cutout and random cutout based on precomputed domain-specific performance profiles. For malware, it would default to random cutout or no cutout; for natural images, it would use low-saliency cutout.
-
Saliency-Aware Mask Shape Generation: Replace fixed square masks with structure-aware masks for malware. The improved system would use executable metadata (PE section boundaries, binary region entropy) to generate irregular, contiguous masks that align with code/data sections, rather than relying on pixel-level saliency. This would preserve semantic chunks of the binary while occluding others.
-
Saliency Reliability Scoring: Add a preprocessing module that computes a
saliency reliability index
for each image (e.g., variance of saliency map, contrast with background). If the index is low (as in grayscale malware), the system automatically reduces the influence of saliency-guided augmentation and falls back to random cutout. This prevents harmful high-saliency masking. -
Multi-Objective Augmentation Tuning: Build an AI trainer that jointly optimizes cutout area, number of augmented copies (M), and saliency placement strategy per dataset using a small validation set. For malware, it would learn that larger M with random placement is better; for CIFAR-100, it would learn that low-saliency with small area is optimal.
-
Saliency-Transfer Detection: Create a meta-learning layer that tests whether saliency maps from a pre-trained natural-image model transfer to a new target domain. If transfer fails (measured by correlation between saliency and classification-relevant regions), the system automatically disables saliency-guided augmentation and instead uses domain-specific priors (e.g., binary structure for malware).
What the Improved AI System Can Do:
-
Achieve higher classification accuracy on malware binaries by avoiding harmful saliency-guided cutout and instead using PE-structure-aware masks, potentially exceeding the 72.83% baseline.
-
Automatically adapt its augmentation strategy across heterogeneous image types (natural photos, medical scans, binary executables) without manual tuning.
-
Prevent performance degradation in low-saliency domains by detecting when saliency maps are unreliable and switching to random or structure-based occlusion.
-
Provide explainable decisions about why a particular augmentation was chosen, based on domain detection and saliency reliability scores.
-
Reduce training time and computational waste by not running ineffective saliency-guided experiments in domains where they are known to fail.
Abstract
Dropout regularization is commonly used to reduce overfitting by removing parts of a neural network during training. For Convolutional Neural Networks (CNN), cutouts serve a somewhat analogous purpose. Cutouts can be implemented as data augmentation: the original training image is retained, and additional copies are created with regions removed. In this chapter, we test whether cutout placement can be improved by using High-Resolution Class Activation Mapping (HiResCAM). We compare four controlled training conditions: no cutout, standard random cutout, low-saliency cutout, and high-saliency cutout. We experiment using grayscale malware images from the RawMal-TF dataset (17 families with 1,000 samples per family), and for comparison to natural images, we experiment with the well-known CIFAR-100 dataset. All experiments are based on ResNet18 with 100 training epochs. For the cutout experiments, we test cutout areas of 5%, 10%, 20%, and 30%, and we consider M in4,8 augmented copies per original training image. The RawMal-TF results are slightly worse for all three cutout cases (random, high and low saliency) as compared to no cutouts. In contrast, our CIFAR-100 experimental results improve slightly under low-saliency cutout. These results suggest that the value of saliency-guided cutout is domain dependent, and that malware images should not be treated as equivalent to natural images.
Sources
- A Comparison of Selected Image Transformation Techniques for Malware Classification
- RawMal-TF: Raw Malware Dataset Labeled by Type and Family
- Improved Regularization of Convolutional Neural Networks with Cutout
- Use HiResCAM instead of Grad-CAM for faithful explanations of convolutional neural networks
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models