Temporal Generalization and Explanation Stability of Control Flow Graph Neural Networks for Malware Detection
cs.CR, cs.AI, cs.LG
Submitted: 2026-09-21
Updated: 2026-09-21
Comments: 47 pages, 9 figures, 17 tables. Code available at https://github.com/Ho9pe/TG-CFG
Code: https://github.com/Ho9pe/TG-CFG
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Malware detection is a critical task in cybersecurity, and graph neural networks over control flow graphs have shown promising results for it.
Terminology
Abstract
Malware detection is a critical task in cybersecurity, and graph neural networks over control flow graphs have shown promising results for it. However, detectors are usually evaluated on a random split of a corpus collected over a single period, which cannot show how well a model generalizes to later samples. This study addresses that limitation with a strict temporal split: every model is trained on one period and scored once on a later one. Two corpora of control flow graphs, each node carrying 37 features, were extracted statically from 1,989 Windows portable executables: 459 graphs from 2024-2025 for training and 223 from 2026 for evaluation. Twelve variants and a flat-feature control were trained on the earlier corpus. The choice of message-passing operator changes robustness to the shift significantly, and every pairwise gap that survives correction separates an aggregating architecture from one built around a learned attentional readout. The ranking also reverses: the flat control, which sees node features but no topology, is the best in-distribution model and among the worst across the boundary, so a conventional benchmark would have rejected message passing. Neither recalibration nor ensembling substitutes for the operator choice. Attributions do not shift, but explanation validity is architecture-specific, and the most accurate operator on the later corpus is the hardest to explain. An architecture derived from the finding matches the best searched operator without search. The shift affects both malware and benign classes alike, so these are results about robustness to distribution shift, not malware evolution.
Sources
- Explainable Attention-Guided Stacked Graph Neural Networks for Malware Detection
- EMBER: An Open Dataset for Training Static PE Malware Machine Learning Models
- Microsoft Malware Classification Challenge
- Recent Advances in Malware Detection: Graph Learning and Explainability
- DeeperGCN: All You Need to Train Deeper GCNs
- Gated Graph Sequence Neural Networks
- Towards A Rigorous Science of Interpretable Machine Learning
- Fast Graph Representation Learning with PyTorch Geometric
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs