HYDRA: Proactive Android Malware Drift Adaptation via Hierarchical Graph Contrastive Learning
cs.CR, cs.LG
Submitted: 2026-09-22
Updated: 2026-09-22
Comments: Accepted at ACM CCS 2026. Author's version with full appendix. 17 pages
Code: https://github.com/hzcheney/HGCL
License: http://creativecommons.org/licenses/by/4.0/
The gist: Concept drift, driven by the rapid evolution of Android malware, severely degrades the performance of machine learning detectors.
Terminology
Abstract
Concept drift, driven by the rapid evolution of Android malware, severely degrades the performance of machine learning detectors. Current adaptation strategies are often reactive, responding only after performance has dropped and imposing a significant manual annotation burden, or they are proactive but rely on unstable adversarial training and incomplete, single-level graph representations. To overcome these limitations, we propose HYDRA (Hybrid Drift Adaptation), a proactive adaptation framework that learns drift-invariant representations from hierarchically structured data. HYDRA first models applications using a hybrid graph structure, combining fine-grained Control Flow Graphs (CFGs) and coarse-grained Function Call Graphs (FCGs) to capture comprehensive behavioral patterns. It then introduces a novel cross-domain contrastive learning objective that aligns historical (source) and new (target) data distributions. By generating pseudo-labels for unlabeled target samples, our method pulls representations of semantically similar applications together, regardless of their domain, within a single, stable optimization process. This approach unifies feature learning and domain alignment, eliminating the need for complex adversarial objectives. Extensive experiments on large-scale, time-ordered malware datasets demonstrate that HYDRA achieves substantially lower False Negative and False Positive Rates than state-of-the-art baselines while requiring up to 87.5% fewer labeled samples. Our work thus offers a robust and efficient solution to combat concept drift in security applications.
Sources
- Evasion Attacks against Machine Learning at Test Time
- A Survey on Malware Detection with Graph Representation Learning
- HiGraph: A Large-Scale Hierarchical Graph Dataset for Malware Analysis
- A Large-Scale Database for Graph Representation Learning
- Semi-Supervised Classification with Graph Convolutional Networks
- TESSERACT: Eliminating Experimental Bias in Malware Classification across Space and Time
- Graph Attention Networks
- GoGNN: Graph of Graphs Neural Network for Predicting Structured Entity Interactions
- How Powerful are Graph Neural Networks?
- Semantic-preserving Reinforcement Learning Attack Against Graph Neural Networks for Malware Detection
- Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs