Concept Drift Detection and Adaptive Retraining of Malware Classification Models

arXiv:2608.13465 · cs.LG, cs.AI, cs.CR · Submitted 2026-08-13 · Read on arXiv

Christofer Washington Berruz Chungata, Martin Jurecek, Katerina Potika, William B. Andreopoulos, Mark Stamp

San Jose State University · Czech Technical University in Prague

cs.LG, cs.AI, cs.CR

Submitted: 2026-08-13

Updated: 2026-08-14

Comments: To appear as a chapter in the book "Artificial Intelligence for Cyber Defense in Emerging Threats", to be published by Springer by early 2027

Code: https://github.com/aleguma/kronodroid

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: Concept drift refers to changes over time in the statistical properties of data, as compared to the data that was used to train a learning model.

Terminology

Summary

Concept drift refers to changes over time in the statistical properties of data, as compared to the data that was used to train a learning model. Machine learning models for malware detection or classification are particularly susceptible to performance degradation caused by concept drift, as attackers constantly modify existing malware. In this chapter, we analyze two machine learning-based approaches to automated concept drift detection—a novel approach based on One-Class Support Vector Machines (OCSVM) and a previously-studied technique based on Minibatch K-Means (MK-Means). For comparison we also consider Maximum Mean Discrepancy (MMD), a statistical technique for detecting changes in multidimensional data. We conduct an extensive series of experiments comparing the effectiveness of four learning models, namely, Multilayer Perceptron (MLP), Random Forest (RF), Support Vector Machines (SVM), and eXtreme Gradient Boosting (XGB). For each of these models, we consider three distinct scenarios: A static scenario where no model retraining occurs, a periodic scenario where models are constantly retrained irrespective of concept drift, and a drift-aware scenario where models are only retrained when concept drift is detected. Under the drift-aware scenario, we analyze the tradeoff between accuracy and training efficiency using Pareto Front analysis. We find that all three concept drift detection techniques achieve classification accuracy comparable to periodic retraining, while offering substantially greater efficiency in terms of the number of models that must be retrained. In addition, drift-aware retraining based on our OCSVM technique generally outperforms the MK-Means and MMD approaches. Overall, these results provide strong evidence that we are able to accurately detect concept drift in malware classification models. Furthermore, our concept drift detection techniques are efficient and practical, and the process of updating learning models can easily be fully automated.

The study uses the KronoDroid dataset, which contains 41,382 Android malware samples belonging to 240 distinct malware families. The research focuses on the five malware families with the largest number of samples: Airpush, SMSReg, Malap, Boxer, and Agent. Each sample includes 200 static features and 289 dynamic features, with preprocessing removing non-numerical features, resulting in feature vectors of dimension 470. Samples are ordered temporally based on the HighestModDate timestamp and partitioned into consecutive temporal batches of size 50, with 30 samples per batch used for training and 20 for testing.

For each of the three scenarios, four classification models are considered, and for each model-scenario combination, 20 distinct combinations of malware families are examined. This results in 80 static experiments, 80 periodic retraining experiments, and 240 drift-aware retraining experiments (80 using each of the OCSVM, MK-Means, and MMD drift detectors), giving a total of 400 distinct experiments. Hyperparameter tuning is performed using Optuna with Tree-Structured Parzen Estimator search, running 100 trials per model, resulting in 737,600 total hyperparameter tuning combinations.

The static scenario trains a model on the initial temporal batch and uses it for all subsequent batches without retraining. The periodic scenario retrains models at regular intervals regardless of concept drift. The drift-aware scenario only retrains when concept drift is detected via one of the three detection methods. For MK-Means, drift is detected by computing the average silhouette coefficient for consecutive pairs of batches and comparing the change against a threshold. For OCSVM, an OCSVM model is trained on the initial batch, and the ratio of outliers to inliers is monitored over subsequent batches, with drift detected when this ratio changes beyond a threshold. For MMD, a two-sample test with a Gaussian kernel is performed between consecutive batches, with drift detected when the p-value falls below a significance level.

The experimental results show that periodic retraining yields the highest accuracy, with MLP achieving an average accuracy of 0.9666, followed by SVM at 0.9500, RF at 0.9529, and XGB at 0.9334. The static scenario achieves lower accuracies, with the best being SVM at 0.8416. The drift-aware scenario with OCSVM achieves accuracies close to periodic retraining, with the median accuracy for MLP being 0.9486, representing a 15% improvement over the static scenario and within 1.9% of the periodic scenario. For MK-Means, the median accuracy for MLP is 0.9480, and for MMD it is 0.9380.

Regarding efficiency, the drift-aware scenario offers substantial savings in the number of models that must be retrained. For OCSVM, the median efficiency for MLP is more than 65%, meaning only about 35% of batches require retraining. For MK-Means, the median efficiency is about 58%, and for MMD it is about 69%. The OCSVM technique also requires the least training time among the three drift detectors, with an average of 0.39 seconds per family pair, compared to 7.71 seconds for MK-Means and 1.85 seconds for MMD.

The Pareto Front analysis demonstrates the tradeoff between accuracy and efficiency, allowing for optimal hyperparameter selection. The results indicate that OCSVM generally outperforms both MK-Means and MMD across most metrics considered, making it the optimal choice among the three concept drift detection techniques analyzed. The paper concludes that concept drift detection techniques are efficient and practical, and the process of updating learning models can easily be fully automated.

Improvements for AI systems

Improvements to AI Systems:

  1. Adaptive Retraining Trigger: Implement a drift-aware retraining mechanism using One-Class SVM (OCSVM) as the primary detector. The AI system monitors the ratio of outliers to inliers in incoming data batches and automatically triggers model retraining only when this ratio shifts beyond a learned threshold. This reduces unnecessary retraining by 65% compared to periodic retraining, while maintaining accuracy within 1.9% of the periodic baseline.

  2. Temporal Batch Partitioning for Continuous Learning: Structure the AI system to process data in ordered temporal batches (e.g., size 50) rather than random shuffling. This enables the system to detect concept drift in real-time as data evolves, and to maintain a rolling training window (e.g., 30 samples per batch) that reflects the most recent statistical properties, improving long-term robustness against adversarial changes.

  3. Pareto-Optimal Hyperparameter Selection: Use Pareto Front analysis during model configuration to balance accuracy and training efficiency. The AI system can automatically select hyperparameters (via Optuna with Tree-Structured Parzen Estimator) that maximize classification performance while minimizing the number of retraining events, ensuring cost-effective operation in resource-constrained environments.

  4. Multi-Model Drift Detection Fusion: Integrate OCSVM as the primary drift detector, with MK-Means and MMD as fallback or ensemble members. The system can compare drift signals from all three methods and use a voting or confidence-weighted scheme to reduce false positives, improving detection reliability across different malware families and data distributions.

  5. Automated Model Lifecycle Management: Build a fully automated pipeline that: (a) trains an initial model on the first temporal batch, (b) continuously evaluates drift using OCSVM, (c) triggers retraining only when drift is detected, and (d) logs accuracy and efficiency metrics for each retraining event. This enables the AI system to self-maintain its performance without human intervention, as demonstrated by the 400-experiment validation.

What the Improved AI System Can Do:

  • Detect concept drift in malware classification with 94.86% median accuracy (MLP) using OCSVM, nearly matching periodic retraining (96.66%) while retraining only 35% of the time.

  • Operate in three modes (static, periodic, drift-aware) and automatically switch to drift-aware mode when efficiency is prioritized, or periodic mode when maximum accuracy is required.

  • Handle high-dimensional feature vectors (470 dimensions) from static and dynamic malware analysis, processing 41,382 samples across 240 families, with focus on top-5 families (Airpush, SMSReg, Malap, Boxer, Agent).

  • Reduce training time by 80% compared to MK-Means (0.39s vs 7.71s per family pair) and 79% compared to MMD (0.39s vs 1.85s), enabling near-real-time adaptation to new malware variants.

  • Provide a tradeoff analysis interface that shows users the accuracy vs. efficiency frontier, allowing them to select an operating point (e.g., 95% accuracy with 50% retraining savings) based on deployment constraints.

Abstract

Concept drift refers to changes over time in the statistical properties of data, as compared to the data that was used to train a learning model. Machine learning models for malware detection or classification are particularly susceptible to performance degradation caused by concept drift, as attackers constantly modify existing malware. In this chapter, we analyze two machine learning-based approaches to automated concept drift detection-a novel approach based on One-Class Support Vector Machines (OCSVM) and a previously-studied technique based on Minibatch K-Means (MK-Means). For comparison we also consider Maximum Mean Discrepancy (MMD), a statistical technique for detecting changes in multidimensional data. We conduct an extensive series of experiments comparing the effectiveness of four learning models, namely, Multilayer Perceptron, Random Forest, Support Vector Machines, and eXtreme Gradient Boosting. For each of these models, we consider three distinct scenarios: A static scenario where no model retraining occurs, a periodic scenario where models are constantly retrained irrespective of concept drift, and a drift-aware scenario where models are only retrained when concept drift is detected. Under the drift-aware scenario, we analyze the tradeoff between accuracy and training efficiency using Pareto Front analysis. We find that all three concept drift detection techniques achieve classification accuracy comparable to periodic retraining, while offering substantially greater efficiency in terms of the number of models that must be retrained. In addition, drift-aware retraining based on our OCSVM technique generally outperforms the MK-Means and MMD approaches. Overall, these results provide strong evidence that we can accurately detect concept drift in malware classification models.

Sources

Related papers