Machine Learning-Based Cyber Defense for Cloud Infrastructure: An Adaptive Deep Q-Network Architecture for Intelligent Intrusion Detection and Automated Threat Mitigation
Md Yassir Mottalib, Md Yousuf, Eklachur Rahman Bhuiyan, S M Ahsan Habib, Sonjoy Kumar Dey, Md. Salahuddin Gazi, Molay Kumar Roy, Asaduzzaman Anik
Wilmington University · University of Houston-Downtown · Washington University of Science and Technology · South Dakota School of Mines & Technology · South Dakota State University · St. Francis College · Stanton University
cs.CR, cs.AI
Submitted: 2026-08-12
Updated: 2026-08-13
Comments: THis paper already peer reviewed
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
Terminology
Summary
Summary
This paper proposes a reinforcement learning-based dynamic cyber defense framework for cloud infrastructures, deploying a Deep Q-Network (DQN) to train effective defensive strategies against evolving cyberattacks. The authors state: "With the increasing complexity of cyber assaults in cloud environments, adaptable security solutions are needed that can support real-time detection and autonomous response. In this paper, we propose a reinforcement learning-based dynamic cyber defense framework. We deploy a Deep Q-Network (DQN) to train effective defensive strategies to counteract the evolving cyberattacks."
The methodology leverages two public datasets: We leverage the CICIDS2017 dataset for model creation and the UNSW-NB15 dataset for external validation, involving preprocessing of data, feature engineering, and adaptive policy learning.
The CICIDS2017 dataset, sourced from Kaggle or the Canadian Institute for Cybersecurity, contains approximately 2.83 million records with around 78 features covering attack categories including DDoS, DoS, Brute Force, Botnet, Port Scan, Web Attack, Infiltration, Heartbleed, and Benign traffic. The UNSW-NB15 dataset, from the UCI Machine Learning Repository, contains 257,673 records with 49 features covering Fuzzers, Exploits, DoS, Generic, Worms, Shellcode, Analysis, Backdoor, and Reconnaissance attack types.
Data preprocessing involved removing duplicates and missing values, handling infinite values and outliers, min-max scaling of numerical variables, one-hot and label encoding of categorical variables, and applying the Synthetic Minority Oversampling Technique (SMOTE) to address class imbalance. The data was split into 70% training, 15% validation, and 15% test sets. Feature extraction captured flow-level statistical features including flow duration, packet counts, packet lengths, inter-arrival times, protocol types, TCP flag statistics, and other network behavior indicators. Feature engineering included Pearson correlation analysis to remove highly correlated features (correlation coefficient > 0.90), feature importance estimation using Random Forest and XGBoost, creation of behavioral features such as packet rate ratios and session entropy, and Principal Component Analysis (PCA) for dimensionality reduction while preserving 95% of variance.
The authors compared the proposed DQN with decision tree, support vector machine, random forest, XGBoost, and multilayer perceptron models. The results show: "The proposed DQN achieves an accuracy of 99.72%, a precision of 99.68%, a recall of 99.65%, an F1-score of 99.66%, and an ROC-AUC of 0.999, while the false positive rate is 0.31%, the false negative rate is 0.35%, and the detection latency is 15 ms. The framework achieved a 99.54% attack mitigation rate,
demonstrating strong adaptive and real-time defensive capabilities."
For comparison, the Decision Tree achieved 96.11% accuracy, 96.20% precision, 96.15% recall, 96.15% F1-score, 0.972 ROC-AUC, 2.90% FPR, 3.89% FNR, and 18 ms detection time. The Support Vector Machine achieved 96.84% accuracy, 97.12% precision, 97.05% recall, 97.08% F1-score, 0.981 ROC-AUC, 2.15% FPR, 2.95% FNR, and 26 ms detection time. Random Forest achieved 99.18% accuracy, 99.05% precision, 99.01% recall, 99.03% F1-score, 0.996 ROC-AUC, 0.81% FPR, 0.99% FNR, and 20 ms detection time. XGBoost achieved 99.36% accuracy, 99.31% precision, 99.24% recall, 99.27% F1-score, 0.997 ROC-AUC, 0.66% FPR, 0.76% FNR, and 19 ms detection time. The Multilayer Perceptron achieved 98.71% accuracy, 98.56% precision, 98.43% recall, 98.49% F1-score, 0.992 ROC-AUC, 1.34% FPR, 1.57% FNR, and 23 ms detection time.
Reinforcement learning-specific performance metrics included an average episodic reward of 982.45, maximum reward of 1000.00, average convergence episodes of 285, policy stability of 99.1%, attack mitigation success rate of 99.54%, adaptive response accuracy of 99.48%, average decision time of 15 ms, and cloud resource utilization efficiency of 96.8%. The authors note: The reinforcement learning agent converged quickly after about 285 training episodes, after which the cumulative rewards were stable with minor fluctuations.
The paper discusses the fundamental limitation of supervised learning approaches: "Supervised learning models such as decision trees, random forests, support vector machines (SVM), artificial neural networks (ANN), and extreme gradient boosting (XGBoost) have been shown to perform very well in detecting intrusions by analyzing historical network data... However, these techniques are fundamentally limited because they need to be retrained when new attack techniques are discovered. Cyberattacks are always developing; thus, models trained on outdated data may become ineffective against new threats."
The authors argue that reinforcement learning differs from supervised learning because "it considers cybersecurity as a sequential decision-making problem, where the system observes the state of the network, decides on an action, receives a reward or a penalty as feedback, and improves its strategy over time. The DQN agent
learns optimal protection strategies through interaction with the cloud network and
doesn't just mimic past attack patterns but rather adapts its protection strategy in real-time depending on input, enabling proactive adaptability to changing cyber-attacks."
The paper identifies research gaps addressed by their work: "Most prior work either only considers a limited number of attack types or only focuses on restricted components such as intrusion detection or firewall rules, instead of establishing all-in-one adaptive security frameworks for cloud systems. Many models are tested using old or generated data that does not represent real-world cloud traffic today. Another problem is that there are no fair side-by-side comparisons of reinforcement learning vs. the best-performing supervised machine learning algorithms on the identical test settings."
The discussion section highlights practical applications: "Cloud computing and real-time cybersecurity organizations in the US can use deep Q-Network reinforcement learning. Intelligent security systems for large cloud infrastructures must respond to evolving cyber threats without operator interaction. Healthcare cloud systems for EHRs, medical imaging, telemedicine, and IoMT devices can integrate the proposed solution... Reinforcement learning can handle account takeover, DDoS, credential stuffing, fraud, and insider threats in banking, insurance, digital payment systems, and investment platforms."
The conclusion states: "We presented a reinforcement learning–based dynamic cyber protection architecture for cloud infrastructure that uses a deep Q-network (DQN) to continually learn optimal security rules from the cloud environment... DQN outperformed decision tree, support vector machine, random forest, XGBoost, and multilayer perceptron across numerous assessment measures in experiments. The framework improved accuracy, precision, recall, F1-score, and ROC-AUC, with decreased false positive and false negative rates and shorter detection latency."
Future work directions mentioned include investigating Double Deep Q-Network (DDQN), Dueling DQN, Proximal Policy Optimization (PPO), Soft Actor-Critic (SAC), and Multi-Agent Reinforcement Learning platforms such(MARL)
and incorporating real-time cloud telemetry, federated learning, explainable artificial intelligence (XAI), graph neural networks, and LLM-assisted threat intelligence.
The authors also suggest validating the framework on enterprise cloud environments and commercial cloud platforms as Amazon Web Services (AWS), Microsoft Azure, and Google Cloud Platform (GCP).
Improvements for AI systems
Improvements to AI Systems Based on This Paper:
- Adaptive Sequential Decision-Making for Threat Response
-
Implement a Deep Q-Network (DQN) layer that treats cybersecurity as a Markov Decision Process, enabling the AI to select defensive actions (e.g., isolate nodes, adjust firewall rules, rate-limit traffic) in real-time based on current network state, rather than relying solely on static classification.
-
The improved system can autonomously reconfigure defenses without human intervention, learning optimal policies from reward signals (e.g., +1 for blocking an attack, -1 for false positive) and adapting to novel attack patterns that supervised models would miss.
- Hybrid Supervised + Reinforcement Learning Pipeline
-
Use supervised models (XGBoost, Random Forest) as a fast initial filter for known attack signatures, then feed their probabilistic outputs as state features into the DQN agent. This reduces the action space and improves sample efficiency.
-
The improved system achieves both high detection accuracy (99.7%+) and proactive response, with the DQN handling zero-day or polymorphic attacks that supervised models fail to flag.
- Real-Time Feature Engineering and Dimensionality Reduction
-
Integrate the paper’s preprocessing pipeline (Pearson correlation >0.90 removal, Random Forest/XGBoost feature importance, PCA retaining 95% variance) into the AI’s online learning loop, so it continuously prunes redundant features and recomputes behavioral metrics (e.g., session entropy, packet rate ratios) as traffic evolves.
-
The improved system maintains low detection latency (15 ms) even under high-throughput cloud traffic, by dynamically reducing input dimensionality and prioritizing the most discriminative features for the current attack phase.
- Reward Shaping for Imbalanced and Evolving Threats
-
Design a reward function that penalizes false negatives more heavily than false positives (e.g., reward = +10 for blocking a real attack, -1 for false positive, -20 for missed attack) to align with the paper’s low FNR (0.35%) and high mitigation rate (99.54%).
-
The improved system becomes more conservative in uncertain states, reducing the risk of successful intrusions while still minimizing alert fatigue—critical for healthcare, banking, and critical infrastructure.
- Multi-Dataset Generalization and Transfer Learning
-
Train the DQN on CICIDS2017 and fine-tune on UNSW-NB15 using a domain-adversarial or progressive neural network approach, as the paper demonstrates cross-dataset validation.
-
The improved system can generalize across different cloud environments (AWS, Azure, GCP) and attack families (DDoS, fuzzers, exploits) without full retraining, reducing deployment time and computational cost.
- Explainable Action Logging for Compliance
-
Augment the DQN with an explainability module (e.g., SHAP values on state features, or a rule extractor from Q-values) that logs why a specific defensive action was chosen (e.g.,
blocked IP due to high packet rate ratio and low session entropy
). -
The improved system provides auditable decision trails, satisfying regulatory requirements in finance and healthcare, and enabling security analysts to trust and verify autonomous actions.
- Federated and Multi-Agent Extensions
-
Extend the single-agent DQN to a Multi-Agent Reinforcement Learning (MARL) framework where each cloud node runs a local agent, sharing distilled policies (not raw data) via federated learning.
-
The improved system scales to large distributed infrastructures, coordinates responses across nodes (e.g., containing a botnet spread), and preserves data privacy—aligning with the paper’s future work on MARL and federated learning.
- LLM-Assisted Threat Intelligence Integration
-
Use a large language model (LLM) to parse real-time threat feeds (e.g., CVE reports, dark web chatter) and convert them into structured reward-shaping signals or new state features for the DQN (e.g.,
increase suspicion score for port scan if new exploit is published
). -
The improved system anticipates attacks before they hit the network, reducing detection latency further and enabling preemptive defense actions, as suggested by the paper’s future work on LLM-assisted intelligence.
What the Improved AI System Can Do:
-
Detect and mitigate both known and zero-day attacks in cloud environments with >99% accuracy and <16 ms response time.
-
Autonomously adapt its defense strategy in real-time, learning from ongoing attacks without retraining on historical data.
-
Operate across heterogeneous cloud platforms and attack types, with minimal human oversight.
-
Provide transparent, auditable decisions for compliance and trust.
-
Scale to distributed cloud architectures via multi-agent coordination and federated learning.
-
Proactively incorporate external threat intelligence to stay ahead of emerging attack techniques.
Sources
- Shapley Value-Guided Adaptive Ensemble Learning for Explainable Financial Fraud Detection with U.S. Regulatory Compliance Validation
- Explainable Graph Neural Networks for Interbank Contagion Surveillance: A Regulatory-Aligned Framework for the U.S. Banking Sector
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs