Dueling Deep Q-Learning for Intrusion Detection

arXiv:2608.11291 · cs.CR, cs.LG · Submitted 2026-08-11 · Read on arXiv

Logan Luna, Matthew P. Berkowitz, Laxima Niure Kandel, Sirio Jansen-S'anchez

Embry-Riddle Aeronautical University · Georgia Institute of Technology

cs.CR, cs.LG

Submitted: 2026-08-11

Updated: 2026-08-13

Comments: 6 pages, 5 figures. Published in Proc. IEEE SoutheastCon 2025, pp. 1192-1197

Journal ref: L. Luna, M. P. Berkowitz, L. Niure Kandel, and S. Jansen-S\'anchez, "Dueling Deep Q-Learning for Intrusion Detection," in Proc. IEEE SoutheastCon 2025, pp. 1192-1197, 2025

DOI: 10.1109/SOUTHEASTCON56624.2025.10971436

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 81/100

The gist: This study proposes a novel approach to intrusion detection systems (IDS) by employing a reward-based, dueling Q-learning model, achieving an average accuracy of 99.68% across multiple attack classes.

Terminology

Summary

This study proposes a novel approach to intrusion detection systems (IDS) by employing a reward-based, dueling Q-learning model, achieving an average accuracy of 99.68% across multiple attack classes. The proposed model has a dueling network architecture which separates its predictions into value and advantage streams, which has the benefit of improving learning efficiency and stability. The model was trained on the CIC-IDS2018, a benchmark dataset based on real-world intrusion detection scenarios, having multiple attack classes such as DDoS, botnets, and brute-force attacks. Furthermore, Explainable AI (XAI), specifically SHAP (SHapley Additive exPlanations), was also integrated into the training and evaluation process to provide interpretability into the model’s predictions.

The study highlights that traditional signature-based detection methods rely on prior knowledge of attack patterns, limiting their ability to recognize novel or evolving threats, while anomaly detection approaches face issues with high false-positive rates. Reinforcement Learning (RL) offers a promising alternative since the model learns from a reward structure rather than being based on purely labeled data, optimizing its policy around the influence of its action on the environment rather than just classifying it correctly, making it more adaptable for environments where attack types constantly evolve.

The contributions of the paper include: employing a dueling architecture which separates the value stream to estimate the value of the current state V(s), independent of actions, and the advantage streams to estimate the advantage A(s,a) of each action, combining both streams to compute the final Q-values, enhancing stability and convergence during training; training on the CIC-IDS2018 dataset spanning multiple attack types such as DDoS, botnets, and brute-force attacks, utilizing 2,177,804 samples; utilizing SHAP to provide insight into the key features influencing the model’s predictions; and achieving an average accuracy of 99.68% across multiple attack types, outperforming previously implemented RL-based methods for IDSs.

The methodology employs a Dueling Q-Network (DQN) Framework with parameters including hidden layers of [128, 64], batch size of 128, learning rate of 0.001, gamma (discount) of 0.99, epsilon start of 1.0, epsilon end of 0.1, epsilon decay of 0.999, memory size of 10,000 experiences, target update frequency of every 1000 steps, episode count of 200, using a CUDA GPU Nvidia 3060, Adam optimizer, and MSE loss function. The agent identifies specific attack types, providing higher accuracy than traditional models but with increased computational requirements. The architecture decomposes the Q-value into a value stream to estimate the state value V(s), and an advantage stream to estimate the action advantage A(s,a), improving learning efficiency and stability.

A custom environment, NetworkClassificationEnv, was designed using OpenAI Gym to facilitate interaction of the agents with network traffic data for both classification tasks and reward-based training. The environment processes network traffic sequentially, treating each data point as an individual network flow based on its timestamp. The reward function is structured as: rt = +1 · Sl · Ca + min(0.5 · log(streak), 2.0) if at = lt, rt = −1 · Sl · Ca if at ≠ lt, and rt = 0 if no action, where Sl is the severity weight based on the true label lt, Ca = 0.5 + confidence/2, and streak is the number of consecutive correct predictions.

The results demonstrate that the proposed DQN Agent achieves an average accuracy of 99.68%, successfully classifying various attack types present in network traffic. The classification report shows precision, recall, and F1-scores for each class: Benign (0.9998, 0.99996, 0.99987), Botnet (0.9979, 0.9886, 0.9932), Brute-force (0.9870, 0.9991, 0.9931), DDoS attack (0.9972, 1.0000, 0.9986), DoS attack (0.9996, 0.9937, 0.9966), and Web attack (0.0000, 0.0000, 0.0000). The Web attack was excluded from analysis due to its limited sample size of just 173 samples. The study notes that while web attacks are rare, approaches such as weighting the reward function higher or generating mock data were considered, but mock data can lead to certain repeated features in the data being used to classify the attack instead of what features should actually be used.

The proposed framework displays notable improvements over similar previous studies, with Alavizadeh et al. achieving 88% accuracy while the DQN attains 99.68% accuracy in multi-class classifications. The enhanced performance is attributed to the incorporation of a Dueling DQN architecture along with more samples being used to model the training environment (2,177,804 samples versus 219,980 samples used by Alavizadeh et al.).

SHAP explainability results reveal key findings: RST Flag Cnt is highly impactful in identifying attacks like DoS or scanning, indicating abnormal termination of connections; PSH Flag Cnt reflects urgency in data transmission, often linked to buffer overflow or data exfiltration attempts; Bwd Pkt Len Max captures the largest backward packet being sent from server-to-client, often signifying exfiltration activities; Init Bwd Win Byts represents the initial window size received in the backward flow, with abnormal values often associated with SYN flood DoS attacks; ACK Flag Cnt measures acknowledgment packets and helps differentiate benign flows from brute-force attacks based on handshake patterns; and Flow Byts/s and Flow Pkts/s measure the rate of packet traffic, with elevated rates correlating with high-speed attacks like DDoS.

The limitations of the current framework include limited classification granularity (performing only high-level classification of attack types rather than fine-grained attack differentiation), deployment in real-world scenarios (the model was trained and evaluated in the same environment), and the need for a simulated enterprise environment to test the model with actual attacks. Future research directions include incorporating temporal features for evolving attack scenarios, exploring advanced techniques such as multi-agent systems, and optimizing the framework for real-time deployment in large-scale networks.

Improvements for AI systems

Improvements to AI Systems:

  1. Adaptive Threat Response via Reward-Shaped Reinforcement Learning
  • Implement a reward function that dynamically weights attack severity (Sl) and model confidence (Ca), enabling the AI to prioritize high-impact threats (e.g., botnets, DDoS) over benign traffic.

  • Add a streak-based bonus to reinforce sustained correct predictions, improving stability in long-running network monitoring.

  • Capability: The system can autonomously adjust its detection sensitivity in real time, reducing false negatives for critical attacks while maintaining low false-positive rates for benign flows.

  1. Dueling Q-Network with Explainable Action Selection
  • Integrate a dueling architecture that separates state-value (V(s)) and action-advantage (A(s,a)) streams, allowing the AI to learn which network states are inherently risky vs. which actions (classifications) add value.

  • Pair this with SHAP-based feature attribution at each decision step, so the AI can justify why it selected a specific attack class (e.g., DDoS due to high Flow Byts/s and ACK Flag Cnt).

  • Capability: The system provides transparent, real-time explanations for every classification, enabling security analysts to trust and audit AI decisions in critical infrastructure.

  1. Handling Rare Attack Classes via Severity-Weighted Oversampling
  • Modify the reward function to assign exponentially higher rewards for correctly classifying rare attacks (e.g., web attacks) without generating synthetic data (which caused feature leakage).

  • Use a dynamic epsilon-greedy policy that explores rare classes more aggressively during training, while exploiting common classes during deployment.

  • Capability: The AI can detect low-frequency threats (e.g., web attacks) with improved recall, without sacrificing performance on dominant classes—addressing the current 0% F1-score for web attacks.

  1. Temporal Sequence Modeling for Evolving Attack Patterns
  • Extend the environment to process network flows as time-series sequences (using recurrent or attention-based layers) rather than independent samples, capturing attack evolution (e.g., slow DDoS ramping or multi-stage botnet behavior).

  • Incorporate the timestamp-based sequential processing from NetworkClassificationEnv into the Q-network’s state representation.

  • Capability: The system can predict and preemptively flag emerging threats before they fully manifest, improving proactive defense in real-world, dynamic networks.

  1. Multi-Agent Collaborative Detection
  • Deploy multiple Dueling Q-Network agents, each specialized for a subset of attack types (e.g., one for volumetric DDoS, another for stealthy botnets), coordinated by a meta-controller that fuses their Q-values.

  • Use the SHAP insights to assign feature importance weights per agent, enabling specialization (e.g., agent A focuses on Flow Byts/s, agent B on Bwd Pkt Len Max).

  • Capability: The system scales to large enterprise networks with distributed detection, reducing computational load per node while maintaining high accuracy across diverse attack vectors.

  1. Confidence-Aware Decision Fusion
  • Use the Ca = 0.5 + confidence/2 term not only in rewards but also as a gating mechanism for final classification—if confidence is low, the system flags the flow for human review rather than making a hard decision.

  • Capability: The AI reduces false alarms in ambiguous traffic, improving operational efficiency for security teams while still catching novel attacks.

  1. Real-Time Deployment Optimization
  • Prune the Dueling Q-Network (e.g., via quantization or knowledge distillation) based on SHAP feature importance, retaining only the most impactful features (e.g., RST Flag Cnt, Init Bwd Win Byts) for inference.

  • Implement a target network update frequency scheduler that adapts based on traffic volatility (e.g., faster updates during attack spikes).

  • Capability: The system runs on edge devices or mid-range GPUs (like the Nvidia 3060 used) with lower latency, enabling inline network traffic inspection without bottlenecking throughput.

  1. Cross-Domain Generalization via Reward Calibration
  • Train the agent on CIC-IDS2018, then fine-tune the reward function’s severity weights (Sl) using a small labeled dataset from a new network environment, while freezing the feature extraction layers.

  • Capability: The AI can be rapidly deployed in new organizational networks with minimal retraining, adapting its detection priorities to local threat profiles (e.g., healthcare vs. finance).

Sources

Related papers