Multi-Level Distributional Entropy from Flow Summary Statistics for Explainable Network Intrusion Detection

arXiv:2606.29797 · cs.CR, cs.AI, cs.LG · Submitted 2026-06-29 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "Multi-Level Distributional Entropy from Flow Summary Statistics for Explainable Network Intrusion Detection".

Elias: Multi-Level Distributional Entropy (MDE) is an analytical framework that derives interpretable entropy features directly from flow-level summary statistics at three levels—within-flow Gaussian differential entropy, crossdirectional Jensen-Shannon divergence (JSD),

Nadia: First, who's behind it and why it matters.

Title and authors: Nadia: So we're looking at "Multi-Level Distributional Entropy from Flow Summary Statistics for Explainable Network Intrusion Detection," and the title itself is pretty dense; it tells us this work is about taking flow statistics and turning them into some sort of entropy measure to explain what an AI model is doing.

Elias: It sounds like they're bridging the gap between having raw packet data, which you can't usually access in a pre-aggregated flow format, and using entropy measures that are already known to be good indicators of traffic structure.

Priya: I’m curious about what that means for us actually seeing the data; is this just another layer of complexity we have to interpret, or is it something fundamentally new?

Nadia: Well, essentially they're proposing a way to get interpretable features from pre-aggregated flow statistics without needing raw packet access or any specific training data.

Elias: That’s the key part; they are deriving these interpretable features directly from statistics like mean and standard deviation of packet sizes or inter-arrival times, which are already in the flow records.

Priya: So instead of having to train a complicated model on raw sequences, you're using these analytically defined entropy measures to characterize the traffic structure itself?

Nadia: Exactly; they’re saying that conventional flow statistics only capture magnitudes like byte counts and duration, but they miss the underlying distributional structure of how the data is actually distributed.

Elias: And by using Gaussian differential entropy for packet sizes or inter-arrival times, they’re trying to capture that structural complexity in a way that's mathematically grounded.

Priya: That’s interesting because conventional methods are often susceptible to those labeling artifacts, as the paper mentions regarding Engelen et al. thirteen <ref:2606.29797#pg1>.

Nadia: Right, and the authors are aiming for features that are inherently interpretable through SHAP, which is a huge deal for security analysts trying to understand alerts.

Elias: They’re leveraging that analytic definition of entropy so it’s grounded in information theory and has known ranges, which makes it more trustworthy than empirically motivated feature engineering.

Priya: I wonder how well this analytical approach holds up when we look at real-world traffic, especially things like encrypted flows where the Gaussian approximation might not be perfect.

Nadia: That’s a fair point; the paper does note that the framework operates under the assumption of Gaussian approximations for ADE and JSD, which means it might underestimate true entropy if the flow is multimodal or heavily encrypted.

Elias: Precisely; that limitation means we need to keep an eye on whether these features still hold up when we encounter traffic types that deviate significantly from a simple bell curve.

Priya: So, the title suggests they’re giving us a new lens through which to view network flow data for intrusion detection systems.

The paper's summary: Nadia: Moving past the title, the core summary of "Multi-Level Distributional Entropy from Flow Summary Statistics for Explainable Network Intrusion Detection" boils down to their proposal of a specific analytical framework called Multi-Level Distributional Entropy, or MDE.

Elias: They are proposing that this MDE framework constructs seven to twelve interpretable entropy features by looking at three distinct levels of flow statistics <ref:2606.29797#pg2>.

Priya: Could you elaborate on what those levels actually are? What's the mechanism they use to derive these features from the summary statistics?

Nadia: They break it down into three complementary levels: first, within-flow Gaussian differential entropy, which uses packet sizes and inter-arrival times to characterize variability.

Elias: Then there’s crossdirectional Jensen-Shannon divergence, which measures asymmetry between forward and backward traffic directions based on packet lengths.

Priya: And the third level involves TCP flag-pattern Shannon entropy, which they use to measure protocol irregularity by looking at the counts of different TCP flags per flow.

Nadia: That’s right; these features are designed to be "grounded in information theory with analytic definitions" and have closed-form expressions, which is what makes them so useful for interpretability via SHAP.

Elias: Because they don't require raw packet access or training data, the method is schema-independent and portable across any flow record that contains means and standard deviations.

Priya: That’s a huge practical win because it means we aren't locked into specific pipeline formats or requiring massive datasets just to generate features.

Nadia: It gives us a way to create structural fingerprints of traffic that is naturally sensitive to the difference between benign and malicious behavior, which is what we discussed earlier.

Elias: The paper essentially states that conventional flow statistics capture magnitudes but fail at capturing the distributional structure, and this MDE framework targets that structural aspect directly.

Priya: So the summary suggests this isn't just adding another feature to an existing pipeline; it’s a new way of characterizing the input data itself.

The paper's improvements: Nadia: The paper details several specific improvements they suggest, focusing on how MDE enhances current methods and what it allows us to do in practice.

Elias: One major improvement is the proposal of MDE as a method that constructs these features analytically from pre-aggregated flow statistics without needing raw packet access or training data for feature construction.

Priya: That eliminates the need for raw packet access, which I think is a significant practical advantage for many organizations working with existing network monitoring infrastructure.

Nadia: And then there's the development of a leakage-free protocol that reports the full operational metric suite alongside standard F1 scores across various settings, designed to surface failure modes that aggregate scores conceal.

Elias: That suggests they aren't just focusing on getting a single score, but on getting a comprehensive view of performance metrics like DR, false alarm rate (FAR), MCC, and precision-recall AUC alongside the F1.

Priya: I’m interested in how this leakage-free protocol helps us identify issues that a standard F1 score might hide about how well the system is actually performing under different operational scenarios.

Nadia: It allows them to expose failure modes like threshold-ranking divergence, which happens when the model's score ranking stays stable but fixed decision thresholds collapse under temporal distribution shift.

Elias: And they also address unseen attack families by showing that aggregate F1 can be driven entirely by the "ninety-nine point nine percent benign majority" in those scenarios <ref:2606.29797#pg2>.

Priya: So, this moves us from just checking if a model scores well to understanding the operational readiness of the system when it's actually deployed in a real environment.

Conclusion: Nadia: To wrap up, the conclusion of "Multi-Level Distributional Entropy from Flow Summary Statistics for Explainable Network Intrusion Detection" summarizes that MDE’s main implication is that we can now derive interpretable features directly from flow summaries using information theory.

Elias: They are confirming that the features derived via SHAP attribution are robust across different environments, with Spearman correlation values ranging from zero point eight zero to zero point nine five for all of them, which speaks to their stability in representation.

Priya: It sounds like the framework provides a stable way to ground these entropy attributions, which is important because it confirms that the features we're seeing aren't just artifacts of the model training process.

Nadia: They are confirming that these entropy attributions are reproducible and domain-coherent across structurally distinct environments, which means security analysts can trust those explanations more when they see them.

Elias: It’s a strong point, especially since they also showed noise robustness experiments where the features maintain discriminative power throughout all tested noise levels.

Priya: I just want to make sure we are clear about the limitations mentioned in the paper; they state that the framework operates under Gaussian approximations for ADE and JSD, which may underestimate true entropy for multimodal or encrypted flows.

Nadia: That’s a crucial caveat we have to keep in mind when deploying this technology because that limitation means we can't just assume perfect accuracy everywhere.

Elias: So, the MDE framework offers a strong analytical foundation, but it provides a way to generate features that are inherently interpretable through SHAP, even with those noted limitations regarding the Gaussian assumptions.

Centre for Intelligent Cloud Computing, CoE for Advanced Cloud, Faculty of Information Science and Technology, Multimedia University · Department of Communication Technology and Networking, Faculty of Computer Science and Information Technology, Universiti Putra Malaysia · Laboratory of Computational Science and Mathematical Physics, Institute for Mathematical Research, Universiti Putra Malaysia

cs.CR, cs.AI, cs.LG

Submitted: 2026-06-29

Updated: 2026-10-03

Comments: 21 pages. v2: retitled from "Multi-Level Distributional Entropy for Explainable Network Intrusion Detection"; the paper now foregrounds the analytic, packet-free derivation of entropy features from flow summary statistics and the leakage-free evaluation methodology that exposes failure modes an aggregate F1 conceals. Code: https://github.com/drbouke/mde

Code: https://github.com/drbouke/mde

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 85/100

The gist: Multi-Level Distributional Entropy (MDE) is an analytical framework that derives interpretable entropy features directly from flow-level summary statistics at three levels—within-flow Gaussian

Key concepts

Within-flow Gaussian differential entropy (L1)
This measures the structural complexity of traffic within a single flow by analyzing the mean and standard deviation of packet sizes or inter-arrival times. High values indicate traffic with highly variable and complex patterns, suggesting structural irregularity.
Crossdirectional Jensen-Shannon divergence (L2)
This concept captures asymmetry in network communication by comparing the distribution of packet lengths in forward versus backward directions. It quantifies how different the packet size characteristics are when traffic moves from source to destination compared to destination back to source.
Transmission Control Protocol (TCP) flag-pattern Shannon entropy (L3)
This feature assesses protocol irregularity by counting the occurrences of TCP flags per flow. Low entropy here signals flows dominated by a single flag type, which is a common indicator of specific attacks like SYN flooding.

Terminology

Summary

Multi-Level Distributional Entropy (MDE) is an analytical framework that derives interpretable entropy features directly from flow-level summary statistics at three levels—within-flow Gaussian differential entropy, crossdirectional Jensen-Shannon divergence (JSD), and Transmission Control Protocol (TCP) flag-pattern Shannon entropy—without requiring raw packet access or training data. This method addresses the limitations of existing intrusion detection systems by providing information-theoretic features that are natively interpretable via SHAP, while simultaneously exposing failure modes concealed by aggregate F1 scores across various operational scenarios.

How it works

MDE constructs 7–12 interpretable entropy features analytically from pre-aggregated network flow statistics at three complementary levels (L1, L2, and L3). These levels are:

  1. Within-flow Gaussian differential entropy (L1): Computed using the mean and standard deviation of packet sizes for both forward and backward directions, or inter-arrival times. High differential entropy indicates variable, structurally complex traffic.

  2. Crossdirectional Jensen-Shannon divergence (L2): Captures directional asymmetry between source-to-destination and destination-to-source traffic, measuring the distributional difference between forward and backward packet lengths.

  3. Transmission Control Protocol (TCP) flag-pattern Shannon entropy (L3): Measures protocol irregularity by analyzing per-flow TCP flag counts, where low entropy identifies flows dominated by a single flag type, such as in SYN flooding.

The framework is designed to be schema-independent and portable. The formulas apply to any flow record containing means, standard deviations, and packet counts without retraining or dataset-specific engineering. This analytical design ensures that the features are grounded in information theory with analytic definitions and possess closed-form expressions, which directly translates into interpretability via SHAP.

Key Contributions

The paper outlines three main contributions to the field:

  1. Proposal of MDE: A method that constructs 7–12 interpretable entropy features analytically from pre-aggregated flow statistics, requiring no raw packet access or training data for feature construction.

  2. Development of a Leakage-Free Protocol: A protocol that reports the full operational metric suite (DR, false alarm rate (FAR), Matthews Correlation Coefficient (MCC), and precision-recall AUC) alongside F1 across various settings, designed to surface failure modes that aggregate scores conceal.

  3. SHAP Attribution Analysis: An analysis quantifying the rank-stability and domain-coherence of analytical differential entropy (ADE) and JSD attributions across structurally distinct environments, confirming that entropy attributions are reproducible and domain-coherent.

Performance and Evaluation

The MDE features are evaluated across four benchmarks (NSL-KDD, CICIDS-2017, CICIDS-2018, UNSW-NB15) using tree-based classifiers like LightGBM and Random Forest. The results demonstrate that combined and conventional conditions achieve statistically indistinguishable within-distribution F1 on all tested datasets, confirming the entropy transformation does not degrade aggregate discriminative power. However, full operational metric reporting reveals critical gaps: for instance, on CICIDS-2018, an aggregate F1 of 0.74 hides a detection rate (DR) of 0.48, meaning fewer than half of attacks are detected.

Operational Insights and Stability

The evaluation protocols expose two failure modes that aggregate F1 conceals:

** Threshold-Ranking Divergence**

Under temporal distribution shift, the model's score ranking is preserved (AUC=0.87) but fixed thresholds collapse (DR=0.082) and recalibration offers no recovery. This suggests that while the relative ordering of flows remains stable, fixed decision boundaries fail under shift.

** Unseen Attack Families**

In unseen-attack-family experiments on CICIDS-2017, F1 exceeds 0.998 while DR=0 for held-out Infiltration and Bot families, showing that the aggregate score is driven entirely by the 99.9% benign majority.

Stability and Robustness

The SHAP fold-stability analysis confirms that entropy attributions are robust across environments, with Spearman correlation values ranging from 0.80 to 0.95 for all features. Furthermore, noise robustness experiments show that the features maintain discriminative power throughout all tested noise levels, with F1 degrading by only a small amount (e.g., ∆ = 0.049 at ε = 0.25), consistent with the information-theoretic nature of the features. The study concludes that MDE's contribution is primarily in representation grounding and evaluation methodology.

Limitations

The framework operates under the assumption of Gaussian approximations for ADE and JSD, which may underestimate true entropy for multimodal or encrypted flows.

Improvements for AI systems

As a fastidious and diligent researcher, I have analyzed the provided paper, Multi-Level Distributional Entropy for Explainable Network Intrusion Detection by Bouke et al. The core innovation lies in deriving interpretable, information-theoretic entropy features directly from pre-aggregated flow statistics without requiring raw packet access or training data.

Here are the specific improvements to AI systems and what the improved system can achieve:


)

Improvement Focus Area 1: Feature Engineering for Distributional Structure

The primary improvement is the integration of Multi-Level Distributional Entropy (MDE) features into existing machine learning pipelines.

  1. Maturity of Feature Space: Enhance current flow-level feature vectors (e.g., packet counts, byte volumes) by adding the 7–12 MDE entropy features derived from:

  2. Within-Flow Gaussian Differential Entropy (L1): Capturing the complexity/variance of packet sizes and inter-arrival times.

  3. Cross-Directional Jensen-Shannon Divergence (L2): Quantifying structural asymmetry between forward and backward traffic flows.

  4. TCP Flag-Pattern Shannon Entropy (L3): Measuring protocol irregularity based on control flag distributions per flow.

  5. What the Improved System Can Do:

The system moves beyond merely counting packets or bytes; it gains a structural fingerprint of the traffic flow that is inherently sensitive to the difference between benign and malicious behavior (e.g., scanners vs. human browsing). This allows for the detection of low-volume, structurally unique attacks (like port scans or C2 beaconing) that are invisible when only aggregate counts are used.

Improvement Focus Area 2: Enhanced Model Interpretability

The system incorporates SHapley Additive Explanations (SHAP) specifically for these entropy features.

  1. Interpretive Attribution: Apply SHAP analysis to the resulting classifier to assign feature importance values that reflect the contribution of MDE-derived entropy features to a specific intrusion prediction. The paper demonstrates high fold-stability (Spearman ρ = 0.80–0.95) for these attributions across different environments.

  2. What the Improved System Can Do:

The system moves from being a black box classifier to an explainable decision engine. Security analysts can immediately see why a specific flow was flagged (e.g., Flagged because the packet size distribution entropy (L1) was unusually low, indicating uniform probe packets). This drastically reduces analyst triage time and builds trust in automated security decisions in regulated environments.

Improvement Focus Area 3: Robust Operational Metrics

The system mandates the reporting of a full operational metric suite alongside standard F1 scores.

  1. Failure Mode Detection: Implement leakage-free evaluation protocols across temporal shifts, hold-out sets, and pseudo-live replaying of traffic (703K flows). This reveals critical failures that aggregate F1 conceals. Specifically, the system can detect threshold-ranking divergence, where a model's internal ranking ability (AUC) is preserved under shift but its fixed decision threshold collapses (Recall/DR drops to near zero).

  2. What the Improved System Can Do:

The system moves from score-driven performance assessment to operational readiness assessment. It identifies when a model is fundamentally brittle due to distribution shift, allowing security teams to proactively implement necessary adaptive recalibration strategies before a real-world deployment fails catastrophically.

Improvement Focus Area 4: Cross-Dataset Generalization Assessment

The framework includes protocols for cross-dataset transfer and unseen attack family testing.

  1. Transfer Gap Quantification: Systematically test the zero-shot performance of models trained on one dataset's entropy features when applied to another (e.g., CICIDS-2017 to CICIDS-2018). This quantifies the expected performance drop when attack taxonomies change, providing a crucial metric for model deployment planning.

  2. What the Improved System Can Do:

The system provides an honest assessment of model portability. It prevents organizations from deploying a model trained on one environment (e.g., NSL-KDD) and assuming it will perform equally well in a modern environment (e.g., UNSW-NB15), highlighting that domain adaptation is necessary for structural generalization, rather than just feature addition.


Summary of Overall System Capability:

The improved AI system will be a high-assurance IDS capable of:

  1. Detecting subtle, low-volume attacks by analyzing the distributional shape (entropy) of flow characteristics, not just their magnitudes.

  2. Providing clear, theoretically grounded justifications for every detection decision via SHAP explanations.

  3. Identifying the specific conditions (like temporal shift or novel attack types) under which its performance degrades from a ranking failure (AUC drop) versus a detection failure (DR collapse).

  4. Quantifying the necessary domain adaptation required when moving models across different network environments.

Abstract

Machine learning network intrusion detection systems (IDS) operate on aggregate flow statistics that discard the distributional structure of traffic, and although information-theoretic measures capture that structure, established entropy estimators require raw packet sequences that pre-aggregated flow datasets do not contain. No prior method derives entropy from the summary statistics those records already hold. We introduce Multi-Level Distributional Entropy (MDE), which computes interpretable information-theoretic features analytically from flow-level summary statistics at three levels, within-flow Gaussian differential entropy, cross-directional Jensen-Shannon divergence (JSD), and Transmission Control Protocol (TCP) flag-incidence Shannon entropy, with closed-form properties and no raw packet access; only imputation medians and score bounds are fitted on the training split. We pair the features with a leakage-free, fold-local evaluation protocol that reports the full operational metric suite across cross-validation, temporal, pseudo-live, cross-dataset, and unseen-attack-family settings, on four benchmarks (NSL-KDD, CICIDS-2017, CICIDS-2018, UNSW-NB15) with tree-ensemble classifiers and SHAP. The protocol exposes failure modes that aggregate weighted F1 conceals: on CICIDS-2018 an F1 of 0.73 hides a detection rate (DR) of 0.44, on held-out attack families F1 exceeds 0.998 while DR falls to zero, and a 703K-flow pseudo-live replay reveals a threshold-ranking divergence in which score ranking is largely preserved (area under the ROC curve, AUC, 0.84 to 0.86) while fixed-threshold detection collapses (DR 0.08). The entropy features match conventional features within 0.1 percentage points of F1 and receive reproducible SHAP attributions (Spearman 0.84 to 0.94), so their contribution is a grounded, interpretable representation and an evaluation methodology rather than an accuracy gain.

Sources

Related papers