Jev-IDS: System One Models for Network Intrusion Detection
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: Today's paper: "Jev-IDS: System One Models for Network Intrusion Detection".
Elias: Machine-learning Network Intrusion Detection Systems (IDS) often depend on substantial labeled datasets, but this work presents JEV-IDS,
Nadia: First, who's behind it and why it matters.
Title and authors: Nadia: So, we're diving into the paper "Jev-IDS: System One Models for Network Intrusion Detection," and it looks like the authors are tackling a real headache in machine learning for security, especially when you don't have tons of labeled data.
Elias: I’m curious about that title; "System One Models" suggests they're moving away from those free-form text generation models we often see, focusing instead on typed probabilistic decisions.
Priya: From a privacy perspective, I wonder how relying on such structured decision interfaces affects the overall data handling and what kind of inferences we can reliably draw from the flow data itself.
Nadia: Exactly what I mean is that they're making a specific choice about how the AI should respond to network traffic, which is interesting because it directly impacts how we get actionable security intelligence.
Elias: Right, and the core idea seems to be serializing one flow per request and asking the JEV system two very specific questions: a binary attack probability and a finite-choice traffic category.
Priya: That sounds like a very controlled way to get classification because it forces the model into predefined decision spaces rather than letting it wander off into potentially misleading interpretations.
Nadia: That’s the key, Priya, because when you are hunting for zero-day intrusions where you have zero prior labels, having a calibrated probability and a specific category is much more useful than just a general text output.
Elias: The paper mentions this system is built around JEV, which they describe as the first model from the System One Model family designed for these typed probabilistic decisions instead of free-form text generation.
Priya: So it’s not just about using an LLM in a new way, but fundamentally changing the interface so that we get structured answers instead of open-ended text which can be hard to parse reliably.
Nadia: Precisely, and this is where the practical value really shines for us as researchers trying to build deployable systems that don't require massive labeled datasets upfront.
Elias: They lay out a four-stage pipeline: first constructing an ordered Flow representation using a Dataset Card, then building the JEV request from a fixed template with optional examples, then evaluating those two typed questions by JEV, and finally extracting the detection outputs.
Title and authors: Priya: That flow of construction sounds very methodical; it implies that the quality of the initial flow representation is crucial because everything downstream depends on that structured input.
Nadia: It is, and I'm particularly interested in how they handle those examples during the request construction phase, since that links directly to their findings regarding label scarcity.
Elias: The experimental setup they use is pretty rigorous; they compare JEV-IDS against a few different approaches like a structured-output LLM such as Gemini three point six Flash, a Random Forest, and an Isolation Forest for unsupervised reference.
Priya: Comparing it to established methods like Random Forest and Isolation Forest gives us a good baseline to see if this new SOM approach actually delivers on its promise without introducing new, unquantifiable risks.
Nadia: I'm looking at the results now, and the numbers they present are pretty compelling because they show JEV-IDS performing better than both of those conventional machine learning baselines under specific conditions.
Elias: Indeed, the paper reports that at k=one JEV was four point eight times faster and three point eight times cheaper than GPT-five point six Luna, while also showing a one point five times higher novel-attack recall compared to that generative model.
Priya: That speed and cost improvement is significant, especially when we think about deploying network detection systems at scale where latency and operational expense are major concerns for organizations.
Nadia: And the false alarm rate reduction is also a big deal; they found that JEV produced fifteen times fewer false alarms than a low-data Random Forest when k=one. That suggests higher precision without sacrificing recall much.
Elias: It’s interesting to see how the cost comparison plays out, as they noted that at k=one JEV exhibited approximately seven point seven times lower request latency and a twenty-two times lower estimated list-price cost than Gemini.
Priya: From a measurement standpoint, those latency and cost figures are concrete metrics we can use to justify the shift away from using very large generative models for every single network flow inspection.
Nadia: The paper also investigates generalization to novel attacks, and they found that once labeled examples are introduced into the system, JEV consistently achieved higher recall than Gemini on novel attack types.
Elias: But they also made a cautious observation about the learning process itself, noting that neither model benefits monotonically from additional examples; this suggests there might be some saturation point for learning in these scenarios.
Title and authors: Priya: That’s a fair caveat, because it shows that we can't just keep adding data and expect performance to always climb indefinitely when dealing with truly novel threats.
Nadia: So, the main implication seems to be that JEV-IDS provides a distinct trade-off where you get strong detection performance without the massive computational overhead of the largest generative models.
Elias: That distinction between predictive performance and inference efficiency is what really sets this work apart when you compare it to purely generative approaches like Gemini three point six Flash, which achieves high overall F1 but at a higher cost profile.
Priya: It really frames the discussion around practical deployment: can we achieve near-top performance using a model that runs much more efficiently than the most powerful generative AI available?
Nadia: Absolutely; this paper shows that System One Models, when applied to flow classification, offer a very efficient way to manage uncertainty in intrusion detection tasks.
Elias: It opens up avenues for designing security systems where we can trade some of the absolute highest recall for much better operational efficiency and lower inference costs.
Priya: That kind of trade-off is what matters when you think about real-world security budgets and the sheer volume of traffic we have to process constantly.
Nadia: So, as we wrap up on this paper, Jev-IDS really proves that structured decision interfaces can be a highly effective tool for detecting zero-day intrusions under label scarcity.
Elias: It’s a solid piece of work because it doesn't just claim accuracy; it quantifies the trade-off between how well you detect and how much it costs to run the system.
Priya: I think what’s most important is that JEV-IDS demonstrates that we can get strong detection performance even when only a few labeled examples are available, which opens up possibilities for rapid response in zero-day situations.
Nadia: It certainly shows how to make those few labeled examples work hard by structuring the AI's decision process correctly through the System One Model approach.
Elias: Well, this study on Jev-IDS is a great example of moving toward more efficient, specialized probabilistic models for complex tasks like network intrusion detection.
Priya: It’s a nice reminder that in applied security research, efficiency and reliability are just as important as achieving the highest possible accuracy scores.
The paper's summary: Nadia: So, to recap, JEV-IDS is this new approach that uses System One Models to classify network traffic flows by asking two specific questions—a probability of attack and a traffic category—which lets them do it even when they don't have many labeled examples.
Elias: That structured output is key; it means the AI isn't just spitting out long text, but giving you something you can actually plug into a security tool immediately.
Priya: And what I’m seeing from the data is that this method maintains high precision while dramatically cutting down on the time and resources needed for inference compared to those big generative models we usually use for classification.
Nadia: Exactly, it’s all about that trade-off between getting a very accurate answer and making sure the system can run fast enough in a real network environment without costing too much.
Elias: I'm looking at the architecture described, and it seems they’ve successfully constrained the model into a probabilistic decision space rather than letting it wander into unstructured text generation, which is what makes this approach so different from other LLM applications.
Priya: From a privacy standpoint, the fact that the system serializes just one flow per request and asks these specific questions means we have a very clear boundary on what data is being processed during the detection decision itself.
Nadia: That’s important because if you can clearly define the input-output structure, it makes auditing that AI’s behavior much more straightforward for security teams trying to understand what it’s actually doing.
Elias: The experimental setup they used, pitting JEV against things like Gemini and Random Forest, really highlights how the System One Model handles the scarcity of labeled data better than those other methods when you're dealing with zero-day threats.
Priya: I’m particularly interested in their findings on generalization; if it can actually perform well on novel attacks just by being given a handful of examples from that new category, that’s huge for real-time security operations.
Nadia: That would mean we could deploy these kinds of detectors in environments where the threats are constantly evolving, without needing to wait weeks for massive retraining cycles on a whole new dataset.
Elias: The implication here is moving away from a monolithic approach where you need endless data just to tune a general-purpose model, toward something more specialized and efficient that can handle specific decision tasks with minimal initial training input.
Priya: So, the world gets a detection system that is both powerful enough to catch new threats and smart enough not to waste massive amounts of compute power on classifying benign traffic repeatedly.
Nadia: That’s the core idea, Priya—a powerful tool that’s lean and fast enough for line-rate inspection while still offering strong probabilistic defense against unknown intrusions.
Elias: It really shows how carefully designing the interface, like JEV’s typed questions, can yield tangible performance benefits when you have resource constraints involved.
Priya: So, the next thing we need to figure out is whether this efficiency gain actually translates into a reliable security posture in complex enterprise networks where traffic patterns are highly variable.
The paper's improvements: Nadia: So, to wrap up on the improvements section, the authors are suggesting a few ways we can take this JEV-IDS concept and make it even more useful in real-world security scenarios.
Elias: They are focusing heavily on making that structured output even more machine-consumable and deterministic, which is really smart because it cuts down on any ambiguity that could be exploited by an attacker trying to manipulate the detection logic.
Priya: I'm seeing a lot of talk about hybrid architectures, where you might use this efficient System One Model as a first line of defense for routine traffic and only bring in a more powerful generative LLM when things get truly weird or suspicious.
Nadia: That tiered response sounds like exactly what we need for massive network monitoring; you want the fast filter handling the common stuff and the heavy AI reserved for those rare, high-value anomalies.
Elias: And I noticed they’re emphasizing grammar-constrained output, like forcing a JSON format, which ensures that whatever decision is made—whether it's from JEV or another model—is immediately readable by other security orchestration tools without any messy text parsing required.
Priya: That structured enforcement is key for measurement because it means we can reliably track the performance metrics across different types of attacks and traffic without getting bogged down in subjective interpretation of the model’s reasoning.
Nadia: It's about making sure that when an alert fires, security teams know exactly what they're looking at, and that level of clarity is vital when trying to figure out how to exploit a weakness or patch a vulnerability quickly.
Elias: And they also touched on few-shot generalization mechanisms, which means the system can learn to spot new attack types using just a small set of examples for that specific type, which directly addresses the label scarcity issue we discussed earlier.
Priya: If we can achieve reliable detection against zero-day variants with minimal labels, it significantly lowers the barrier to entry for deploying intrusion detection systems in environments where threat intelligence is constantly being generated on the fly.
Nadia: That makes these systems much more accessible to smaller organizations or even individual researchers who don't have access to massive security datasets.
Elias: The implication is that we can design models that are inherently adaptive and efficient, rather than static classifiers that break when the threat landscape shifts even slightly.
Priya: Ultimately, it feels like they’re pushing toward systems that are not just accurate on paper but robust enough to handle the messy, unpredictable nature of real network data in a way that is both fast and measurable.
Conclusion: Nadia: So, to wrap up, we've seen how JEV-IDS uses System One Models to provide robust network intrusion detection even when labeled data is scarce, focusing on that efficiency trade-off between performance and cost.
Elias: It really shows a solid approach to making probabilistic decisions in security contexts without relying on the massive training sets that most current generative models demand.
Priya: From my side, the results confirm that this method doesn't just deliver high scores; it provides a measurable way to balance detection capability with operational cost, which is what matters for real-world deployment.
Nadia: It’s a great demonstration of how we can build practical tools that are fast and reliable enough to run continuously on live network traffic inspection.
Elias: I'm still thinking about the constraints they put on the JEV interface; it really proves that you can constrain an AI's output format to ensure it serves a deterministic purpose in a high-throughput environment.
Priya: And when we look at how well it generalizes to novel attacks, even with limited examples, that suggests this technology could be very useful for catching zero-day threats before they become widespread.
Nadia: That means we’re talking about security systems that can actually keep up with the speed of modern attack creation without needing a constant stream of new labeled data.
Elias: I think the real impact here is in how it changes the design philosophy; instead of always chasing higher F1 scores at any cost, you start prioritizing inference efficiency alongside accuracy metrics.
Priya: That shift in priority is what makes this work important for researchers focused on practical security applications where resources are always finite and every computation counts.
Nadia: So, we’ve seen how JEV-IDS handles the zero-day challenge under label scarcity by focusing on structured, efficient decision-making.
Elias: It's a strong piece of work because it doesn't just claim better accuracy; it proves that a specialized model architecture can offer significant operational savings compared to its larger counterparts.
Priya: It leaves us wondering how this kind of efficient detection could integrate into larger security ecosystems and what the long-term privacy implications are as these models become more embedded in network infrastructure.
Nadia: That's a big question for our next discussion, because scaling this efficiency while maintaining that low latency is going to be the real test.
Paulo Severo, Silvio E. Quincozes, Amanda Dias
Federal University of Pampa
cs.CR, cs.AI
Submitted: 2026-10-01
Updated: 2026-10-01
Code: https://github.com/jev-ids/jev-ids
Project page: https://jev-ids.github.io
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 90/100
The gist: Machine-learning Network Intrusion Detection Systems (IDS) often depend on substantial labeled datasets, but this work presents JEV-IDS, an open experimental general NIDS based on the Jev System One
Key concepts
- JEV
- JEV is the core model from the System One Model family used by JEV-IDS. Unlike models that generate free text, JEV is designed for making typed probabilistic decisions based on a structured program state. It asks specific questions about network flows, such as whether a flow is an attack or normal traffic.
- System One Model (SOM)
- SOM is the family of models used to build JEV-IDS. This model structure is specialized for handling typed probabilistic decisions rather than general text generation. It allows the system to receive a structured input and provide pre-specified, finite outputs, making it suitable for structured tasks like intrusion detection.
- Flow Representation
- This refers to how a network traffic flow is organized into data that the IDS can understand. JEV-IDS constructs this representation using a 'Dataset Card' that defines specific features and categories for each flow. This structured input is necessary for JEV to accurately analyze the traffic.
- Labeled Examples (k)
- This refers to the limited amount of labeled data available during testing or training, quantified by 'k'. The experiment tests how well JEV-IDS performs when k is very small (zero-shot inference) versus when a few labeled examples from specific traffic categories are provided. This helps assess performance under real-world data scarcity.
Terminology
Summary
Machine-learning Network Intrusion Detection Systems (IDS) often depend on substantial labeled datasets, but this work presents JEV-IDS, an open experimental general NIDS based on the Jev System One Model (SOM) to detect zero day intrusions under label scarcity.
How it works
JEV-IDS serializes one flow per request and asks JEV two questions: a binary attack probability and a finite-choice traffic category. The system is built around JEV, which is the first model from the System One Model (SOM) family designed for typed probabilistic decisions instead of free-form text generation. This interface receives a structured program state, where the input is a structured program state, and possible outputs are specified in advance.
The pipeline comprises four main stages:
-
The construction of an ordered Flow representation, defined by a Dataset Card that specifies features and categories.
-
The construction of the JEV request from a fixed request template and optional labeled Examples.
-
The evaluation of two typed questions by JEV: one asking whether the Flow represents an intrusion rather than normal traffic, and another asking for a traffic category choice.
-
The extraction of the corresponding detection outputs from these answers.
Experimental Methodology
The experimental methodology was designed to assess JEV-IDS under different levels of labeled-data availability and to compare its behavior with representative generative, supervised, and unsupervised approaches. All detectors are evaluated on the same target flows and feature representation whenever applicable, while the number of labeled examples available during prediction or training is explicitly controlled.
The experiments consider three factors:
**: The detector (JEV-IDS, structured-output LLM like Gemini 3.6 Flash, Random Forest, and Isolation Forest). **
**: The number of labeled examples per category, denoted as **
**: The sampling seed. **
For JEV and the generative baseline, the budget for labeled examples considered was a set of values:
-
The condition where zero-shot inference is used (k = 0).
-
Conditions where k > 0 provides k labeled flows from each of the five traffic categories, with specific values being considered as k ∈ 0, 1, 2, 4, 8.
The Random Forest was evaluated with k ∈ 1, 2, 4, 8 and additionally with the complete labeled training pool. The Isolation Forest serves as an unsupervised reference and is trained exclusively on benign flows from the training pool.
Comparison Methods
The comparison methods involve evaluating three distinct approaches:
-
Structured-output LLM (Gemini 3.6 Flash) as the generative baseline, which receives the same ordered flow representation, traffic-category descriptions, and labeled examples supplied to JEV under the corresponding experimental condition.
-
Random Forest as a supervised machine-learning baseline, trained using exactly the same labeled flows made available to JEV and the generative baseline for each finite value of k.
-
Isolation Forest as an unsupervised reference, trained exclusively on benign flows from the training pool, whose anomaly score is used to distinguish normal from anomalous traffic.
Results and Trade-offs
The results position JEV differently from both the generative LLM and the conventional machine-learning baselines by providing a distinct trade-off between predictive performance and inference efficiency. Gemini achieved higher overall F1 across the evaluated labeled-example budgets, largely through higher attack recall, whereas JEV maintains comparable precision with substantially lower inference latency and estimated cost.
Key findings include:
**: At k = 1, JEV was 4.8 times faster and 3.8 times cheaper than GPT-5.6 Luna, with 1.5 times higher novel-attack recall; it also produced 15 times fewer false alarms than a low-data Random Forest. **
**: The results position JEV differently from both the generative LLM and the conventional machine-learning baselines by providing a distinct trade-off between predictive performance and inference efficiency. **
**: The cost behavior also reflects differences in token usage and provider accounting, with JEV exhibiting approximately 7.7× lower request latency and a 22× lower estimated list-price cost than Gemini at k = 1. **
The study also investigates generalization to novel attacks, finding that once labeled Examples are introduced, JEV consistently obtains higher recall than Gemini on novel attack types, although the statistical evidence is strongest at the smallest Example budgets. The results also show that neither model benefits monotonically from additional Examples.
Conclusion
The work indicates that JEV can achieve strong detection performance even when only a few labeled Examples are available, obtaining the best overall results among evaluated approaches for considered settings. The main finding is not that a SOM consistently surpasses a generative LLM in detection accuracy, but that it provides a distinct trade-off between predictive performance and inference efficiency.
Improvements for AI systems
Here are specific improvements to existing AI systems based on the Jev-IDS paper, along with what those improved systems can achieve:
-
Inference Efficiency and Cost Optimization for Flow-Based NIDS:
-
Adoption of System One Models (SOM) over Autoregressive LLMs for Low-Latency, High-Throughput Detection:
-
Few-Shot Generalization and Zero-Day Attack Detection Under Label Scarcity:
-
Structured Output Control for Programmatic Consumption of Intrusion Decisions:
-
Hybrid Detection Architectures Combining Efficient SOMs with Generative LLMs for Cascaded Response Systems:
Specific capabilities of the improved systems:
-
A flow-based Network Intrusion Detection System (NIDS) that processes traffic asynchronously, achieving significantly lower inference latency (up to 8.6x faster than GPT-5.6 Luna at low label budgets) and substantially reduced operational costs (up to 25x lower list price cost), making it suitable for high-rate, line-rate processing where the primary constraint is invocation frequency and monetary expenditure rather than raw predictive power.
-
A detection component that uses a System One Model (like JEV) instead of a general LLM for per-flow classification. This system can provide two distinct, typed outputs: a calibrated attack probability and a finite-choice traffic category, allowing downstream security policies to make deterministic decisions based on these structured values without parsing unstructured text.
-
A few-shot learning mechanism integrated into the NIDS that enables the system to accurately classify flows belonging to attack types it has never seen during training (zero-day or unseen attacks), provided a small set of labeled examples from those categories are supplied per request. This allows security teams to rapidly deploy detection capabilities against emerging threats without requiring retraining on massive, task-specific datasets.
-
A system that enforces grammar-constrained, structured output (e.g., JSON format) from the decision model (whether SOM or LLM). This ensures that the classification output is machine-readable and immediately consumable by other security tools or orchestration systems for automated alerting and response triggering, eliminating ambiguity caused by free-form text generation.
-
A hybrid detection pipeline where an efficient, low-cost SOM acts as a fast filter for high-confidence or straightforward traffic, while a more powerful generative LLM is selectively invoked only when the SOM output indicates high uncertainty or when investigating novel/unseen attack patterns. This architecture allows for a tiered response: low-latency, cheap handling of common threats and expensive, high-recall analysis for edge cases.
Abstract
Machine-learning Network Intrusion Detection Systems (IDS) depend on substantial labeled datasets and task-specific training, whereas Large Language Models (LLMs) detection can analyze flow records directly but incurs higher inference cost and latency, with less constrained outputs. This paper presents JEV-IDS, an open experimental general NIDS based on the Jev System One Model (SOM) to detect zero day intrusions Under label scarcity. JEV-IDS serializes one flow per request and asks JEV two questions: a binary attack probability and a finite-choice traffic category. Our results show that, at k=1, JEV was 4.8 times faster and 3.8 times cheaper than GPT-5.6 Luna, with 1.5 times higher novel-attack recall; it also produced 15 times fewer false alarms than a low-data Random Forest. Across 5,400 decisions on a 300-flow NSL-KDD pilot split, JEV achieved F1-Score 0.859, precision 0.941, recall 0.790, and novel-attack recall 0.838. Increasing k to 2 reduced its F1-Score to 0.839.
Sources
- From Flows to Words: Can Zero-/Few-Shot LLMs Detect Network Intrusions? A Grammar-Constrained, Calibrated Evaluation on UNSW-NB15
- Are We Shooting Flies with Cannons? Trade-off Analysis for AI-based 5G Intrusion Detection
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs