Evolution of Log-Based Detection Rules in Public Repositories
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Evolution of Log-Based Detection Rules in Public Repositories".
Jane: The paper was written by Minjun Long and David Evans from University of Virginia.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Jane, this new paper is going to change how we think about security rule maintenance.
Jane: You're talking about "Evolution of Log-Based Detection Rules in Public Repositories," aren't you?
Tom: That's the one, and it comes from Minjun Long and David Evans at the University of Virginia.
Jane: I love that they're looking at public repositories like Sigma and Splunk's Security Content.
Tom: It's a smart move because we can't really see the rules being used inside private companies.
Jane: Since we can't see those, these public repositories give us a window into how experts actually refine their work.
Lu: That window is actually a goldmine for training future autonomous security agents, don't you think?
Tom: Lu, are you suggesting we could use these evolution traces to teach AI how to self-correct?
Lu: Precisely, because these repositories show the exact path from a rough idea to a polished, working rule.
Meng: That sounds great in theory, but how much of this "evolution" is actually useful for a real-world engineer?
Jane: Well, Meng, the paper suggests that even though it's public, it mirrors the real pressures analysts face.
Meng: I suppose if they are fighting the same false positives we see in the field, then the data is highly relevant.
Lalam: It's also a beautiful example of how collective human intelligence is being recorded for the future.
Tom: That's a deep way to put it, Lalam, but it really does show a growing culture of shared security knowledge.
Jane: It's definitely more than just a collection of files; it's a historical record of defensive thinking.
Tom: We should look closer at what they actually found in those files next.
Summary: Jane: We're moving into the actual results of "Evolution of Log-Based Detection Rules in Public Repositories" now.
Tom: The most striking number is that roughly fifty-six percent of the rules undergo at least one revision to their detection logic.
Jane: So, more than half of these rules aren't just "set it and forget it" tools.
Tom: Right, and the researchers found that the evolution is mostly non-monotonic.
Jane: That sounds complicated, but it just means they aren't just constantly adding more and more conditions, right?
Tom: Exactly, they're adding clauses and then later removing them to balance things out.
Lu: It's like watching a sculptor constantly chipping away and then adding more clay to find the perfect shape.
Meng: That constant chipping away sounds like the struggle to keep alert volumes from exploding in a SOC.
Jane: It really is, Meng, because the paper shows they are constantly swinging between expanding coverage and reducing false positives.
Meng: That tug-of-war is exactly what my team deals with every single day.
Lalam: This constant oscillation shows that security isn't a destination, but a continuous, living process.
Tom: It really highlights that there isn't one "perfect" rule that stays perfect forever.
Jane: Instead, they are always adjusting to stay effective against new threats.
Tom: Let's talk about how they actually managed to track all these subtle changes.
Methodology: Tom: We've talked about what changed, but Jane, how did they actually compare these different versions?
Jane: They used something called a predicate graph intermediate representation, or PGIR.
Tom: So instead of just looking at the text, they look at the logical skeleton of the rule?
Jane: Yes, that way they can tell the difference between a simple text edit and a real change in detection behavior.
Tom: They also used a tree alignment procedure to match up the different versions of the rules.
Jane: That's really clever because it allows them to ignore superficial changes like reordering lines.
Lu: If we can represent logic as these clean graphs, we could build AI that understands the "why" behind a rule.
Meng: Can they actually use an LLM to figure out the reason for a change, though?
Tom: They actually did, Meng, and they used GPT-five to classify the operational intent of each revision.
Jane: They found out if a change was meant to catch more bad actors or to stop the rule from being too noisy.
Meng: Using an LLM to automate that kind of semantic analysis would save engineers so much time.
Lalam: It moves us from just seeing "what" changed to understanding the human intent behind the security decision.
Tom: It's a massive step up from just looking at a git diff.
Jane: We're reaching the end of our discussion, so let's wrap this all up.
Conclusion: Tom: Well, we've spent a lot of time on "Evolution of Log-Based Detection Rules in Public Repositories."
Jane: It's clear that detection rules are dynamic, living things that require constant, non-linear maintenance.
Tom: The researchers showed us that the fight between coverage and precision is never truly over.
Jane: And they've given us a much better way to measure that fight using structural logic.
Lu: I see a future where these evolution patterns help us create self-healing security systems.
Meng: For me, the impact is in the tools; we need better ways to manage this constant rule churn.
Lalam: Ultimately, this research helps us bridge the gap between human expertise and automated defense.
Tom: Thanks to everyone for joining us today.
Jane: We'll see you next time for the next paper!
University of Virginia
cs.CR, cs.SE
Submitted: 2026-05-06
Updated: 2026-09-28
Code: https://github.com/Elena6918/Evolution-of-Log-Based-Detection-Rules
Importance score: 86/100
The gist: This paper presents "the first longitudinal analysis of detection rule evolution across two widely used repositories: the community-driven Sigma project and the curated Splunk Security Content
Key concepts
- Log-Based Detection Rules
- These are security rules used to monitor logs for suspicious activity. The paper examines how these rules evolve in public repositories like Sigma and Splunk's Security Content, offering insight into real-world security maintenance.
- Non-monotonic Evolution
- This refers to the pattern of rule changes where developers don't just add conditions. Instead, they constantly add clauses and then later remove them to refine the rule, balancing effectiveness and complexity.
- PGIR (Predicate Graph Intermediate Representation)
- A technical method used by researchers to analyze rules. Instead of looking at raw text, it examines the logical skeleton of the rule, allowing detection of real changes in behavior rather than just superficial edits.
Terminology
Summary
This paper presents the first longitudinal analysis of detection rule evolution across two widely used repositories: the community-driven Sigma project and the curated Splunk Security Content (SSC).
The study aims to understand when rules change in ways that alter detection behavior, how those changes manifest structurally, and what operational pressures they reflect such as false-positive reduction or coverage expansion.
To facilitate this analysis, the authors introduce a predicate graph intermediate representation (PGIR) that canonicalizes the logical structure of a rule, together with a tree alignment procedure for analyzing changes across revisions.
This PGIR isolates a rule’s detection logic, allowing us to compare rules across renames and reorganizations, distinguish behavior-altering changes from maintenance edits, and analyze how detection logic evolves longitudinally.
This structural analysis is paired with an LLM-based classification step that annotates each adjacent-version pair with a semantic direction and operational rationale.
Applying this method to nine years (2017–2026) of commit history across the Sigma and SSC repositories,
the researchers find that roughly 56% of rules undergo at least one revision on detection logic.
The study reveals that evolution is predominantly non-monotonic, with over half of rules both adding and removing clauses over time.
Regarding maintenance, revisions persist long after creation: 8.1% of edited rules in Sigma and 14.2% in SSC do not see their first predicate change until more than two years post-introduction.
Structurally, the authors find that structural edits are coordinated rather than isolated: 42% of predicate-changing steps carry multiple co-occurring operations.
In terms of specific logic changes, conjunctive additions outpace disjunctive additions by 3 to 6 times.
At the lineage level, more than half of predicate-changing rules in both datasets evolve non-monotonically, mixing expansion and contraction over their lifetime,
with the MIXED
pattern being the most common (56.3% in Sigma and 56.1% in SSC). Additionally, the study identifies structural reversion (A–B–A patterns),
which occur in 9.2% of all predicate-changing lineages in Sigma... and 25.4% in SSC.
By using LLM-based inference to determine semantic intent, the study finds that coverage expansion is the most common rationale in both corpora (41.3% Sigma; 35.9% SSC).
At the lineage level, more than half of multi-revision lineages reverse direction at least once across their history, and roughly a quarter oscillate repeatedly between coverage expansion and false-positive reduction.
The authors identify three distinct mechanisms driving this non-monotonicity:
-
Oscillating lineages: These
expose unresolved precision–coverage tension
wherecompeting implementations repeatedly replace one another because the better operational choice depends on deployment feedback.
-
Coupled rules: These
expose a stable maintenance regime in which coverage expansion and exclusion growth are intrinsically coupled by the detection strategy itself.
-
IE-only rules: These
expose a representational limit
wherethe directional effect cannot be inferred from rule text alone without external knowledge of schemas, parsers, or telemetry populations.
Ultimately, the research demonstrates that detection rule evolution in public repositories reflects ongoing operational trade-offs rather than steady convergence,
and that a substantial share carry unresolved precision–coverage tensions rather than converging to a stable form.
Improvements for AI systems
1. PGIR-Integrated Semantic Rule Generation (Text-to-Logic-to-Syntax)
-
The Improvement: Shift LLM-based rule authoring from raw text generation to a multi-stage pipeline that utilizes the Predicate Graph Intermediate Representation (PGIR). The model would first generate a canonicalized logical AST (Boolean operators and atomic predicates) before translating it into target syntax (Sigma, SPL, etc.).
-
What the improved system can do: It would eliminate
syntactic hallucinations
where a model produces valid-looking code that is logically broken. Crucially, it would enableCoupled-Aware Generation,
where the AI proactively suggests necessary exclusions (e.g.,Since you are adding this expansion predicate, I am also adding these three exclusion predicates to maintain the precision-coverage balance
).
2. Automated Detection Engineering Lifecycle Management (The Evolutionary Monitor
)
-
The Improvement: Implement an AI agent that monitors rule lineages using the Weighted Predicate-Logic Change Cost Model (d pred) and the identified structural evolution patterns.
-
What the improved system can do: Instead of merely alerting on rule changes, the agent would perform
Design Loop Detection.
It would identifyOscillating
rules—those repeatedly switching between competing implementations (e.g., transaction-based vs. stats-based)—and instead of suggesting more tuning, it would recommend architectural refactors, such as splitting a single unstable rule into two distinct, specialized rules to resolve the precision-coverage tension.
3. Semantic-Logic Diffing and Impact Analysis Tools
-
The Improvement: Replace traditional line-based
git difftools with a PGIR-based structural alignment engine for security operations. -
What the improved system can do: When an analyst reviews a rule update, the AI would provide a
Semantic Impact Report.
This report would distinguish betweenMaintenance Edits
(renames, whitespace, formatting) andBehavior-Altering Edits
(Boolean restructuring). It would explicitly predict the direction of the change: "This update is a Coverage Expansion (CE) that will increase detection surface but likely increase false positives,or
This is a False-Positive Reduction (FPR) via structural enrichment."
4. Evolutionary Trajectory Data Augmentation for LLM Training
-
The Improvement: Use the paper’s taxonomy of rule histories (Coupled, Oscillating, Phased, and Mixed) to create a synthetic training dataset for fine-tuning security-specific LLMs.
-
What the improved system can do: Current models are trained on static snapshots of rules. An improved model trained on
Evolutionary Chains
would understand the temporal intent of detection engineering. It would learn to recognize that a rule's logic is not a static truth but a moving target, allowing it to better assist analysts in predicting how a rule should evolve in response to new telemetry or observed false-positive trends.
Sources
- LSTM-Based System-Call Language Modeling and Robust Ensemble Method for Designing Host-Based Intrusion Detection Systems
- RuleGenie: SIEM Detection Rule Set Optimization
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs