Operationalizing Cyber Threat Intelligence with GraphRAG
Atul Kabra, Prakhar Paliwal, Manjesh K. Hanawal
Indian Institute of Technology Bombay
cs.CR, cs.AI
Submitted: 2026-08-13
Updated: 2026-08-14
Comments: 12 pages
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: This paper investigates whether using a knowledge-graph-based retrieval system (Microsoft GraphRAG) instead of a standard vector-similarity retrieval system (Naive RAG) produces threat hunting plans
Terminology
Summary
This paper investigates whether using a knowledge-graph-based retrieval system (Microsoft GraphRAG) instead of a standard vector-similarity retrieval system (Naive RAG) produces threat hunting plans that rely more on durable, hard-to-evade clues (TTPs) rather than fragile indicators like IP addresses and file hashes. The authors build a fully on-premise pipeline that converts CTI PDFs into threat hunting plans through three retrieval back-ends—GraphRAG Local Search, GraphRAG Global Search, and Naive vector RAG—all driven by the same generation prompt and the same locally-hosted language model (gpt-oss-20b). A second, cybersecurity-tuned language model (Foundation-Sec-8B-Instruct) grades every plan against a ten-point rubric split across two tiers: foundation quality (60 points) and Pyramid-of-Pain resilience (40 points).
The central finding is that GraphRAG retrieval pushes the dominant Pyramid level of generated plans upward and increases the fraction of detections that survive IOC rotation. In a detailed case study of one APT28 report, the GraphRAG plan kept firing at 100% of its detections after every IP address, domain, and file hash in the report was rotated, while the Naive RAG plan kept firing at only 29%.
Repeating the comparison across nine real CTI reports from four vendors confirms the same pattern: GraphRAG plans consistently reach higher, harder-to-evade levels of the pyramid, even when the two systems end up close on total score.
The paper makes four contributions: (1) a reproducible, fully on-premise pipeline that converts CTI PDFs into threat hunting plans through three retrieval back-ends; (2) a multi-turn LLM-as-Judge harness built specifically for detection-engineering output, using a cybersecurity-domain judge model with explicit score floors, ceilings, and a step-by-step counting procedure designed to resist reward-hacking; (3) two empirical evaluations—a single-report deep-dive producing criterion-level scores for all three pipelines on an APT28 advisory, and a breadth experiment across nine real CTI reports drawn from four vendors; (4) a characterisation of two failure modes—silent failure of GraphRAG Global Search on short reports, and JSON parsing failures in the small judge model on long output.
The results show that on the APT28 deep-dive, GraphRAG Local plans retain 100% of their detection surface after a complete IOC rotation, while Naive RAG plans retain 29%. On the full nine-report breadth experiment, "GraphRAG reaches the durable L4–L7 layer on every successful report and reaches L5 or higher on three of seven Run 1 reports; Naive RAG clusters at L4 with a long fragile L1–L3 tail and reaches L5 or higher on zero reports excluding the FancyBear special case. The paper concludes that
GraphRAG is the architecturally superior approach for IOC-resilient threat hunting plan generation, while noting that
the wording of the generation prompt matters almost as much as the retrieval back-end itself"—the V1-to-V2 prompt change produced a fifty-five to seventy-four point swing on identical retrieval back-ends.
The paper also identifies three operational implications for SOC teams: (1) deploy GraphRAG Local as the primary retrieval back-end, as it is the only mode whose plans reach dominant L7-TTP on multi-page CTI reports; (2) deploy GraphRAG Global as a secondary mode with an output-length fallback to detect silent failures; (3) retain Naive RAG only for fact-lookup queries, as it produces plans whose useful life is measured in hours. Future work includes scaling to fifty to one hundred reports, replacing the 8B judge with a 13B-class cybersecurity judge, and implementing a Local-Search fallback when GraphRAG Global returns too little output.
Improvements for AI systems
Improvements to AI Systems:
-
Add a retrieval-backend selector to RAG pipelines. The improved system automatically chooses between graph-based local search, graph-based global search, and vector similarity based on document length and query type. For multi-page threat reports, it defaults to graph local search; for short reports, it triggers a fallback to local search when global search produces under a threshold of output tokens (e.g., <200 tokens). This prevents silent failures and ensures durable, TTP-level answers.
-
Integrate a
Pyramid-of-Pain resilience scorer
as a post-generation filter. The improved system runs every generated plan through a lightweight, domain-tuned judge model (e.g., 8B-class) that scores each detection on a 1–7 scale (L1=hash, L7=TTP). It automatically rejects or re-generates any plan where >50% of detections fall below L4, and it re-ranks final outputs to prioritize plans with higher median pyramid level, not just total score. -
Implement an IOC-rotation stress test as a built-in evaluation step. Before finalizing a threat hunting plan, the improved system automatically rotates all IPs, domains, and file hashes in the source report (using placeholder values) and re-simulates the detection logic. If the plan’s detection rate drops below 80% after rotation, the system flags it as fragile and triggers a graph-based retrieval retry with a modified prompt emphasizing behavioral indicators (e.g.,
focus on registry keys, process trees, and network patterns
). -
Add a prompt-versioning and reward-hacking-resistant judge harness. The improved system maintains two prompt templates (V1 and V2) and runs both on the same retrieval backend, then uses a judge with explicit score floors/ceilings (e.g., min 0, max 10 per criterion) and a step-by-step counting procedure that requires the judge to list each detection before scoring. This prevents the judge from giving inflated scores for verbose or irrelevant output, and it automatically selects the prompt version that yields higher median pyramid level across multiple runs.
-
Create a hybrid retrieval mode that fuses graph local and vector results. The improved system generates candidate plans from both backends, then uses a cross-encoder reranker (trained on detection-engineering rubrics) to select the best plan per query. This hybrid mode outperforms either alone by combining graph’s durable TTP extraction with vector’s recall for niche facts, and it falls back to pure graph local if vector results are all below L4.
-
Add a
durability metadata
layer to every generated detection. The improved system tags each detection with its Pyramid level, the source evidence (e.g.,from CTI PDF page 3, section 'Tactics'
), and a fragility score (0–1) computed from how many IOC types it depends on. This allows SOC teams to instantly filter for L5+ detections and to auto-generate aresilient-only
version of the plan with fragile detections removed.
What the improved AI system can do:
-
Generate threat hunting plans that retain >90% detection efficacy after full IOC rotation, even for APT-level adversaries, by prioritizing TTPs, registry keys, and process behaviors over IPs and hashes.
-
Automatically detect when a source report is too short for global graph search and switch to local search or vector retrieval, avoiding empty or hallucinated outputs.
-
Produce a single plan that combines the best of graph and vector retrieval, with a visible breakdown of which detections come from which backend and their respective durability scores.
-
Self-evaluate its own output using a domain-specific judge that cannot be gamed by verbose answers, and automatically re-run with a different prompt if the initial plan scores below a resilience threshold.
-
Provide SOC analysts with a
resilience report
alongside each plan, showing exactly which detections survive IOC rotation and which are fragile, enabling rapid triage and deployment of durable hunts.
Abstract
When a security researcher publishes a report on a cyberattack, detection engineers are supposed to turn it into working detection rules. In practice, most automated attempts at this only extract the simplest clues from the report --- bad IP addresses, domain names, and file hashes --- and turn them into block lists. This is a weak strategy, because attackers can change these simple clues within hours or days, so the resulting detections stop working almost as soon as they are deployed. Security teams describe this idea with the Pyramid of Pain. This project asks whether feeding a report into a knowledge-graph retrieval system, Microsoft GraphRAG, rather than a standard vector-similarity retrieval system (Naive RAG), produces detection plans that rely more on these durable, top-of-pyramid clues. Both systems are given the same report, the same generation instructions, and the same language model to write the final plan; only the retrieval step differs. In a detailed case study of one APT28 report, the GraphRAG plan kept firing at 100% of its detections after every IP address, domain, and file hash in the report was rotated, while the Naive RAG plan kept firing at only 29%. Repeating the comparison across nine real CTI reports from four vendors confirms the same pattern: GraphRAG plans consistently reach higher, harder-to-evade levels of the pyramid, even when the two systems end up close on total score. The results support treating knowledge-graph-aware retrieval as the architecturally correct foundation for automatically generating SOC-deployable hunting plans, while showing that the wording of the generation prompt matters almost as much as the retrieval back-end itself.
Sources
- SecureBERT 2.0: Advanced Language Model for Cybersecurity Intelligence
- SecureBERT: A Domain-Specific Language Model for Cybersecurity
- Looking Beyond IoCs: Automatically Extracting Attack Patterns from External CTI
- Evaluating LLM Generated Detection Rules in Cybersecurity
- CTI-REALM: Benchmark to Evaluate Agent Performance on Security Detection Rule Generation Capabilities
- CTINexus: Automatic Cyber Threat Intelligence Knowledge Graph Construction Using Large Language Models
- Cyber Defense Benchmark: Agentic Threat Hunting Evaluation for LLMs in SecOps
- From Local to Global: A Graph RAG Approach to Query-Focused Summarization
- Large Language Models for Security Operations Centers: A Comprehensive Survey
- Beyond RAG for Cyber Threat Intelligence: A Systematic Evaluation of Graph-Based and Agentic Retrieval
- Retrieval-Augmented Generation with Graphs (GraphRAG)
- CyberLLM-FINDS 2025: Instruction-Tuned Fine-tuning of Domain-Specific LLMs with Retrieval-Augmented Generation and Graph Integration for MITRE Evaluation
- Llama-3.1-FoundationAI-SecurityLLM-Base-8B Technical Report
- Benchmarking LLM-Assisted Blue Teaming via Standardized Threat Hunting
- FALCON: Transforming Cyber Threat Intelligence into Deployable IDS Rules with Self-Reflection
- Towards Automated Classification of Attackers' TTPs by combining NLP with ML Techniques
- LLMCloudHunter: Harnessing LLMs for Automated Extraction of Detection Rules from Cloud-Based CTI
- ThreatPilot: Attack-Driven Threat Intelligence Extraction
- AttacKG+:Boosting Attack Knowledge Graph Construction with Large Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs