Ranked by the Matcher: A Reproducibility Audit of Knowledge Graph Extraction from Threat Reports
cs.CR, cs.AI, cs.CL
Submitted: 2026-09-01
Updated: 2026-09-05
Comments: Code and configurations are available at https://github.com/sbhakim/CTIForge
Code: https://github.com/sbhakim/CTIForge
License: http://creativecommons.org/licenses/by/4.0/
The gist: Security teams and researchers choose knowledge-graph extraction tooling for threat reports on the strength of published triple-F1 scores, yet those scores depend on how predicted triples are matched
Terminology
Abstract
Security teams and researchers choose knowledge-graph extraction tooling for threat reports on the strength of published triple-F1 scores, yet those scores depend on how predicted triples are matched to gold annotations. We could reimplement the stated matching rule for only five of twelve inspected systems. Re-scoring ten system outputs on shared documents under eight protocols reverses eleven of forty-five pairwise orderings; one fixed prediction set spans 0.16-0.70 F1. On GRID's external 378-item calibration set, no mechanical matcher (lexical, embedding, or entailment) agrees with multi-reviewer adjudication above 71%, whereas an LLM judge reaches 86%. To separate component effects from matcher rewards, we build CTIForge, whose deterministic validation layer can vary while extraction is held byte-identical. Across seven tested deployment configurations, validation raises precision for all four hosted backbones and lowers it for all three offline backbones. Because backbone, decoding, and backend-specific prompting covary, this is a descriptive split rather than an isolated serving effect. It coincides with a roughly 2.8-fold increase in actions explicitly disputing entity type, consistent with hand-written rules encoding the conventions of the extractor against which they were developed. We release the pipeline, protocol suite, and per-triple audit records.
Sources
- CTINexus: Automatic Cyber Threat Intelligence Knowledge Graph Construction Using Large Language Models
- GRID: Graph Representation of Intelligence Data for Security Text Knowledge Graph Construction
- CTIBench: A Benchmark for Evaluating LLMs in Cyber Threat Intelligence
- SEvenLLM: Benchmarking, Eliciting, and Enhancing Abilities of Large Language Models in Cyber Threat Intelligence
- Evaluating Open-Source LLMs for Multi-Label ATT&CK Technique Classification on CTI Reports
- Beyond Single Reports: Evaluating Automated ATT&CK Technique Extraction in Multi-Report Campaign Settings
- Schema-Agnostic Knowledge Graph Construction via Hybrid Ontology Discovery for Cyber Threat Intelligence
- TACTIC-KG: Toward Small Agent Teams for Cyber Threat Intelligence Knowledge Graph Construction
- Context-aware Entity-Relation Extraction for Threat Intelligence Knowledge Graphs
- TTPrint: Evidence-Grounded TTP Extraction via Diverge-then-Converge Verification
- Better Later Than Sooner: Neuro-Symbolic Knowledge Graph Construction via Ontology-grounded Post-extraction Correction
- TRACE-CTI: Auditable Post-Extraction Governance of TTP Claims with Knowledge Graphs
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs