CTIFoundry: An Agent-Native Corpus Scaffold for Cyber Threat Intelligence
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "CTIFoundry: An Agent-Native Corpus Scaffold for Cyber Threat Intelligence".
Jane: Cyber threat intelligence (CTI) investigation is increasingly consumed by LLM agents,
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So, circling back to where we are, this paper introduces CTIFoundry as an agent-native corpus scaffold specifically for cyber threat intelligence <ref:2608.18613#pg0>. The central thesis is that the current way threat reports are packaged for retrieval-augmented generation results in opaque chunks behind a similarity search interface <ref:2608.18613#pg1>.
Jane: That packaging leads to vendor aliases not being resolved and official cross-references surviving only as text inside record blobs, which means derived claims lack span-level provenance <ref:2608.18613#pg0>. This is what the authors call the substrate bottleneck, arguing that this packaging is inherited from RAG methods <ref:2608.18613#pg1>.
Lu: Essentially, they claim that putting an agent on top of this flat corpus doesn't fix these fundamental issues; it just means iterating over an opaque substrate only re-retrieves what the indexing discarded <ref:2608.18613#pg1>.
Meng: So, the authors are proposing a solution that materializes the latent structure of this CTI corpus at build time, creating a deterministic ontology graph and other layers <ref:2608.18613#pg0>.
Lalam: CTIFoundry achieves this by materializing three core artifacts: an ontology graph over four authoritative knowledge bases—CVE, CWE, CAPEC, and ATT andCK—where official cross-references become typed, traversable edges <ref:2608.18613#pg0>.
Tom: That ontology graph is where they solve the challenge of resolving aliases and sibling confusion by making those official cross-references explicit as typed edges <ref:2608.18613#pg0>.
Jane: They also introduce a span-grounded report layer, which indexes chunks using canonical entities that carry vendor-specific alias sets and have grounding in the graph <ref:2608.18613#pg0>.
Lu: This layer ensures that every chunk has exact character-offset provenance, so an agent can verify claims by grounding them directly in specific text spans <ref:2608.18613#pg0>.
Meng: Furthermore, they include hybrid retrieval surfaces that fuse dense and lexical search under a steerable term filter to give the agent flexible access to the data <ref:2608.18613#pg0>.
Lalam: These components—the ontology graph, the report layer, and the hybrid surfaces—form what they define as an agent-native corpus scaffold <ref:2608.18613#pg0>.
Tom: So, the main point is that this scaffold allows agents to perform multi-step investigations with high accuracy because it materializes the latent structure instead of relying on a flat similarity search <ref:2608.18613#pg0>.
Jane: It matters because it directly addresses the failure modes where vendor aliases are never resolved and claims lack proper source separation in the original RAG packaging <ref:2608.18613#pg1>.
Lu: The core argument is that for domains with authoritative reference structure, this approach ensures an autonomous investigator can traverse and verify information anchored to those taxonomies <ref:2608.18613#pg0>.
Tom: It really reframes the research question from how to build a better agent to how to deliberately build a corpus that is perfectly structured for investigation <ref:2608.18613#pg0>.
Jane: And that’s what makes it important, Tom; it shifts the highest-leverage investment toward the data organization itself <ref:2608.18613#pg0>.
Conclusion: Tom: So, we've walked through the mechanics of CTIFoundry, looking at how it builds that structure at build time to solve the substrate bottleneck <ref:2608.18613#pg0>. Now we need to look at what this means for the bigger picture regarding threat intelligence and AI agents <ref:2608.18613#pg0>.
Jane: The authors of CTIFoundry are proposing a way to make CTI corpora inherently navigable for autonomous agents by materializing their latent structure beforehand <ref:2608.18613#pg0>. It’s about giving the agent a map that goes beyond simple text similarity <ref:2608.18613#pg0>.
Lu: The implication is that we can finally move toward truly autonomous CTI investigators who can traverse and verify information anchored to authoritative taxonomies, rather than just synthesizing from opaque text blobs <ref:2608.18613#pg0>.
Meng: For practical impact, this means the development process for intelligence data needs to incorporate this kind of structural mapping upfront, ensuring the AI tools it uses are built on a foundation that is inherently verifiable <ref:2608.18613#pg0>.
Lalam: Culturally, I see this as a significant step toward building better AI systems by focusing on the quality and organization of the intelligence sources themselves, rather than just chasing bigger model capabilities <ref:2608.18613#pg2>.
Tom: So, in simple terms, CTIFoundry suggests that the highest-leverage investment isn't necessarily a better agent, but rather a corpus deliberately built to be investigated by those agents <ref:2608.18613#pg0>.
Jane: Precisely; it’s about ensuring an autonomous investigator can follow and verify information anchored to established taxonomies like CVEs and ATT andCK <ref:2608.18613#pg0>.
Lu: This moves the focus onto structuring the data in a way that respects the authoritative relationships between different knowledge bases, making cross-referencing reliable <ref:2608.18613#pg2>.
Meng: From an engineering view, it means we're designing for structure first, and then building the agent tools to utilize that structure efficiently <ref:2608.18613#pg0>.
Lalam: It really points toward a future where the data scaffolding itself is as intelligent as the reasoning engine operating on top of it <ref:2608.18613#pg2>.
Tom: That's a powerful way to look at it, Jane; focusing on building the investigative environment rather than just optimizing the execution engine <ref:2608.18613#pg0>.
Virginia Tech University · Amazon.com Inc. · Northwestern University
cs.AI, cs.CR
Submitted: 2026-08-19
Updated: 2026-10-04
Comments: Preprint
Code: https://github.com/SWE-agent/mini-swe-agent
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 92/100
The gist: Cyber threat intelligence (CTI) investigation is increasingly consumed by LLM agents, but current Retrieval-Augmented Generation (RAG) methods fail because they package threat reports as opaque
Key concepts
- Ontology Graph
- This is a map of the CTI knowledge based on four authoritative sources (CVE, CWE, CAPEC, ATT&CK). It explicitly defines how these sources cross-reference each other as 'typed edges.' This structure helps agents resolve confusing aliases and understand the official relationships between different threat concepts clearly.
- Span-Grounded Report Layer
- This layer indexes CTI chunks by canonical entities, tracking exactly where in the text each chunk originates ('exact character-offset provenance'). This allows an agent to verify any piece of information by directly pointing to the specific text span where it was found, ensuring accuracy.
- Hybrid Retrieval Surfaces
- These surfaces combine two search methods: dense vector search (for meaning) and lexical search (for keywords), filtered by a steerable term. This helps agents find relevant information even when a perfect match isn't in the embedding space, addressing failures where rare tokens confuse standard similarity searches.
Terminology
Summary
Cyber threat intelligence (CTI) investigation is increasingly consumed by LLM agents, but current Retrieval-Augmented Generation (RAG) methods fail because they package threat reports as opaque chunks behind a similarity search interface. This paper introduces CTIFoundry, an agent-native corpus scaffold designed to solve this substrate bottleneck by materializing the latent structure of a CTI corpus at build time, thereby enabling agents to perform multi-step investigations with high accuracy.
How it works
CTIFoundry functions by creating a derived representation of the CTI corpus through a build-time pipeline that materializes three core artifacts: an ontology graph, a report layer, and hybrid retrieval surfaces. The ontology graph is a "deterministic ontology graph over four authoritative knowledge bases (CVE, CWE, CAPEC, ATT&CK) whose official cross-references become typed, traversable edges. This structure addresses the challenge of resolving aliases and sibling confusion by making official cross-references explicit as
typed edges."
The report layer handles the narrative layer of CTI. It is a span-grounded report layer whose canonical, alias-resolved cross-vendor entities index provenance-carrying chunks.
This ensures that every chunk has exact character-offset provenance,
allowing the agent to verify claims by grounding them in specific text spans.
The hybrid retrieval surfaces provide flexible access. These surfaces fuse dense and lexical search under a term filter, allowing the agent to steer its search based on feedback, which addresses the failure mode where a gold entry sharing rare discriminative tokens with the query yet sitting far from it in embedding space.
The Scaffold Components
The scaffold is defined as a derived representation through which an agent accesses the corpus. It consists of:
-
An ontology graph (G) over knowledge base records and their typed official cross-reference edges (C1).
-
A span-grounded report layer (X), where chunks are indexed by canonical entities with vendor-attributed alias sets and grounding in the graph (C2).
-
Hybrid retrieval surfaces (I), fusing dense and lexical search under a steerable term filter (C3).
The Query-Time Interface
At query time, this structure is exposed through seven typed tools and three procedural skills mounted on a stock, widely-used open-source agent harness. The seven tools cover non-overlapping capabilities such as resolve entity name/alias/id,
ontology neighbors traverse official cross-reference edges,
and search kb hybrid dense+BM25 search over one KB.
These are governed by rules like Structure before similarity: descriptions encode the substrate’s priority order (resolve names before querying, prefer an authoritative edge over text search whenever an identifier is known).
Complementing the tools are three procedural skills that encode investigation discipline. These skills are advice, not workflow engines,
and they bind to structure that exists. For example, a skill for entity linking dictates: do NOT search the TARGET taxonomy first,
because the description paraphrases one specific SOURCE entry; the authoritative cross-reference edge from that source entry gives the answer.
Empirical Validation
Evaluation on CTIConnect showed significant gains by swapping only the action surface. Swapping only the action surface lifts overall F1 from 0.610 to 0.829 for gpt-5.4 and from 0.470 to 0.745 for claude-haiku-4-5, demonstrating that structure makes the right investigation possible, procedure makes it reliable, and neither substitutes for the other.
The ablation study shows that procedural skills recover +0.062 and structure alone recovers +0.136, resulting in a super-additive gain of +0.219 when both are used together.
Conclusion
The paper concludes that the binding constraint is the substrate, not the agent.
CTIFoundry demonstrates that for any domain whose corpora carry authoritative reference structure, the highest-leverage investment in agent quality may not be a better agent at all, but a corpus deliberately built to be investigated.
This approach ensures that an autonomous investigator can traverse and verify information anchored to authoritative taxonomies.
The gist
CTIFoundry materializes the latent structure of a CTI corpus—a deterministic ontology graph, span-grounded report layer, and hybrid retrieval surfaces—at build time, exposing it at query time through seven typed tools and three procedural skills to solve the substrate bottleneck in agentic CTI investigation.
Key Procedural Skills
The three procedural skills are task-specific playbooks injected into the user turn:
-
Entity Linking (RCM, ATD, ESD): Focuses on mapping behavioral descriptions to taxonomy entries by prioritizing authoritative cross-reference edges over direct target taxonomy searches.
Improvements for AI systems
Based on the provided research paper, CTIFOUNDRY: AN AGENT-NATIVE CORPUS SCAFFOLD FOR CYBER THREAT INTELLIGENCE,
here are specific improvements for AI systems and what those improved systems can achieve:
The core improvement is shifting the bottleneck from model capability to corpus structure. Instead of relying on general retrieval-augmented generation (RAG) over flat, opaque chunks, the system should use CTIFoundry as an agent-native corpus scaffold.
Here are the specific improvements and capabilities:
-
Use a build-time pipeline that materializes a deterministic ontology graph over four authoritative knowledge bases (CVE, CWE, CAPEC, ATT&CK) into typed, traversable edges (C1).
-
Enforce a zero-fabrication invariant during the build process by checking all identifiers against pinned snapshots at write time. This guarantees that every identifier generated is verifiable and prevents fabricated entities from entering the system (RQ1).
-
Implement a span-grounded report layer (C2) that indexes vendor reports with exact character-offset provenance, resolving canonical cross-vendor entities based on deterministic signals before indexing chunks. This ensures derived claims are auditable back to the specific text and source span (C2).
-
Utilize hybrid dense+lexical retrieval surfaces (C3) that fuse embedding similarity search with term filtering, allowing the agent to steer the search pool size based on query feedback. This closes the failure mode where gold entries share rare tokens but are far from them in embedding space (C3).
-
Integrate a query-time interface consisting of seven typed tools (e.g., resolve entity name/alias/id, ontology neighbors traverse official cross-reference edges) and three procedural skills (e.g., resolve before searching, consult an alias set). This structure is exposed through a stock agent harness (C4).
-
Ensure the agent's action surface prioritizes authoritative structure over textual similarity when an identifier is known, explicitly teaching it to prefer recorded edges over semantic retrieval (RQ4/E.5).
The improved AI system, powered by CTIFoundry, can perform the following specific tasks:
-
Perform multi-step CTI investigations that require resolving aliases across different vendor reports (e.g.,
Identify the malware family of 'Hidden Cobra'
). The system will use its built-in ontology graph to follow official cross-references, moving beyond simple textual similarity to find the authoritative answer. -
Ground derived assertions in precise evidence by citing specific text spans from original reports (e.g.,
Which CTI report states that the malware uses CVE-2024-21412?
). This allows for verifiable attribution of claims, which is critical for defensive triage and attribution (C2). -
Accurately map behavioral descriptions in narrative prose to official attack patterns (e.g., "Map this description to an ATT&CK technique"). The system will use procedural skills that enforce a strict decomposition of the passage into atomic behaviors before searching the taxonomy, ensuring it selects the most specific, authoritative technique rather than a plausible but wrong one (H.3).
-
Precisely resolve and distinguish between near-miss entities with similar names across different vendors (e.g., distinguishing between two different ransomware families that share a similar name). The system uses deterministic logic (union-find over exact/span-verified evidence) to prevent state actors from collapsing into incorrect mega-clusters during entity linking (G.4).
-
Synthesize multi-document campaign profiles by systematically reading all listed reports, resolving every alias across vendors, and ordering events strictly by the dates stated in the text, attributing claims precisely to their source vendor (H.5). This allows for a complete, auditable timeline of a threat actor's activity.
-
Execute complex research workflows that scale efficiently: at 1.7× the question volume, accuracy is maintained because the structure provides reliable anchors for retrieval recall, and cost remains linear (E.3).
In summary, the improved AI system moves from being a sophisticated search engine to an agent that performs high-stakes CTI investigation by leveraging a pre-built, deterministic knowledge scaffold that enforces structural validity and procedural discipline.
Abstract
Cyber threat intelligence (CTI) is increasingly consumed not by human analysts but by LLM agents that compose multi-step investigations at query time. The harness side of this shift has matured rapidly, but the corpus side has not: threat reports and vulnerability databases are still packaged for retrieval-augmented generation, as opaque chunks behind an embedding index. We argue that this substrate, not model capability, is the bottleneck on agentic CTI investigation, and present CTIFoundry, an agent-native corpus scaffold. At build time, CTIFoundry materializes the latent structure of a CTI corpus: a deterministic ontology graph over four authoritative knowledge bases (CVE, CWE, CAPEC, ATT&CK) whose official cross-references become typed, traversable edges; a span-grounded report layer whose canonical, alias-resolved cross-vendor entities index provenance-carrying chunks; and hybrid dense+lexical retrieval surfaces. At query time this structure is exposed through seven typed tools and three procedural skills mounted on a stock, widely-used open-source agent harness. On the public CTIConnect benchmark, swapping only the action surface lifts the identically-harnessed agent from 0.610 to 0.829 overall F1 with gpt-5.4 and from 0.470 to 0.745 with claude-haiku-4-5: a small model on CTIFoundry surpasses a flagship model on the flat substrate. The scaffolded agent is simultaneously more accurate and more efficient: on both Claude models it answers with roughly half the tool calls per question. The ablation distills design principles for matching corpus scaffolding to data modality, in CTI and beyond. Build-time validation guarantees zero fabricated identifiers by construction, and the scaffold sustains 1,168 investigations end-to-end at about 2.6 cents each.
Sources
- HarnessForge: Joint Harness and Policy Evolution for Adaptive Agent Systems
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- From Local to Global: A Graph RAG Approach to Query-Focused Summarization
- LightRAG: Simple and Fast Retrieval-Augmented Generation
- Linking Threat Tactics, Techniques, and Patterns with Defensive Weaknesses, Vulnerabilities and Affected Platform Configurations for Cyber Hunting
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- Meta-Harness: End-to-End Optimization of Model Harnesses
- Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents
- Towards Accurate and Efficient Document Analytics with Large Language Models
- AutoHarness: improving LLM agents by automatically synthesizing a code harness
- MemGPT: Towards LLMs as Operating Systems
- Zep: A Temporal Knowledge Graph Architecture for Agent Memory
- Abacus: A Cost-Based Optimizer for Semantic Operator Systems
- Agentic Retrieval-Augmented Generation: A Survey on Agentic RAG
- R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning
- ZeroSearch: Incentivize the Search Capability of LLMs without Searching
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- Agent Workflow Memory
- Self-Harness: Harnesses That Improve Themselves
- Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection