CTIFoundry: An Agent-Native Corpus Scaffold for Cyber Threat Intelligence

summary

Video file (mp4)

The gist

Cyber threat intelligence (CTI) investigation is increasingly consumed by LLM agents, but current Retrieval-Augmented Generation (RAG) methods fail because they package threat reports as opaque

In short

CTIFoundry solves a problem where LLM agents struggle with complex Cyber Threat Intelligence (CTI) reports because they are presented as opaque text chunks. It creates an agent-native scaffold by materializing the corpus's underlying structure—an ontology graph, report layer, and hybrid search—at build time. This allows agents to perform accurate, multi-step investigations by using structured tools instead of relying solely on similarity search.

Key concepts

Ontology Graph
This is a map of the CTI knowledge based on four authoritative sources (CVE, CWE, CAPEC, ATT&CK). It explicitly defines how these sources cross-reference each other as 'typed edges.' This structure helps agents resolve confusing aliases and understand the official relationships between different threat concepts clearly.
Span-Grounded Report Layer
This layer indexes CTI chunks by canonical entities, tracking exactly where in the text each chunk originates ('exact character-offset provenance'). This allows an agent to verify any piece of information by directly pointing to the specific text span where it was found, ensuring accuracy.
Hybrid Retrieval Surfaces
These surfaces combine two search methods: dense vector search (for meaning) and lexical search (for keywords), filtered by a steerable term. This helps agents find relevant information even when a perfect match isn't in the embedding space, addressing failures where rare tokens confuse standard similarity searches.

Terminology used across episodes

This episode discusses

The paper

CTIFoundry: An Agent-Native Corpus Scaffold for Cyber Threat Intelligence · Read on arXiv

Virginia Tech University · Amazon.com Inc. · Northwestern University

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "CTIFoundry: An Agent-Native Corpus Scaffold for Cyber Threat Intelligence".

Jane: Cyber threat intelligence (CTI) investigation is increasingly consumed by LLM agents,

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So, circling back to where we are, this paper introduces CTIFoundry as an agent-native corpus scaffold specifically for cyber threat intelligence <ref:2608.18613#pg0>. The central thesis is that the current way threat reports are packaged for retrieval-augmented generation results in opaque chunks behind a similarity search interface <ref:2608.18613#pg1>.

Jane: That packaging leads to vendor aliases not being resolved and official cross-references surviving only as text inside record blobs, which means derived claims lack span-level provenance <ref:2608.18613#pg0>. This is what the authors call the substrate bottleneck, arguing that this packaging is inherited from RAG methods <ref:2608.18613#pg1>.

Lu: Essentially, they claim that putting an agent on top of this flat corpus doesn't fix these fundamental issues; it just means iterating over an opaque substrate only re-retrieves what the indexing discarded <ref:2608.18613#pg1>.

Meng: So, the authors are proposing a solution that materializes the latent structure of this CTI corpus at build time, creating a deterministic ontology graph and other layers <ref:2608.18613#pg0>.

Lalam: CTIFoundry achieves this by materializing three core artifacts: an ontology graph over four authoritative knowledge bases—CVE, CWE, CAPEC, and ATT andCK—where official cross-references become typed, traversable edges <ref:2608.18613#pg0>.

Tom: That ontology graph is where they solve the challenge of resolving aliases and sibling confusion by making those official cross-references explicit as typed edges <ref:2608.18613#pg0>.

Jane: They also introduce a span-grounded report layer, which indexes chunks using canonical entities that carry vendor-specific alias sets and have grounding in the graph <ref:2608.18613#pg0>.

Lu: This layer ensures that every chunk has exact character-offset provenance, so an agent can verify claims by grounding them directly in specific text spans <ref:2608.18613#pg0>.

Meng: Furthermore, they include hybrid retrieval surfaces that fuse dense and lexical search under a steerable term filter to give the agent flexible access to the data <ref:2608.18613#pg0>.

Lalam: These components—the ontology graph, the report layer, and the hybrid surfaces—form what they define as an agent-native corpus scaffold <ref:2608.18613#pg0>.

Tom: So, the main point is that this scaffold allows agents to perform multi-step investigations with high accuracy because it materializes the latent structure instead of relying on a flat similarity search <ref:2608.18613#pg0>.

Jane: It matters because it directly addresses the failure modes where vendor aliases are never resolved and claims lack proper source separation in the original RAG packaging <ref:2608.18613#pg1>.

Lu: The core argument is that for domains with authoritative reference structure, this approach ensures an autonomous investigator can traverse and verify information anchored to those taxonomies <ref:2608.18613#pg0>.

Tom: It really reframes the research question from how to build a better agent to how to deliberately build a corpus that is perfectly structured for investigation <ref:2608.18613#pg0>.

Jane: And that’s what makes it important, Tom; it shifts the highest-leverage investment toward the data organization itself <ref:2608.18613#pg0>.

Conclusion: Tom: So, we've walked through the mechanics of CTIFoundry, looking at how it builds that structure at build time to solve the substrate bottleneck <ref:2608.18613#pg0>. Now we need to look at what this means for the bigger picture regarding threat intelligence and AI agents <ref:2608.18613#pg0>.

Jane: The authors of CTIFoundry are proposing a way to make CTI corpora inherently navigable for autonomous agents by materializing their latent structure beforehand <ref:2608.18613#pg0>. It’s about giving the agent a map that goes beyond simple text similarity <ref:2608.18613#pg0>.

Lu: The implication is that we can finally move toward truly autonomous CTI investigators who can traverse and verify information anchored to authoritative taxonomies, rather than just synthesizing from opaque text blobs <ref:2608.18613#pg0>.

Meng: For practical impact, this means the development process for intelligence data needs to incorporate this kind of structural mapping upfront, ensuring the AI tools it uses are built on a foundation that is inherently verifiable <ref:2608.18613#pg0>.

Lalam: Culturally, I see this as a significant step toward building better AI systems by focusing on the quality and organization of the intelligence sources themselves, rather than just chasing bigger model capabilities <ref:2608.18613#pg2>.

Tom: So, in simple terms, CTIFoundry suggests that the highest-leverage investment isn't necessarily a better agent, but rather a corpus deliberately built to be investigated by those agents <ref:2608.18613#pg0>.

Jane: Precisely; it’s about ensuring an autonomous investigator can follow and verify information anchored to established taxonomies like CVEs and ATT andCK <ref:2608.18613#pg0>.

Lu: This moves the focus onto structuring the data in a way that respects the authoritative relationships between different knowledge bases, making cross-referencing reliable <ref:2608.18613#pg2>.

Meng: From an engineering view, it means we're designing for structure first, and then building the agent tools to utilize that structure efficiently <ref:2608.18613#pg0>.

Lalam: It really points toward a future where the data scaffolding itself is as intelligent as the reasoning engine operating on top of it <ref:2608.18613#pg2>.

Tom: That's a powerful way to look at it, Jane; focusing on building the investigative environment rather than just optimizing the execution engine <ref:2608.18613#pg0>.

More episodes

← Home