TopoIntent: Compiling Security Intent into Executable, Compliance-Checked Network Topologies

arXiv:2608.13389 · cs.AI, cs.CR, cs.NI · Submitted 2026-08-13 · Read on arXiv

Xiaokang Qu, Jianliang Ma, Zao Fan, Tianshu Chu, Tianlong Fan, Linyuan Lü

University of Science and Technology of China · Topsec Technologies Group Inc.

cs.AI, cs.CR, cs.NI

Submitted: 2026-08-13

Updated: 2026-08-14

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: TopoIntent is a system that compiles natural-language security intent into executable, compliance-checked network topologies.

Terminology

Summary

TopoIntent is a system that compiles natural-language security intent into executable, compliance-checked network topologies. It addresses the problem that existing NetOps automation tools operate after the security topology design has been fixed, translating policies into configurations or interpreting existing diagrams, but provide limited support for producing a structured security topology from an underspecified natural-language request.

The system follows a compiler-like workflow with five sequential stages. Stage 1, Intent Analysis, transforms free-form natural language requirements into an initial topology template conforming to a schema contract (SCHEMACONTRACT v2.1), using Qwen2.5-72B-Instruct in a zero-shot setting. Stage 2, Template Retrieval, queries a curated Template Library of 22 reference security topologies using dense-vector search with BGE-M3 embeddings and Qdrant, retrieving the top-k most similar templates via cosine similarity. Stage 3, Intent Fusion, combines the user intent JSON with retrieved templates through two sequential LLM calls: Semantic Fusion (structural alignment without inference, preserving user-specified fields verbatim) and Intelligent Completion (filling missing fields according to the target CIS Controls Implementation Group level, producing separate IG2 and IG3 instances if no level is specified). Stage 4, CIS-Aware Compliance Check & Repair, checks the topology against 22 topology-visible CIS Controls v8.1.2 safeguards, labels each as satisfied, unsatisfied, or manual review, and performs iterative additive repair that only adds new zones, nodes, and links without modifying existing elements. Stage 5, Emulation-Based Validation, converts the topology into Mininet scripts with kernel-level iptables ACLs, testing reachability and allow/deny behavior, and feeds diagnostic failure categories back to Stage 4 for targeted repair.

The system is evaluated on a held-out set of 7 templates and 14 intents from finance and government scenarios excluded from the retrieval index. On this held-out set, additive repair raises first-pass topology-visible CIS satisfaction from 0.78/0.78 (IG2/IG3) to 1.00 in fewer than 1.5 rounds on average. In the end-to-end ablation, one feedback round raises the post-ACL policy pass rate from 0.78 to 0.88. The paper reports that Stage 1 achieves scenario accuracy of 0.98, zone type recall of 0.68, and node type recall of 0.53 on the reference set. Stage 2 single-vector retrieval achieves Top-1 accuracy of 0.75 and MRR of 0.84, outperforming multi-field variants. Stage 5 validation shows 100% node-link reachability, with post-ACL pass rates improving from 0.87 to 0.86 (IG2) and 0.87 to 0.89 (IG3) after one feedback round.

The paper makes four main contributions: a security-design compilation pipeline that maps natural-language requirements into a schema-constrained topology representation; template-grounded topology synthesis using dense retrieval from a curated multi-industry template library; CIS-guided structural checking and repair using a scoped set of topology-visible safeguards with additive schema-preserving edits; and executable validation with diagnostic feedback that separates topology failures from ACL-enforcement errors.

The authors note limitations including that TopoIntent does not certify full organizational CIS compliance (only topology-visible safeguards), LLM judgments can make errors, Mininet emulation does not model production deployment details, the dataset is modest in scale, additive repair may add redundant devices, and feedback convergence is not guaranteed.

Improvements for AI systems

Improvements to AI Systems:

  1. Schema-Constrained Generative Pipeline: Implement a multi-stage compiler architecture (analysis → retrieval → fusion → compliance check → validation) that enforces a formal schema contract at each stage, preventing hallucinated or malformed outputs. The improved system can generate structured, executable artifacts (e.g., network topologies, infrastructure code) directly from free-form natural language, with guaranteed schema compliance.

  2. Template-Grounded Retrieval-Augmented Generation (RAG): Use dense-vector retrieval (e.g., BGE-M3 embeddings) over a curated library of domain-specific reference templates to ground LLM synthesis. The improved system can retrieve the most relevant prior solutions (e.g., security architectures, cloud designs) and fuse them with user intent, reducing novel but invalid outputs and improving domain accuracy (e.g., Top-1 retrieval accuracy 0.75, MRR 0.84).

  3. Additive, Schema-Preserving Repair Loop: Integrate an iterative repair mechanism that only adds new elements (zones, nodes, links) without modifying existing ones, guided by a compliance checker (e.g., CIS Controls). The improved system can self-correct its own outputs to meet external standards (raising satisfaction from 0.78 to 1.00 in <1.5 rounds) while preserving user-specified constraints verbatim.

  4. Two-Stage Fusion with Explicit Semantic and Completion Calls: Separate structural alignment (no inference) from missing-field completion (inference) into distinct LLM calls. The improved system can merge retrieved templates with user intent without overwriting user choices, then intelligently fill gaps based on target compliance levels (e.g., IG2 vs. IG3), producing multiple compliant variants when unspecified.

  5. Executable Validation with Diagnostic Feedback: Convert generated artifacts into emulation scripts (e.g., Mininet with kernel-level ACLs) to test functional behavior (reachability, allow/deny). The improved system can distinguish topology-level failures from policy-enforcement errors and feed categorized diagnostics back into the repair stage, improving end-to-end pass rates (e.g., from 0.78 to 0.88 after one feedback round).

  6. Scoped Compliance Checking with Manual-Review Triage: Label each requirement as satisfied, unsatisfied, or manual-review, rather than binary pass/fail. The improved system can prioritize automated repair only for topology-visible, machine-checkable constraints, while flagging ambiguous or non-verifiable items for human oversight, reducing false confidence.

  7. Zero-Shot Intent Parsing with High Recall on Structural Elements: Use a large instruction-tuned LLM (e.g., Qwen2.5-72B) for initial intent extraction, but combine it with retrieval and fusion to compensate for lower recall on fine-grained types (e.g., zone recall 0.68, node recall 0.53). The improved system can achieve high scenario accuracy (0.98) while using templates to fill in missing structural details.

  8. Iterative Convergence with Bounded Rounds: Implement a feedback loop that tracks the number of repair rounds and stops after a threshold (e.g., <1.5 rounds average). The improved system can guarantee bounded iteration for cost control, while still improving compliance and functional correctness without infinite loops.

What the Improved AI System Can Do:

  • Accept a natural-language request (e.g., Build a segmented network for a finance app with DMZ and internal DB) and output a fully specified, schema-valid topology with zones, nodes, links, and ACLs.

  • Automatically retrieve and adapt the most similar proven security architecture from a library, preserving user-specific details.

  • Self-check against industry compliance frameworks (e.g., CIS Controls) and repair its own design additively until all topology-visible safeguards are met.

  • Generate executable network emulation scripts (e.g., Mininet) and run them to verify that the intended allow/deny policies actually work, then fix any failures.

  • Produce multiple compliance-level variants (e.g., IG2 and IG3) when the user does not specify a target, allowing trade-off analysis.

  • Clearly separate automated fixes from items requiring human judgment, avoiding over-automation on ambiguous requirements.

  • Operate with a predictable number of repair iterations, making it suitable for real-time or cost-sensitive deployment.

Sources

Related papers