FlowGuard: From Signals to Evidence for MCP Security Detection
Baichao An, Pei Chen, Geng Hong, Yueyue Chen, Mengying Wu
cs.CR
Submitted: 2026-07-16
Code: https://github.com/antgroup/MCPScan
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: The Model Context Protocol (MCP) enables LLM agents to interact with external tools through metadata exchange, tool invocation, and response consumption.
Terminology
Abstract
The Model Context Protocol (MCP) enables LLM agents to interact with external tools through metadata exchange, tool invocation, and response consumption. Existing MCP security scanners primarily reason about suspicious semantic signals rather than real execution behaviors, which can lead to unreliable risk assessment. For example, credential-like strings may simply be placeholders rather than actual leakage. This gap requires runtime evidence for execution-related risks and careful semantic analysis for risks carried in metadata or returned content. We present FlowGuard, an evidence-grounded MCP security detection system. FlowGuard combines semantic risk triage, recon-guided payload narrowing, schema-valid probe generation, evidence adjudication, and history-guided refinement. It verifies execution-related risks through runtime evidence and detects semantic risks in tool metadata and returned content. We evaluate FlowGuard on an executable benchmark containing 1,880 MCP cases across five vulnerability categories. FlowGuard achieves F1 scores of 0.879 and 0.942 on the execution-related Command Injection and File System Access categories, respectively. Compared with existing dynamic scanners, FlowGuard reduces end-to-end latency by up to 2.23x. In the real-world evaluation, FlowGuard reports 523 findings across 326 servers. These results show that evidence-grounded detection can assess both execution-related and semantic risks in MCP interactions.
Sources
- MCPSecBench: A Systematic Security Benchmark and Playground for Testing Model Context Protocols
- MCP-SafetyBench: A Benchmark for Safety Evaluation of Large Language Models with Real-World MCP Servers
- MCPTox: A Benchmark for Tool Poisoning Attack on Real-World MCP Servers
- MCPZoo: A Large-Scale Dataset of Runnable Model Context Protocol Servers for AI Agent
- MCPXKIT: The Unified Toolkit for Analyzing Model Context Protocol Security
- MPMA: Preference Manipulation Attack Against Model Context Protocol
- MCP-ITP: An Automated Framework for Implicit Tool Poisoning in MCP
- Attractive Metadata Attack: Inducing LLM Agents to Invoke Malicious Tools
- Breaking the Protocol: Security Analysis of the Model Context Protocol Specification and Prompt Injection Vulnerabilities in Tool-Integrated LLM Agents
- Compatibility at a Cost: Systematic Discovery and Exploitation of MCP Clause-Compliance Vulnerabilities
- Don't believe everything you read: Understanding and Measuring MCP Behavior under Misleading Tool Descriptions
- A First Look at the Security Issues in the Model Context Protocol Ecosystem
- Automatic Red Teaming LLM-based Agents with Model Context Protocol Tools
- T-MAP: Red-Teaming LLM Agents with Trajectory-aware Evolutionary Search
- MindGuard: Intrinsic Decision Inspection for Securing LLM Agents Against Metadata Poisoning
- MCP-Guard: A Multi-Stage Defense-in-Depth Framework for Securing Model Context Protocol in Agentic AI
- Evaluating a Layered Prompt-Injection Defence for the Model Context Protocol: A Record-Level Audit of Decision Conventions, Corpus Provenance and Reproducibility
- MCP Safety Audit: LLMs with the Model Context Protocol Allow Major Security Exploits
- Auditing MCP Servers for Over-Privileged Tool Capabilities
- MCP-SandboxScan: WASM-based Secure Execution and Runtime Analysis for MCP Tools
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs