Rethinking MCP Security: A Large-Scale Study of Runtime MCP Servers and Security Scanner Reliability
Pei Chen, Baichao An, Mengying Wu, Binwang Wan, Geng Hong, Jinsong Chen, Xudong Pan, Jiarun Dai, Min Yang
cs.CR
Submitted: 2026-07-13
Comments: 18 pages, 11 figures, and 10 tables. This article substantially extends the preliminary 3-page MCPZoo dataset release arXiv:2512.15144. Includes appendices
Code: https://github.com/aira-security/mcp-armor
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: The Model Context Protocol (MCP) has rapidly established itself as a standard interface for enabling LLM-based agents to interact with external tools and services.
Terminology
Abstract
The Model Context Protocol (MCP) has rapidly established itself as a standard interface for enabling LLM-based agents to interact with external tools and services. As MCP servers are increasingly entrusted with security-sensitive operations, understanding their real-world risks has become critical. In practice, due to the absence of large-scale runtime MCP servers, such understanding largely relies on security scanners applied to a small number of cases, yet the reliability of these assessments remains unclear. In this study, we revisit how MCP security is measured. We present MCPZoo, the largest collection of MCP servers for dynamic analysis to date. MCPZoo is constructed through a multi-agent framework for transforming in-the-wild static repositories into dynamic services. The framework emulates how human experts build, diagnose, and iteratively repair deployment and runtime defects by combining environment inference with feedback-driven refinement. To ensure practical interactivity at runtime, the servers are validated via real protocol interactions. As a result, MCPZoo contains 64,611 unique MCP servers (113,927 in total), with more than 37,288 supporting dynamic analysis. Leveraging MCPZoo, we conduct the first ecosystem-scale measurement of MCP servers and the scanners that analyze them. While existing scanners report that 96.89% of servers are risky, we find that these signals are unreliable. In particular, manual validation shows that less than 50% of sampled alerts are true positives, and scanner outputs exhibit clear inconsistency across scanners. Overall, MCPZoo enables large-scale, reproducible measurement of MCP server security and exposes limitations of current scanning practices. We further release a public query interface to support practical risk assessment of MCP servers.
Sources
- Servant, Stalker, Predator: How An Honest, Helpful, And Harmless (3H) Agent Unlocks Adversarial Skills
- MCP Safety Audit: LLMs with the Model Context Protocol Allow Major Security Exploits
- MCPThreatHive: Automated Threat Intelligence for Model Context Protocol Ecosystems
- Quantifying Conversation Drift in MCP via Latent Polytope
- Beyond the Protocol: Unveiling Attack Vectors in the Model Context Protocol (MCP) Ecosystem
- How are AI agents used? Evidence from 177,000 MCP tools
- Who Grants the Agent Power? Defending Against Instruction Injection via Task-Centric Access Control
- Vexed by VEX tools: Consistency evaluation of container vulnerability scanners
- Trivial Trojans: How Minimal MCP Servers Enable Cross-Tool Exfiltration of Sensitive Data
- AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
- A Measurement Study of Model Context Protocol Ecosystem
- MCPXKIT: The Unified Toolkit for Analyzing Model Context Protocol Security
- Model Context Protocol (MCP) at First Glance: Studying the Security and Maintainability of MCP Servers
- Model Context Protocol (MCP): Landscape, Security Threats, and Future Research Directions
- MCP-in-SoS: Risk assessment framework for open-source MCP servers
- A First Look at the Security Issues in the Model Context Protocol Ecosystem
- Model Context Protocol for Vision Systems: Audit, Security, and Protocol Extensions
- MCPTox: A Benchmark for Tool Poisoning Attack on Real-World MCP Servers
- MCPZoo: A Large-Scale Dataset of Runnable Model Context Protocol Servers for AI Agent
- MCPSecBench: A Systematic Security Benchmark and Playground for Testing Model Context Protocols
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs