MCPAgentBench: A Real-world Task Benchmark for Evaluating LLM Agent MCP Tool Use
cs.AI
Submitted: 2025-12-31
Updated: 2026-09-15
Comments: Accepted to REALM, EMNLP 2026
Code: https://github.com/Brunestuder/MCPAgentBench
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large Language Models (LLMs) are increasingly serving as autonomous agents, and their utilization of external tools via the Model Context Protocol (MCP) is considered a future trend.
Terminology
Abstract
Large Language Models (LLMs) are increasingly serving as autonomous agents, and their utilization of external tools via the Model Context Protocol (MCP) is considered a future trend. Current MCP evaluation sets suffer from issues such as reliance on external MCP services and a lack of difficulty awareness. To address these limitations, we propose MCPAgentBench, a benchmark based on real-world MCP definitions designed to evaluate the tool-use capabilities of agents. We construct a dataset containing authentic tasks and simulated MCP tools. The evaluation employs a dynamic sandbox environment that presents agents with candidate tool lists containing distractors, thereby testing their tool selection and discrimination abilities. Furthermore, we introduce comprehensive metrics to measure both task completion rates and execution efficiency. Experiments conducted on state-of-the-art LLMs reveal significant performance differences in handling complex, multi-step tool invocations. All code is open-source at https://github.com/Brunestuder/MCPAgentBench.
Sources
- MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework
- Model Context Protocol (MCP): Landscape, Security Threats, and Future Research Directions
- Generative Agents: Interactive Simulacra of Human Behavior
- A Survey on Large Language Model based Autonomous Agents
- AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
- The Rise and Potential of Large Language Model Based Agents: A Survey
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection