DynamicMCPBench: A Trace-Grounded, Effect-Scored Benchmark for LLM Agents over Live MCP Servers
cs.AI
Submitted: 2026-07-10
Updated: 2026-08-31
Comments: Accepted to EMNLP 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language model (LLM) agents are increasingly deployed over Model Context Protocol (MCP) servers, yet the benchmarks used to evaluate them score the final answer or a fixed "ground-truth" list
Terminology
Abstract
Large language model (LLM) agents are increasingly deployed over Model Context Protocol (MCP) servers, yet the benchmarks used to evaluate them score the final answer or a fixed "ground-truth" list of tools, both of which are fragile once the underlying data is live and stateful. We present DynamicMCPBench, a reusable framework rather than a fixed dataset. A practitioner can run it on their own MCP servers to test models on their own tasks, or let it collect servers automatically to measure a model's general ability to solve agentic tasks. Given the servers and any set of models, it generates realistic goals, pursues each one live to record a successful trajectory, distills that trajectory into path-agnostic effect checkpoints, and scores an agent on whether it reproduces those effects, never on the final answer. To show what the framework reveals, we run it at scale: 24 models over 121 servers and 750 tasks spread evenly over 15 task categories (50 each), where each category targets a distinct tool-use challenge of the generated questions. Each task is scored by pass 3: it counts as solved only if all three independent attempts succeed. Even the strongest agents solve only about half of the tasks, 31% of tasks are solved by no model at all, and accuracy collapses as the required tool chain grows longer (from 39% on the shortest chains to 13% on the longest). A human validation study confirms the automatic scoring is reliable (chance-corrected agreement of 0.76). DynamicMCPBench thus turns benchmark construction into something practitioners can rerun on their own servers and models, while exposing a consistent inability of current agents to handle long, multi-step agentic tasks.
Sources
- MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- Beyond Itinerary Planning-A Real-World Benchmark for Multi-Turn and Tool-Using Travel Tasks
- Toward Scalable Verifiable Reward: Proxy State-Based Evaluation for Multi-turn Tool-Calling LLM Agents
- MCPToolBench++: A Large Scale AI Agent Model Context Protocol MCP Tool Use Benchmark
- RAG-MCP: Mitigating Prompt Bloat in LLM Tool Selection via Retrieval-Augmented Generation
- MCP-RADAR: A Multi-Dimensional Benchmark for Evaluating Tool Use Capabilities in Large Language Models
- OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use Agents
- GLM-5: from Vibe Coding to Agentic Engineering
- Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents
- MiniMax Sparse Attention
- MCPVerse: An Expansive, Real-World Benchmark for Agentic Tool Use
- Hammer: Robust Function-Calling for On-Device Language Models via Function Masking
- Ministral 3
- ToolACE: Winning the Points of LLM Function Calling
- Model Context Protocol (MCP) Tool Descriptions Are Smelly! Towards Improving AI Agent Efficiency with Augmented MCP Tool Descriptions
- MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers
- OpenAI GPT-5 System Card
- LiveMCPBench: Can Agents Navigate an Ocean of MCP Tools?
- Scaling Granite Code Models to 128K Context
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection