PeakBench: Benchmarking Resource-Aware Tool Invocation in LLM Agents
cs.AI, cs.SE
Submitted: 2026-08-25
Updated: 2026-08-25
Code: https://github.com/Czzzk/Staggering-the-Peaks
Terminology
Sources
- MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers
- MCPToolBench++: A Large Scale AI Agent Model Context Protocol MCP Tool Use Benchmark
- R-ConstraintBench: Evaluating LLMs on NP-Complete Scheduling
- LLM+P: Empowering Large Language Models with Optimal Planning Proficiency
- MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers
- AIOS: LLM Agent Operating System
- LiveMCPBench: Can Agents Navigate an Ocean of MCP Tools?
- COMPASS: Benchmarking Constrained Optimization in LLM Agents
- OpenAI GPT-5 System Card
- Kimi K2.5: Visual Agentic Intelligence
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
- MCPMark: A Benchmark for Stress-Testing Realistic and Comprehensive MCP Use
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- ReAct: Synergizing Reasoning and Acting in Language Models
- CCTU: A Benchmark for Tool Use under Complex Constraints
- GLM-5: from Vibe Coding to Agentic Engineering
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection