UniACE: A Unified Framework for Evaluating LLM Agentic Capabilities
cs.AI
Submitted: 2026-05-27
Updated: 2026-09-01
Code: https://github.com/whfeLingYu/A-Unified-Framework-for-the-Evaluation-of-LLM-Agentic-Capabilities
License: http://creativecommons.org/licenses/by/4.0/
The gist: Agent benchmarks are increasingly used to compare large language models (LLMs) across domains, yet a reported score reflects a complete model--harness--environment configuration rather than the model
Terminology
Abstract
Agent benchmarks are increasingly used to compare large language models (LLMs) across domains, yet a reported score reflects a complete model--harness--environment configuration rather than the model alone. Benchmark packages couple native tasks with specific prompts, tool protocols, orchestration logic, and sometimes dynamic external resources, making cross-benchmark comparisons sensitive to implementation and resource conditions. We present UniACE, a unified framework for model-centric evaluation under an explicit, common execution condition. UniACE represents each benchmark as an instruction--tool--environment triplet, executes LLMs through a shared, task-agnostic harness in isolated per-task runtimes, and preserves native success criteria. For tasks that rely on dynamic resources, an optional offline mode replaces live access with fixed, pre-collected snapshots. Its evaluation protocol further standardizes efficiency measurement, execution records, and trace-based failure attribution. We migrate 7 benchmarks spanning 24 domains and evaluate 15 models in more than 400K rollouts consuming 5B tokens. Comparisons with source implementations show large bidirectional score changes and model-ranking reversals, while matched online and offline runs reveal substantial sensitivity to accessible evidence and its representation. Under the shared UniACE configuration, efficiency and failure profiles expose task-dependent model behaviors hidden by task-success scores alone. These findings motivate reporting agent benchmark outcomes as properties of an explicit evaluation configuration, enabling more interpretable and reproducible cross-benchmark comparisons. Codes and benchmarks at are available at https://github.com/whfeLingYu/A-Unified-Framework-for-the-Evaluation-of-LLM-Agentic-Capabilities, https://huggingface.co/datasets/whfeLingYu/Unified Agent Framework.
Sources
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- DeepSeek-V3 Technical Report
- Agentic Reinforced Policy Optimization
- Large Language Model Agent: A Survey on Methodology, Applications and Challenges
- AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
- OpenAI GPT-5 System Card
- MedAgents: Large Language Models as Collaborators for Zero-shot Medical Reasoning
- Gemini: A Family of Highly Capable Multimodal Models
- BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
- AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
- Qwen3 Technical Report
- Survey on Evaluation of LLM-based Agents
- Agent-SafetyBench: Evaluating the Safety of LLM Agents
- "LLM Agent Performance" Is Not a Single Evaluation Target
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection