EvoHarnessBench: Can Your Agents Keep Pace with an Evolving Harness?
cs.MA, cs.CL
Submitted: 2026-09-03
Updated: 2026-09-26
Comments: https://mas-orchestra.salesforceresearch.ai/evoharness/
Code: https://github.com/openai/skills
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: Modern LLM-based agents operate through a harness of tools, reusable skills, and specialist agents that shapes what they observe and what they can do.
Terminology
Abstract
Modern LLM-based agents operate through a harness of tools, reusable skills, and specialist agents that shapes what they observe and what they can do. In practice, this harness continually evolves as new capabilities are added. We introduce EVOHARNESSBENCH, a benchmark for evaluating agents under controlled harness evolution across three axes (tools, skills, and agents). Unlike existing continual-learning benchmarks for agents, which typically place non-stationarity (i.e., what changes over time) in the task stream while keeping the harness fixed, EVOHARNESSBENCH places non-stationarity in the externally supplied harness itself. It contains 17 multi-stage harness streams constructed deterministically from verifier-based benchmarks, comprising 802 tasks, 520 tools, 42 skills, and 62 agents. We evaluate two complementary settings corresponding to the central challenges of harness evolution: deployment evaluation, which isolates retention of previously accessible competence as the harness expands, and self-evolving adaptation evaluation, which tests whether accumulated experience remains useful as new capabilities are introduced. Our results reveal three persistent gaps. First, harness expansion alone can degrade performance on previously solved tasks, producing harness-induced forgetting. Second, gains from self-evolving adaptation remain inconsistent across stages of harness evolution, capability axes, and environments. Third, retention and adaptation can pull in different directions: preserving earlier competence does not necessarily improve adaptation to newly introduced capabilities, and vice versa. These results establish harness evolution as a distinct challenge for building agents that can keep pace with an evolving harness while preserving previously effective behavior.
Sources
- Continual Learning Bench: Evaluating Frontier AI Systems in Real-World Stateful Environments
- MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
- SkillCraft: Can LLM Agents Learn to Use Tools Skillfully?
- ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
- MemToolAgent: Leveraging Memory for Tool Using Agents Based on Environment and User Feedback
- Harness Engineering for Agentic AI Coding Tools: An Exploratory Study
- From Model Scaling to System Scaling: Scaling the Harness in Agentic AI
- LEGOMem: Modular Procedural Memory for Multi-agent LLM Systems for Workflow Automation
- Automated Design of Agentic Systems
- A Survey of Frontiers in LLM Reasoning: Inference Scaling, Learning to Reason, and Agentic Systems
- API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
- SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
- AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration
- How Well Do Agentic Skills Work in the Wild: Benchmarking LLM Skill Usage in Realistic Settings
- Stop Hand-Holding Your Coding Agent: Engineering the Loops that Replace Step-by-Step Prompting
- EnterpriseOps-Gym: Environments and Evaluations for Stateful Agentic Planning and Tool Use in Enterprise Settings
- Decentralized Multi-Agent Systems with Shared Context
- AgentCL: Toward Rigorous Evaluation of Continual Learning in Language Agents
- CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark
- PaperBench: Evaluating AI's Ability to Replicate AI Research
Related papers
- Highway Congestion Reduction through Reinforcement Learning Based Eulerian Headway Control
- You Only Align Once: Propagating Cooperative Behaviors in Multi-Agent Systems through Seed Agents
- Deny Without Disabling: Authorization-Paired Evaluation and Control for Multi-Agent Systems
- MA-SAPO: Multi-Agent Reasoning for Score-Aware Prompt Optimization
- PeroMAS: A Multi-agent System of Perovskite Material Discovery
- StitchCUDA: An Automated Multi-Agents End-to-End GPU Programing Framework with Rubric-based Agentic Reinforcement Learning