Beyond Endpoint Performance: Process-Level Evaluation of Self-Evolving Agents
cs.AI
Submitted: 2026-09-21
Updated: 2026-09-21
Code: https://github.com/HQ-Lin/EvoPathBench
License: http://creativecommons.org/licenses/by/4.0/
The gist: Self-evolving agents convert interaction feedback into persistent artifacts, such as memories or skills, which in turn guide subsequent decisions.
Terminology
Abstract
Self-evolving agents convert interaction feedback into persistent artifacts, such as memories or skills, which in turn guide subsequent decisions. As these artifacts are iteratively updated throughout an experience stream, the capabilities they support may evolve. Consequently, endpoint performance alone offers an incomplete view of self-evolution. Process-level evaluation is therefore essential to identify when a target capability emerges and whether later updates strengthen, preserve, or weaken it. Motivated by this, we propose EvoPathBench, a benchmark that tracks individual capabilities during artifact-level self-evolution. EvoPathBench fixes the base model, tools, freezes evolving artifacts at successive checkpoints, and evaluates the target capability on held-out episodes. This benchmark evaluates agent self-evolution using public trading data and calibrated trajectories. It tests three capabilities: generalization to unseen tasks, retention after unrelated learning, and rule adaptation to new evidence. Experimental results show that gains on similar unseen tasks often weaken under distribution shift, retention losses are concentrated in a minority of evolution paths, and no method achieves reliable rule adaptation. Moreover, while self-evolution enables agents to generate candidate artifacts with substantial held-out gains, the selected updates consistently fall short of realizing this potential. Together, these findings establish capability-level process evaluation as a foundation for analyzing self-evolution, identifying candidate evaluation and selection as key targets for improvement.
Sources
- SkillZip Pro: Execution-Aware Dynamic Compression of Progressively Loaded Skills for Self-Evolving Agents
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- FinEvo-Bench: A Longitudinal Benchmark for Self-Evolving Agents in Professional Financial Workflows
- LiveMathematicianBench: A Live Benchmark for Research-Level Mathematical Reasoning with Proof Sketches
- SEA-Eval: A Benchmark for Evaluating Self-Evolving Agents Beyond Episodic Assessment
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- Rethinking Self-Evolution: A Constrained Exploration-Exploitation Process for Mitigating Skill Overfitting
- MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation
- Practice Makes Unsafe: Skill Misevolution in Self-Improving LLM Agents
- Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills
- SkillOS: Learning Skill Curation for Self-Evolving Agents
- WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution
- SkillGrad: Optimizing Agent Skills Like Gradient Descent
- SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning
- OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
- EvoArena: Tracking Memory Evolution for Robust LLM Agents in Dynamic Environments
- Rethinking Self-Evolving Agents: Do We Still Need Prescribed Optimization Pipelines?
- PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents
- AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?
- SkillOpt: Executive Strategy for Self-Evolving Agent Skills
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection