ANCHOR: An External LLM-Driven Supervisory Module Facilitating Healthy Evolution in Self-Evolving Systems
cs.AI
Submitted: 2026-06-04
Updated: 2026-09-15
License: http://creativecommons.org/licenses/by/4.0/
The gist: Self-evolving agents improve through continual self-play and self-generated learning signals, but their internally generated tasks and verifier signals provide limited coverage of phase-level errors,
Terminology
Abstract
Self-evolving agents improve through continual self-play and self-generated learning signals, but their internally generated tasks and verifier signals provide limited coverage of phase-level errors, allowing capability degradation and safety drift to accumulate. We introduce ANCHOR, an LLM-based supervisory framework that delivers evaluative feedback at multiple phases of self-evolution and aggregates reviewed signals into context for subsequent steps. We retrofit two representative open-source self-evolving agent frameworks with ANCHOR, and evaluate them across coding, mathematical reasoning, and safety. Our results show that ANCHOR substantially improves safety performance while maintaining stable performance on the core capabilities of the underlying self-evolving agents. Further analyses provide practical insights for future research, showing that execution-result-based supervision is particularly effective and that increasing supervision frequency yields diminishing returns. Together, these results support external LLM-based supervision as a practical approach to developing safer, more stable, and controllable self-evolving agent systems.
Sources
- A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems
- "You Are Rejected!": An Empirical Study of Large Language Models Taking Hiring Evaluations
- Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models
- A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence
- Enhancing Human Experience in Human-Agent Collaboration: A Human-Centered Modeling Approach Based on Positive Human Gain
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Reinforcement Learning from User Feedback
- Measuring Mathematical Problem Solving With the MATH Dataset
- R-Zero: Self-Evolving Reasoning LLM from Zero Data
- Qwen2.5-Coder Technical Report
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
- Solving Quantitative Reasoning Problems with Language Models
- MM-Zero: Self-Evolving Multi-Model Vision Language Models From Zero Data
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
- Qwen2.5 Technical Report
- Your Agent May Misevolve: Emergent Risks in Self-evolving LLM Agents
- Collaborative Gym: A Framework for Enabling and Evaluating Human-Agent Collaboration
- AI & Human Co-Improvement for Safer Co-Superintelligence
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection