The Blind Curator: How a Biased Judge Silently Disables Skill Retirement in Self-Evolving Agents
Xing Zhang, Yanwei Cui, Guanghui Wang, Ziyuan Li, Wei Qiu, Bing Zhu, Peiyang He
cs.AI, cs.CL, cs.CR
Submitted: 2026-07-08
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: A self-evolving agent retires its bad skills by watching them fail, so what happens when the judge cannot see the failures? Skill retirement is the structural constraint that keeps a growing library
Terminology
Abstract
A self-evolving agent retires its bad skills by watching them fail, so what happens when the judge cannot see the failures? Skill retirement is the structural constraint that keeps a growing library from drifting below the no-skill baseline, but its guarantee assumes an unbiased reward, which is false for the LLM judges that reference-free tasks force upon us. We show that a biased judge does not merely add noise; it silently switches off the curator. We make this precise with a corrupted-reward analysis and, isolating the causal channel by injecting corruption on top of a deterministic reward, a behavioral study on a reference-free report-writing testbed with a code-generation cross-check. Symmetric noise leaves retirement intact, but false-pass bias (failures slipping through as passes) disables contribution-based retirement past a sharp threshold that no amount of data can cross. Separating genuine retirement from cap-eviction churn shows this mechanism failure is universal, holding across domains and failure rates and sparing only near-zero-false-pass, verifier-like graders. The downstream outcome, though, is regime-dependent: eval quality degrades only where the same corruption also starves skill synthesis, and otherwise holds steady, so the disabled curator is silent, surfacing in no aggregate metric. The contribution is a behavioral safety result, not a performance one. A cheap defect-injection audit then tells an operator, before deployment, which side of the threshold their judge occupies.
Sources
- Program Synthesis with Large Language Models
- Constitutional AI: Harmlessness from AI Feedback
- Closing the Feedback Loop: From Experience Extraction to Insight Governance in Verbal Reinforcement Learning
- CASCADE: Cumulative Agentic Skill Creation through Autonomous Development and Evolution
- LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods
- SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
- SkillsVote: Lifecycle Governance of Agent Skills from Collection, Recommendation to Evolution
- Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills
- MemGPT: Towards LLMs as Operating Systems
- AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs
- Dynamic Skill Lifecycle Management for Agentic Reinforcement Learning
- Large Language Models are Inconsistent and Biased Evaluators
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- Agent Workflow Memory
- EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle
- SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning
- AutoSkill: Experience-Driven Lifelong Learning via Skill Self-Evolution
- SkillAdaptor: Self-Adapting Skills for LLM Agents from Trajectories
- TextGrad: Automatic "Differentiation" via Text
- Library Drift: Diagnosing and Fixing a Silent Failure Mode in Self-Evolving LLM Skill Libraries
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection