Rethinking Self-Evolution: A Constrained Exploration-Exploitation Process for Mitigating Skill Overfitting
Hongqiang Lin, Chao Liu, Xiaofan Bai, Xuan Jin, Yuhong Li, Nenggan Zheng, Xipeng Cao
cs.AI, cs.LG
Submitted: 2026-08-19
Updated: 2026-08-20
Terminology
Sources
- The Scaling Laws of Skills in LLM Agent Systems
- A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems
- LiveMathematicianBench: A Live Benchmark for Research-Level Mathematical Reasoning with Proof Sketches
- A Survey of Agent Memory in the Second Half: Towards Self-Evolving and Long-Horizon Agents
- From Raw Experience to Skill Consumption: A Systematic Study of Model-Generated Agent Skills
- SoK: Agentic Skills -- Beyond Tool Use in LLM Agents
- Organizing, Orchestrating, and Benchmarking Agent Skills at Ecosystem Scale
- MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation
- SkillForge: Forging Domain-Specific, Self-Evolving Agent Skills in Cloud Technical Support
- SkillClaw: Let Skills Evolve Collectively with Agentic Evolver
- Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills
- SkillOS: Learning Skill Curation for Self-Evolving Agents
- AutoRefine: Compiling Trajectories into Validated Typed Agent Artifacts
- SkillOpt-Lite: Better and Faster Agent Self-evolution via One Line of Vibe
- Skill-R1: Agent Skill Evolution via Reinforcement Learning
- SkillX: Automatically Constructing Skill Knowledge Bases for Agents
- SkillGrad: Optimizing Agent Skills Like Gradient Descent
- Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory
- SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning
- Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection