WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution
cs.AI, cs.CL
Submitted: 2026-08-27
Updated: 2026-08-27
License: http://creativecommons.org/licenses/by/4.0/
The gist: Agent skills package specialized knowledge and workflows into reusable resources that extend AI agent capabilities.
Terminology
Abstract
Agent skills package specialized knowledge and workflows into reusable resources that extend AI agent capabilities. Recent work automatically discovers such skills from agent experience, which enables agents to progressively adapt through interaction. However, the insights that guide skill development typically remain scattered across optimization histories, limiting their systematic reuse across iterations. We introduce WikiSkill, a framework that co-evolves agent skills with a persistent knowledge base (wiki). At a high level, WikiSkill separates raw execution experience, accumulated knowledge, and executable skills, while continuously consolidating experience into the wiki, which subsequent skill updates can build on. Across diverse benchmarks and models, WikiSkill consistently outperforms state-of-the-art skill-evolution methods and improves over no-skill baselines in most model-benchmark settings. We find that skill evolution complements model scaling: larger models generally benefit more from evolved skills, while smaller models with skills can outperform substantially larger models without them. We also find that evolved skills transfer effectively across models and model families, and skills evolved by other models can outperform self-evolved skills. Finally, our ablation studies confirm that persistent knowledge accumulation in the wiki is critical for effective skill evolution. These results demonstrate the benefits of systematically accumulating and refining agent experience for developing reusable and transferable skills.
Sources
- EvoSkill: Automated Skill Discovery for Multi-Agent Systems
- SkillCraft: Can LLM Agents Learn to Use Tools Skillfully?
- HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry
- SkillRet: A Large-Scale Benchmark for Skill Retrieval in LLM Agents
- Gemma 4 Technical Report
- LiveMathematicianBench: A Live Benchmark for Research-Level Mathematical Reasoning with Proof Sketches
- AA-Omniscience: Evaluating Cross-Domain Knowledge Reliability in Large Language Models
- SoK: Agentic Skills -- Beyond Tool Use in LLM Agents
- Meta-Harness: End-to-End Optimization of Model Harnesses
- SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
- SkillNet: Create, Evaluate, and Connect AI Skills
- Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses
- How Well Do Agentic Skills Work in the Wild: Benchmarking LLM Skill Usage in Realistic Settings
- AutoHarness: improving LLM agents by automatically synthesizing a code harness
- SKILL0: In-Context Agentic Reinforcement Learning for Skill Internalization
- Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills
- SkillOS: Learning Skill Curation for Self-Evolving Agents
- Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning
- Skill Retrieval Augmentation for Agentic AI
- SkillGrad: Optimizing Agent Skills Like Gradient Descent
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection