From Procedural Skills to Strategy Genes: Towards Experience-Driven Test-Time Evolution
cs.SE, cs.CL
Submitted: 2026-04-16
Updated: 2026-09-17
Comments: Technical Report
Code: https://github.com/EvoMap/skill2gep
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Skill-Pro: Learning Reusable Skills from Experience via Non-Parametric PPO for LLM Agents
- MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents
- Memento-Skills: Let Agents Design Agents
- CoWork-X: Experience-Optimized Co-Evolution for Multi-Agent Collaboration System
- Memp: Exploring Agent Procedural Memory
- AutoRefine: Compiling Trajectories into Validated Typed Agent Artifacts
- Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory
- TAME: A Trustworthy Test-Time Evolution of Agent Memory with Systematic Benchmarking
- UMEM: Unified Memory Extraction and Management Framework for Generalizable Memory
- MemEvolve: Meta-Evolution of Agent Memory Systems
- SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
- Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward
- Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?
- Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties