SkillLift: Learning Dense Rubrics from Sparse Oracles for Efficient Skill Evolution
cs.AI
Submitted: 2026-09-14
Updated: 2026-09-14
Code: https://github.com/WalteR-MittY-pro/SkillLift
License: http://creativecommons.org/licenses/by/4.0/
The gist: LLM-based agents increasingly rely on persistent skills, i.e., reusable procedural prompts, to adapt without weight updates.
Terminology
Abstract
LLM-based agents increasingly rely on persistent skills, i.e., reusable procedural prompts, to adapt without weight updates. Existing skill self-evolution methods directly revise skill text based on execution feedback, but each oracle evaluation requires a full agent rollout, creating a supervision bottleneck that confines search to failure-patching updates. Our key insight is that ranking is a smoother supervision target than absolute outcome regression: identifying which skill is better requires fewer oracle evaluations than predicting exact scores. Building on this insight, we propose SkillLift, which decouples skill search from oracle cost by learning an oracle-aligned rubric as a structured evaluation space. We formalize this as a bilevel optimization problem solved via alternating optimization: an inner loop uses the frozen rubric as a cheap surrogate to guide skill revision at no oracle cost, while an outer loop invokes a small number of oracle rollouts to re-align the rubric via rank correlation, amortizing oracle cost and stabilizing text-space updates. Experiments on complex agent task benchmarks show that our method outperforms existing auto-skill methods with 40--70% less token cost compared to frontier evolving methods. Codes are available at https://github.com/WalteR-MittY-pro/SkillLift.
Sources
- EvoSkill: Automated Skill Discovery for Multi-Agent Systems
- SkillCAT: Contrastive, Assessment-Augmented and Topology-AwareSkill Self-Evolution for LLM Agents
- WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
- Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
- SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
- MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation
- SkillClaw: Let Skills Evolve Collectively with Agentic Evolver
- Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills
- DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research
- SKILLFOUNDRY: Building Self-Evolving Agent Skill Libraries from Heterogeneous Scientific Resources
- SkillGrad: Optimizing Agent Skills Like Gradient Descent
- OpenHands: An Open Platform for AI Software Developers as Generalist Agents
- SkillOpt: Executive Strategy for Self-Evolving Agent Skills
- CoEvoSkills: Self-Evolving Agent Skills via Co-Evolutionary Verification
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection