SKILLER: Language-Level Reinforcement Learning for Reusable Skill Extraction in Small Language Models
Chenhao Dang, Siyuan Xiong, Conghui He, Weijia Li
Shanghai Jiao Tong University · Shanghai Artificial Intelligence Laboratory · Harbin Institute of Technology, Shenzhen · Tsinghua Shenzhen International Graduate School, Tsinghua University
cs.AI
Submitted: 2026-08-14
Updated: 2026-08-17
Comments: 15 pages, 6 figures, and 8 tables. Submitted to AAAI 2027. Project and code: https://github.com/DANG-ai/SKILLER
Code: https://github.com/DANG-ai/SKILLER
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 95/100
The gist: SKILLER is a natural-language-driven reinforcement learning framework designed to automatically generate and optimize executor-specific skills for small-scale language models (LVLMs).
Terminology
Summary
SKILLER is a natural-language-driven reinforcement learning framework designed to automatically generate and optimize executor-specific skills for small-scale language models (LVLMs). The framework addresses the model-mismatch problem,
where skills crafted for strong frontier models do not natively transfer to small-scale LVLMs, often causing catastrophic task failure due to cognitive overload, argument hallucination, or derailment by complex instructions. SKILLER treats the textual skill itself as the optimizable policy, employing a strong frontier model (e.g., GPT-5.4) as the actor and critic, while the environment is the agent loop driven by the target open-source small-scale LVLM. All reinforcement learning signals—states, diagnostic rewards, and policy update actions—are propagated entirely via structured natural language, with no neural weight updates.
The framework operates through an iterative optimization loop: at each step, the environment executes the current skill with the compact model, producing a trajectory, scalar reward, and verifier diagnostics. The critic module evaluates the current skill, compares the observed trajectory with a reference trajectory, locates the earliest causal error, and generates natural-language modification suggestions. The actor module then applies bounded edits (Insert, Replace, Create, Delete) to the skill, potentially synthesizing task-local helper scripts to offload complex procedural reasoning into deterministic external tools. A replay memory preserves failure signatures, critic diagnoses, and accepted edits across steps to prevent repeated failures and protect effective behavior.
Extensive evaluations were conducted using Qwen3.5-9B and Qwen3.5-4B across five benchmarks: SkillsBench, SkillLearnBench, SWE-Skills-Bench, GAIA, and EarthBench. SKILLER consistently outperformed three open-source baselines (AutoSkill, EvoSkill, SkillX) and one closed-source baseline (Manus), achieving absolute gains ranging from 4.3 to 20.4 percentage points for the 9B model and 1.8 to 13.3 points for the 4B model. Notably, on single-skill tasks in SkillsBench, Qwen3.5-9B with SKILLER skills matched or exceeded the performance of strong closed-source models like Claude Opus 4.7 Max, while being 167x cheaper in output-token cost. On SWE-Skills-Bench, Qwen3.5-4B with SKILLER surpassed the performance of the larger Qwen3.5-9B model when deployed with human-authored, AutoSkill, EvoSkill, SkillX, or Manus-generated skills, demonstrating that optimized procedural control can be more valuable than raw parameter scaling.
Zero-shot results on held-out halves of GAIA and EarthBench showed that SKILLER extracts reusable procedural rules rather than overfitting to generation instances. Learning dynamics revealed that complex software engineering workflows benefit from cumulative policy specialization across all five iterations, while simpler tasks converge within the first two steps. Structural analysis showed that SKILLER produces concise instructions with low TF-IDF similarity (matching human developers) while maximizing code-level execution capabilities through helper scripts, accommodating the limited context windows of small models. Cost-effectiveness analysis showed that SKILLER achieves a commanding performance advantage (62.86 average score) while remaining more cost-efficient (8.95) than exhaustive generation methods like SkillX (14.55), proving that targeted behavioral alignment offers the highest return on investment.
Ablation studies confirmed the importance of each state component (task instance, current trajectory, reference trajectory, verifier diagnostics), critic operations (evaluate, compare, locate error, generate suggestions), and actor operations (bounded edits, script synthesis, content preservation). The framework's progressive skill evolution was illustrated through a qualitative example on the springboot-tdd skill, showing how feedback converts open-ended exploration into bounded decisions and binds edits to observed interfaces and verifier contracts. The project is available at https://github.com/DANG-ai/SKILLER.
Improvements for AI systems
Improvements to AI Systems:
-
Add a Natural-Language Policy Layer for Small Models: Integrate SKILLER’s core mechanism—treating the textual skill as an optimizable policy—into any small-scale LVLM (e.g., 4B–9B) deployment. This allows the system to dynamically rewrite its own instruction prompts based on environmental feedback, without updating neural weights, preventing cognitive overload and argument hallucination.
-
Implement a Bounded Edit Operator for Instruction Refinement: Use SKILLER’s actor operations (Insert, Replace, Create, Delete) to enable an AI system to make minimal, targeted changes to its task instructions after each failure. This reduces catastrophic derailment from overly complex or ambiguous prompts, ensuring stable execution on hardware-constrained devices.
-
Add a Replay Memory for Failure-Signature Prevention: Incorporate SKILLER’s replay memory into an agent’s control loop, storing failure signatures, critic diagnoses, and accepted edits. The improved system will avoid repeating known mistakes in future tasks, accelerating convergence on new but similar problems (e.g., in software engineering or tool use).
-
Enable Verifier-Coupled Self-Correction: Build an AI system that uses SKILLER’s critic module—evaluate, compare with reference trajectory, locate earliest causal error, generate suggestions—to produce structured, natural-language feedback. This allows the system to correct its own reasoning path mid-task, not just at the end, improving robustness in multi-step workflows.
-
Synthesize Helper Scripts for Procedural Offloading: Integrate SKILLER’s script-synthesis capability, where the system generates deterministic external tools (e.g., Python functions) to handle complex procedural logic. The improved AI can then focus its limited context window on high-level reasoning, reducing token usage and error rates in tasks like code generation or data analysis.
-
Implement Cost-Aware Skill Optimization: Use SKILLER’s cost-effectiveness analysis to build a system that automatically selects between exhaustive generation (e.g., SkillX) and targeted behavioral alignment (SKILLER) based on task complexity and budget. The system will maximize performance per dollar, making advanced AI capabilities feasible for low-resource deployments.
What the Improved AI System Can Do:
-
Run on small models (4B–9B) with performance matching or exceeding frontier models (e.g., Claude Opus 4.7) on single-skill tasks, at 167x lower output-token cost.
-
Self-improve its own instructions in real time, adapting to new environments or verifier contracts without retraining, using only natural-language feedback.
-
Avoid repeated failures by recalling past error signatures and applying proven edits, leading to faster convergence on complex benchmarks (e.g., SWE-Skills-Bench) even with smaller parameter counts.
-
Generate and invoke helper scripts on the fly, enabling it to solve procedural tasks (e.g., Spring Boot TDD, GAIA-style reasoning) that would otherwise exceed its context window.
-
Operate cost-efficiently, achieving high average scores (e.g., 62.86) while spending less (8.95) than alternative methods, making it viable for edge devices, real-time applications, or budget-constrained research.
-
Transfer learned procedural rules to unseen tasks (zero-shot generalization) by extracting reusable, concise instructions rather than memorizing training instances.
Abstract
Agent skills represent a standardized format for packaging procedural knowledge and domain expertise, serving within agent harness systems as an essential mechanism to continually constrain a language model's behavior space for repeatable, high-quality task execution. However, because strong closed-source models entail high inference costs, current popular agent harnesses, such as Codex and OpenClaw, remain prohibitively expensive when deploying these skills to accomplish real-world tasks. The rapid capability enhancement of open-source models deployable on consumer-grade GPUs presents a compelling opportunity to drastically reduce these costs by leveraging skill-based behavioral constraints. Nevertheless, automatically generating effective skills tailored specifically for such compact models remains a significant practical challenge. To address this, we propose SKILLER, a natural-language-driven reinforcement learning framework designed to automatically generate executor-specific skills for small models, which employs a strong model as the actor and critic, treats the small-model agent system as the environment, and propagates all reinforcement learning signals entirely via natural language. Extensive experimental evaluations across five relevant benchmarks using Qwen3.5-9B and Qwen3.5-4B demonstrate that SKILLER outperforms three open-source and one closed-source skill generation or evolution methods, achieving absolute gains ranging from 4.3 to 20.4 percentage points for the 9B model and 1.8 to 13.3 points for the 4B model, while remarkably matching the performance of strong closed-source models on single-skill tasks in SkillsBench. The project is available at https://github.com/DANG-ai/SKILLER.
Sources
- EvoSkill: Automated Skill Discovery for Multi-Agent Systems
- Automating Skill Acquisition through Large-Scale Mining of Open-Source Agentic Repositories: A Framework for Multi-Agent Procedural Knowledge Extraction
- SkillCraft: Can LLM Agents Learn to Use Tools Skillfully?
- Earth-Agent: Unlocking the Full Landscape of Earth Observation with Agents
- SWE-Skills-Bench: Do Agent Skills Actually Help in Real-World Software Engineering?
- From Raw Experience to Skill Consumption: A Systematic Study of Model-Generated Agent Skills
- SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
- SkillNet: Create, Evaluate, and Connect AI Skills
- Agent Skills: A Data-Driven Analysis of Claude Skills for Extending Large Language Model Functionality
- How Well Do Agentic Skills Work in the Wild: Benchmarking LLM Skill Usage in Realistic Settings
- SkillClaw: Let Skills Evolve Collectively with Agentic Evolver
- Skill-Pro: Learning Reusable Skills from Experience via Non-Parametric PPO for LLM Agents
- Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills
- AutoRefine: Compiling Trajectories into Validated Typed Agent Artifacts
- SkillX: Automatically Constructing Skill Knowledge Bases for Agents
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- Reinforcement Learning for Self-Improving Agent with Skill Library
- SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning
- Agent Skill Framework: Perspectives on the Potential of Small to Medium Language Models in Industrial Environments
- AutoSkill: Experience-Driven Lifelong Learning via Skill Self-Evolution
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection