SkillGraph: Skill-Augmented Reinforcement Learning for Agents via Evolving Skill Graphs

arXiv:2605.12039 · cs.CL · Submitted 2026-05-12 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "SkillGraph: Skill-Augmented Reinforcement Learning for Agents via Evolving Skill Graphs".

Tom: The gist The proposed framework, SKILLGRAPH, introduces a graph-structured formulation of skill library for LLM agents where skills are connected by explicit prerequisite, enhancement, and co-occurrence relations.

Jane: First, who's behind it and why it matters.

Paper summary: Tom: So wrapping up, SkillGraph proposes this closed-loop system where graph construction distills skills from trajectories, graph-aware retrieval produces ordered skill sequences for the policy, and then graph evolution refines the structure based on training feedback.

Jane: It’s about making the skill library dynamic rather than static. The authors emphasize that this structure allows for dependency-aware retrieval, which is what gets agents past those tricky multi-step problems they struggle with right now.

Meng: From an engineering standpoint, it means we design the learning process not just around the policy update, but around how the skill graph itself can be updated and pruned based on empirical success. It’s a more holistic way to manage agent knowledge.

Lu: The authors are showing that explicitly modeling prerequisite, enhancement, and co-occurrence relations gives agents a powerful mechanism for planning complex actions in an environment.

Lalam: For the future, this means agents can become much better at tasks that require chaining many small actions together because they aren't just guessing which skill comes next; they have a structure guiding them.

Tom: It really shows that for agents to tackle the harder, longer tasks out there, we need to give them a map of how their skills relate to each other, not just a giant pile of things they’ve seen before.

Conclusion: Tom: So SkillGraph is basically taking all those messy skills an agent learns while doing tasks and turning them into a structured map, right?

Jane: Yeah, that’s right. They’re building this graph where skills are connected by things like "this skill helps with that one," or "you need this before you can do that."

Lu: It’s really about giving the AI a memory structure instead of just a giant list of random actions.

Meng: So, the main idea is that when the agent learns something new, it doesn't just add it to a bag; it updates this graph.

Lalam: Because if a skill gets used successfully more often, or if its related skills get used together often, the connection between them gets stronger in this structure.

Tom: And they show that when you use this graph to pick skills for a task, you get an ordered list of skills that actually makes sense for the whole process.

Jane: It moves away from the agent just guessing what to do next and giving it a guided path based on its existing knowledge map.

Lu: The real power here is how they evolve that graph during training, so the agent keeps getting smarter by refining its own way of thinking skills.

Meng: From an engineering view, that means the learning isn't just about the policy changing; it’s also about restructuring the knowledge base itself based on what worked.

Lalam: It gives agents a way to learn from their successes and failures in a much more organized, traceable way than just trial and error.

Tom: And the results show it works really well across different tasks, proving this structured approach is actually effective for complex problem-solving.

University of Science and Technology of China

cs.CL

Submitted: 2026-05-12

Updated: 2026-10-08

Comments: Under Review

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 92/100

The gist: The gist The proposed framework, SKILLGRAPH, introduces a graph-structured formulation of skill library for LLM agents where skills are connected by explicit prerequisite, enhancement, and

Key concepts

Skill Library Limitations
Current skill libraries fail because they cannot handle complex tasks that require a specific sequence of actions. They also lack structure for updating skills, meaning they cannot easily merge redundant skills or identify obsolete ones.
SKILLGRAPH Framework
SKILLGRAPH organizes agent skills into a graph and co-evolves it with the agent's policy using reinforcement learning. It has three stages: building the initial graph from interactions, using it for dependency-respecting retrieval, and evolving it during training to refine skill nodes and edges.
Graph Construction
Skills are distilled into general (domain-independent) and task-specific skills. The graph uses typed edges like Prerequisite (A prereq → B), Enhances (A enhance → B), and Co-occurs. Nodes track usage statistics, which drive the system's evolution decisions.
Graph Evolution
This stage updates the skill graph during training by refining skills and relations based on success rates. It includes operations like merging, splitting, and pruning weak edges, while progressive unlocking activates higher-level skills when prerequisite performance thresholds are met.

Terminology

Summary

The gist The proposed framework, SKILLGRAPH, introduces a graph-structured formulation of skill library for LLM agents where skills are connected by explicit prerequisite, enhancement, and co-occurrence relations.

Skill Library Limitations

Existing skill libraries suffer from two key limitations: first, retrieval is not compositional because complex tasks often require an ordered sequence of skills; for example, a “heat and place” task in ALFWorld may require locating an object, picking it up, heating it with an appliance, and then placing it at the target destination Second, skill updates are not structured because the library lacks explicit evidence for merging redundant skills, splitting overly broad skills, deprecating obsolete skills, or strengthening useful relations between skills

SKILLGRAPH Framework

SKILLGRAPH is a framework that organizes agent skills into a structured graph and co-evolves it with the agent’s policy through reinforcement learning (RL) The framework consists of three stages:

  1. Graph construction builds an initial skill graph from interaction trajectories, making inter-skill relations explicit

  2. Graph-aware retrieval starts from task-relevant seed skills, expands along graph edges, and orders retrieved skills according to their dependencies

  3. Graph evolution updates the graph during training by refining skill nodes and adjusting edge relations according to skill usage and success rate

Graph Construction Details

During graph construction, skills are distilled from trajectories into two types: general skills, which capture domain-independent reasoning strategies applicable across tasks, and task-specific skills, which encode strategies tied to particular task types Each skill is represented as a compact record containing a title, a core principle describing the strategy, an applicability condition, and a category label indicating its type The graph structure is defined by three typed edges: Prerequisite (A prereq → B), Enhances (A enhance → B), and Co-occurs (A co occur ←→ B) Each node v maintains running statistics—usage count nuse(v), success count nsucc(v), and empirical success rate pˆ(v) = nsucc(v)/nuse(v)—that drive both evolution decisions and progressive unlocking Based on the directed prerequisite and enhancement edges, each node is assigned a topological level l(v) indicating its position in the dependency hierarchy

Graph-Aware Retrieval

Graph-aware retrieval produces a dependency-respecting sequence of skills by traversing the skill graph given a task description d with task type t(d) The procedure involves three steps:

  1. Seed selection identifies task-relevant entry points from the currently active skill set Vactive, selecting all general skills and task-type-matched skills as seed nodes

  2. Graph expansion explores outgoing edges via beam search with beam width B, producing the forward-expanded set Rbeam

  3. Topological ordering sorts the union of seeds, backward-expanded, and forward-expanded skills according to the graph’s dependency edges, producing an ordered skill sequence This sequence is prepended to the task prompt as structured guidance for the policy

Graph Evolution

The graph evolution stage updates both skill nodes and their edges at each validation step, driven by trajectory-level feedback Node-level operations include Insert, Merge, Split, and Deprecate skills based on diagnostic signals from the training process Edge-level operations include Path reinforcement, Co-occurrence discovery, and Decay and pruning of weak edges Progressive unlocking exposes skills based on their topological level, activating level-(L+1) skills when the average success rate of level-L skills exceeds an unlocking threshold θunlock

Policy Optimization and Training Loop

The agent policy πθ is optimized using GRPO, where the loss function includes a KL penalty anchored to the reference policy πref The closed-loop training procedure involves sampling G rollouts from πθ conditioned on the task description d and the retrieved skill sequence Rret, updating θ via GRPO, and then executing the full graph evolution pipeline This closed training loop ensures that the improving policy generates richer trajectories that further refine the graph through node- and edge-level updates, while the refined graph provides higher-quality structured retrieval that accelerates subsequent policy learning

Experimental Results

Experiments on ALFWorld, WebShop, and seven search-augmented QA tasks show that SKILLGRAPH achieves state-of-the-art performance across benchmarks Notably, SKILLGRAPH with a 7B open-source model substantially outperforms closed-source LLMs by surpassing GPT-4o by 42.6 points on ALFWorld and Gemini-2.5-Pro by 30.3 points SKILLGRAPH also achieves the best overall performance on both benchmarks, with a 7B open-source model surpassing GPT-4o by 42.6 points and Gemini-2.5-Pro by 30.3 points on ALFWorld, and exceeds both by over 48 points on WebShop

Analysis of Contributions

Ablation study shows that removing graph-aware retrieval causes the largest single drop (−31.2) on ALFWorld, confirming that the rigid multi-step subtasks critically depend on prerequisite-ordered skill sequences On WebShop, graph evolution (−14.1) and graph structure (−11.

Improvements for AI systems

  1. Bold header: Graph-structured skill memory for compositional planning

SKILLGRAPH represents reusable skills as nodes in a directed graph, with typed edges encoding prerequisite, enhancement, and co-occurrence relations. This allows agents to retrieve an ordered skill subgraph that can guide multistep decision making, enabling them to execute complex sequences like a heat and place task by identifying necessary dependencies.

  1. Bold header: Closed-loop skill library maintenance

The framework implements a continuous learning loop where the agent's policy and the graph evolve jointly, allowing for principled evolution that uses training feedback to refine both individual skills and their relations. This includes node-level operations like Insert Merge Split Deprecate to manage skill redundancy and obsolescence automatically.

  1. Bold header: Dependency-aware skill retrieval

The system moves beyond flat collections by implementing a procedure that traverses the graph to produce a dependency-respecting sequence of skills. This ensures that the agent receives skills in a natural simple-to-complex order that mirrors how sub-tasks should be composed, directly addressing the failure of flat retrievers to indicate execution order.

  1. Bold header: Progressive skill unlocking as an automatic curriculum

SKILLGRAPH utilizes Progressive Unlocking where level-(L+1) skills are activated only when the average success rate of level-L skills exceeds a threshold, ensuring that agents build competence from the ground up, with advanced compositional skills becoming available only when their foundations are reliable.

  1. Bold header: Enhanced generalization to novel tasks

The framework demonstrates strong generalization because the graph structure improves skill reuse, reduces redundancy compared with flat libraries, and enables transfer of compositional knowledge from simpler tasks to more complex ones, confirming that the structured representation learns inter-skill relations transferable across different environments like ALFWorld and WebShop.

Abstract

Skill libraries enable large language model agents to reuse experience from past interactions, but most existing libraries store skills as isolated entries and retrieve them only by semantic similarity. This leads to two key challenges for compositional tasks. Firstly, an agent must identify not only relevant skills but also how they depend on and build upon each other. Secondly, it also makes library maintenance difficult, since the system lacks structural cues for deciding when skills should be merged, split, or removed. We propose SKILLGRAPH, a framework that represents reusable skills as nodes in a directed graph, with typed edges encoding prerequisite, enhancement, and co-occurrence relations. Given a new task, SKILLGRAPH retrieves not just individual skills, but an ordered skill subgraph that can guide multi-step decision making. The graph is continuously updated from agent trajectories and reinforcement learning feedback, allowing both the skill library and the agent policy to improve together. Experiments on ALFWorld, WebShop, and seven search-augmented QA tasks show that SKILLGRAPH achieves state-of-the-art performance against memory-augmented RL methods, with especially large gains on complex tasks that require composing multiple skills.

Sources

Related papers