SkillGraph: Skill-Augmented Reinforcement Learning for Agents via Evolving Skill Graphs
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "SkillGraph: Skill-Augmented Reinforcement Learning for Agents via Evolving Skill Graphs".
Tom: The gist The proposed framework, SKILLGRAPH, introduces a graph-structured formulation of skill library for LLM agents where skills are connected by explicit prerequisite, enhancement, and co-occurrence relations.
Jane: First, who's behind it and why it matters.
Paper summary: Tom: So wrapping up, SkillGraph proposes this closed-loop system where graph construction distills skills from trajectories, graph-aware retrieval produces ordered skill sequences for the policy, and then graph evolution refines the structure based on training feedback.
Jane: It’s about making the skill library dynamic rather than static. The authors emphasize that this structure allows for dependency-aware retrieval, which is what gets agents past those tricky multi-step problems they struggle with right now.
Meng: From an engineering standpoint, it means we design the learning process not just around the policy update, but around how the skill graph itself can be updated and pruned based on empirical success. It’s a more holistic way to manage agent knowledge.
Lu: The authors are showing that explicitly modeling prerequisite, enhancement, and co-occurrence relations gives agents a powerful mechanism for planning complex actions in an environment.
Lalam: For the future, this means agents can become much better at tasks that require chaining many small actions together because they aren't just guessing which skill comes next; they have a structure guiding them.
Tom: It really shows that for agents to tackle the harder, longer tasks out there, we need to give them a map of how their skills relate to each other, not just a giant pile of things they’ve seen before.
Conclusion: Tom: So SkillGraph is basically taking all those messy skills an agent learns while doing tasks and turning them into a structured map, right?
Jane: Yeah, that’s right. They’re building this graph where skills are connected by things like "this skill helps with that one," or "you need this before you can do that."
Lu: It’s really about giving the AI a memory structure instead of just a giant list of random actions.
Meng: So, the main idea is that when the agent learns something new, it doesn't just add it to a bag; it updates this graph.
Lalam: Because if a skill gets used successfully more often, or if its related skills get used together often, the connection between them gets stronger in this structure.
Tom: And they show that when you use this graph to pick skills for a task, you get an ordered list of skills that actually makes sense for the whole process.
Jane: It moves away from the agent just guessing what to do next and giving it a guided path based on its existing knowledge map.
Lu: The real power here is how they evolve that graph during training, so the agent keeps getting smarter by refining its own way of thinking skills.
Meng: From an engineering view, that means the learning isn't just about the policy changing; it’s also about restructuring the knowledge base itself based on what worked.
Lalam: It gives agents a way to learn from their successes and failures in a much more organized, traceable way than just trial and error.
Tom: And the results show it works really well across different tasks, proving this structured approach is actually effective for complex problem-solving.
University of Science and Technology of China
cs.CL
Submitted: 2026-05-12
Updated: 2026-10-08
Comments: Under Review
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 92/100
The gist: The gist The proposed framework, SKILLGRAPH, introduces a graph-structured formulation of skill library for LLM agents where skills are connected by explicit prerequisite, enhancement, and
Key concepts
- Skill Library Limitations
- Current skill libraries fail because they cannot handle complex tasks that require a specific sequence of actions. They also lack structure for updating skills, meaning they cannot easily merge redundant skills or identify obsolete ones.
- SKILLGRAPH Framework
- SKILLGRAPH organizes agent skills into a graph and co-evolves it with the agent's policy using reinforcement learning. It has three stages: building the initial graph from interactions, using it for dependency-respecting retrieval, and evolving it during training to refine skill nodes and edges.
- Graph Construction
- Skills are distilled into general (domain-independent) and task-specific skills. The graph uses typed edges like Prerequisite (A prereq → B), Enhances (A enhance → B), and Co-occurs. Nodes track usage statistics, which drive the system's evolution decisions.
- Graph Evolution
- This stage updates the skill graph during training by refining skills and relations based on success rates. It includes operations like merging, splitting, and pruning weak edges, while progressive unlocking activates higher-level skills when prerequisite performance thresholds are met.
Terminology
Summary
The gist The proposed framework, SKILLGRAPH, introduces a graph-structured formulation of skill library for LLM agents where skills are connected by explicit prerequisite, enhancement, and co-occurrence relations.
Skill Library Limitations
Existing skill libraries suffer from two key limitations: first, retrieval is not compositional because complex tasks often require an ordered sequence of skills; for example, a “heat and place” task in ALFWorld may require locating an object, picking it up, heating it with an appliance, and then placing it at the target destination Second, skill updates are not structured because the library lacks explicit evidence for merging redundant skills, splitting overly broad skills, deprecating obsolete skills, or strengthening useful relations between skills
SKILLGRAPH Framework
SKILLGRAPH is a framework that organizes agent skills into a structured graph and co-evolves it with the agent’s policy through reinforcement learning (RL) The framework consists of three stages:
-
Graph construction builds an initial skill graph from interaction trajectories, making inter-skill relations explicit
-
Graph-aware retrieval starts from task-relevant seed skills, expands along graph edges, and orders retrieved skills according to their dependencies
-
Graph evolution updates the graph during training by refining skill nodes and adjusting edge relations according to skill usage and success rate
Graph Construction Details
During graph construction, skills are distilled from trajectories into two types: general skills, which capture domain-independent reasoning strategies applicable across tasks, and task-specific skills, which encode strategies tied to particular task types Each skill is represented as a compact record containing a title, a core principle describing the strategy, an applicability condition, and a category label indicating its type The graph structure is defined by three typed edges: Prerequisite (A prereq → B), Enhances (A enhance → B), and Co-occurs (A co occur ←→ B) Each node v maintains running statistics—usage count nuse(v), success count nsucc(v), and empirical success rate pˆ(v) = nsucc(v)/nuse(v)—that drive both evolution decisions and progressive unlocking Based on the directed prerequisite and enhancement edges, each node is assigned a topological level l(v) indicating its position in the dependency hierarchy
Graph-Aware Retrieval
Graph-aware retrieval produces a dependency-respecting sequence of skills by traversing the skill graph given a task description d with task type t(d) The procedure involves three steps:
-
Seed selection identifies task-relevant entry points from the currently active skill set Vactive, selecting all general skills and task-type-matched skills as seed nodes
-
Graph expansion explores outgoing edges via beam search with beam width B, producing the forward-expanded set Rbeam
-
Topological ordering sorts the union of seeds, backward-expanded, and forward-expanded skills according to the graph’s dependency edges, producing an ordered skill sequence This sequence is prepended to the task prompt as structured guidance for the policy
Graph Evolution
The graph evolution stage updates both skill nodes and their edges at each validation step, driven by trajectory-level feedback Node-level operations include Insert, Merge, Split, and Deprecate skills based on diagnostic signals from the training process Edge-level operations include Path reinforcement, Co-occurrence discovery, and Decay and pruning of weak edges Progressive unlocking exposes skills based on their topological level, activating level-(L+1) skills when the average success rate of level-L skills exceeds an unlocking threshold θunlock
Policy Optimization and Training Loop
The agent policy πθ is optimized using GRPO, where the loss function includes a KL penalty anchored to the reference policy πref The closed-loop training procedure involves sampling G rollouts from πθ conditioned on the task description d and the retrieved skill sequence Rret, updating θ via GRPO, and then executing the full graph evolution pipeline This closed training loop ensures that the improving policy generates richer trajectories that further refine the graph through node- and edge-level updates, while the refined graph provides higher-quality structured retrieval that accelerates subsequent policy learning
Experimental Results
Experiments on ALFWorld, WebShop, and seven search-augmented QA tasks show that SKILLGRAPH achieves state-of-the-art performance across benchmarks Notably, SKILLGRAPH with a 7B open-source model substantially outperforms closed-source LLMs by surpassing GPT-4o by 42.6 points on ALFWorld and Gemini-2.5-Pro by 30.3 points SKILLGRAPH also achieves the best overall performance on both benchmarks, with a 7B open-source model surpassing GPT-4o by 42.6 points and Gemini-2.5-Pro by 30.3 points on ALFWorld, and exceeds both by over 48 points on WebShop
Analysis of Contributions
Ablation study shows that removing graph-aware retrieval causes the largest single drop (−31.2) on ALFWorld, confirming that the rigid multi-step subtasks critically depend on prerequisite-ordered skill sequences On WebShop, graph evolution (−14.1) and graph structure (−11.
Improvements for AI systems
- Bold header: Graph-structured skill memory for compositional planning
SKILLGRAPH represents reusable skills as nodes in a directed graph, with typed edges encoding prerequisite, enhancement, and co-occurrence relations.
This allows agents to retrieve an ordered skill subgraph that can guide multistep decision making,
enabling them to execute complex sequences like a heat and place
task by identifying necessary dependencies.
- Bold header: Closed-loop skill library maintenance
The framework implements a continuous learning loop where the agent's policy and the graph evolve jointly, allowing for principled evolution that uses training feedback to refine both individual skills and their relations.
This includes node-level operations like Insert Merge Split Deprecate
to manage skill redundancy and obsolescence automatically.
- Bold header: Dependency-aware skill retrieval
The system moves beyond flat collections by implementing a procedure that traverses the graph to produce a dependency-respecting sequence of skills.
This ensures that the agent receives skills in a natural simple-to-complex order that mirrors how sub-tasks should be composed,
directly addressing the failure of flat retrievers to indicate execution order.
- Bold header: Progressive skill unlocking as an automatic curriculum
SKILLGRAPH utilizes Progressive Unlocking
where level-(L+1) skills are activated
only when the average success rate of level-L skills exceeds a threshold, ensuring that agents build competence from the ground up, with advanced compositional skills becoming available only when their foundations are reliable.
- Bold header: Enhanced generalization to novel tasks
The framework demonstrates strong generalization because the graph structure improves skill reuse, reduces redundancy compared with flat libraries, and enables transfer of compositional knowledge from simpler tasks to more complex ones,
confirming that the structured representation learns inter-skill relations
transferable across different environments like ALFWorld and WebShop.
Abstract
Skill libraries enable large language model agents to reuse experience from past interactions, but most existing libraries store skills as isolated entries and retrieve them only by semantic similarity. This leads to two key challenges for compositional tasks. Firstly, an agent must identify not only relevant skills but also how they depend on and build upon each other. Secondly, it also makes library maintenance difficult, since the system lacks structural cues for deciding when skills should be merged, split, or removed. We propose SKILLGRAPH, a framework that represents reusable skills as nodes in a directed graph, with typed edges encoding prerequisite, enhancement, and co-occurrence relations. Given a new task, SKILLGRAPH retrieves not just individual skills, but an ordered skill subgraph that can guide multi-step decision making. The graph is continuously updated from agent trajectories and reinforcement learning feedback, allowing both the skill library and the agent policy to improve together. Experiments on ALFWorld, WebShop, and seven search-augmented QA tasks show that SKILLGRAPH achieves state-of-the-art performance against memory-augmented RL methods, with especially large gains on complex tasks that require composing multiple skills.
Sources
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- From Local to Global: A Graph RAG Approach to Query-Focused Summarization
- Memp: Exploring Agent Procedural Memory
- GPT-4o System Card
- OpenAI o1 System Card
- SimpleMem: Efficient Lifelong Memory for LLM Agents
- ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- ZeroSearch: Incentivize the Search Capability of LLMs without Searching
- Agent KB: Leveraging Cross-Domain Experience for Agentic Problem Solving
- Mem-{\alpha}: Learning Memory Construction via Reinforcement Learning
- EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle
- SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning
- Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward
- Qwen3 Technical Report
- MemEvolve: Meta-Evolution of Agent Memory Systems
- MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering