SkillGraph: Skill-Augmented Reinforcement Learning for Agents via Evolving Skill Graphs
summary
The gist
The gist The proposed framework, SKILLGRAPH, introduces a graph-structured formulation of skill library for LLM agents where skills are connected by explicit prerequisite, enhancement, and
In short
SKILLGRAPH introduces a graph-structured skill library for LLM agents, addressing limitations in existing libraries where skills are not ordered or structured for updates. It organizes skills into a graph with prerequisite, enhancement, and co-occurrence relations. This framework uses reinforcement learning to co-evolve the skill graph with the agent's policy, leading to state-of-the-art performance on tasks like ALFWorld.
Key concepts
- Skill Library Limitations
- Current skill libraries fail because they cannot handle complex tasks that require a specific sequence of actions. They also lack structure for updating skills, meaning they cannot easily merge redundant skills or identify obsolete ones.
- SKILLGRAPH Framework
- SKILLGRAPH organizes agent skills into a graph and co-evolves it with the agent's policy using reinforcement learning. It has three stages: building the initial graph from interactions, using it for dependency-respecting retrieval, and evolving it during training to refine skill nodes and edges.
- Graph Construction
- Skills are distilled into general (domain-independent) and task-specific skills. The graph uses typed edges like Prerequisite (A prereq → B), Enhances (A enhance → B), and Co-occurs. Nodes track usage statistics, which drive the system's evolution decisions.
- Graph Evolution
- This stage updates the skill graph during training by refining skills and relations based on success rates. It includes operations like merging, splitting, and pruning weak edges, while progressive unlocking activates higher-level skills when prerequisite performance thresholds are met.
Terminology used across episodes
This episode discusses
- SkillGraph: Skill-Augmented Reinforcement Learning for Agents via Evolving Skill Graphs · Paper Radio
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- From Local to Global: A Graph RAG Approach to Query-Focused Summarization
- Memp: Exploring Agent Procedural Memory
- GPT-4o System Card
- OpenAI o1 System Card
- SimpleMem: Efficient Lifelong Memory for LLM Agents
- ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- ZeroSearch: Incentivize the Search Capability of LLMs without Searching
- Agent KB: Leveraging Cross-Domain Experience for Agentic Problem Solving
- Mem- alpha: Learning Memory Construction via Reinforcement Learning
- EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle
- SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning
- Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward
- Qwen3 Technical Report
- MemEvolve: Meta-Evolution of Agent Memory Systems
- MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory
The paper
SkillGraph: Skill-Augmented Reinforcement Learning for Agents via Evolving Skill Graphs · Read on arXiv
University of Science and Technology of China
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "SkillGraph: Skill-Augmented Reinforcement Learning for Agents via Evolving Skill Graphs".
Tom: The gist The proposed framework, SKILLGRAPH, introduces a graph-structured formulation of skill library for LLM agents where skills are connected by explicit prerequisite, enhancement, and co-occurrence relations.
Jane: First, who's behind it and why it matters.
Paper summary: Tom: So wrapping up, SkillGraph proposes this closed-loop system where graph construction distills skills from trajectories, graph-aware retrieval produces ordered skill sequences for the policy, and then graph evolution refines the structure based on training feedback.
Jane: It’s about making the skill library dynamic rather than static. The authors emphasize that this structure allows for dependency-aware retrieval, which is what gets agents past those tricky multi-step problems they struggle with right now.
Meng: From an engineering standpoint, it means we design the learning process not just around the policy update, but around how the skill graph itself can be updated and pruned based on empirical success. It’s a more holistic way to manage agent knowledge.
Lu: The authors are showing that explicitly modeling prerequisite, enhancement, and co-occurrence relations gives agents a powerful mechanism for planning complex actions in an environment.
Lalam: For the future, this means agents can become much better at tasks that require chaining many small actions together because they aren't just guessing which skill comes next; they have a structure guiding them.
Tom: It really shows that for agents to tackle the harder, longer tasks out there, we need to give them a map of how their skills relate to each other, not just a giant pile of things they’ve seen before.
Conclusion: Tom: So SkillGraph is basically taking all those messy skills an agent learns while doing tasks and turning them into a structured map, right?
Jane: Yeah, that’s right. They’re building this graph where skills are connected by things like "this skill helps with that one," or "you need this before you can do that."
Lu: It’s really about giving the AI a memory structure instead of just a giant list of random actions.
Meng: So, the main idea is that when the agent learns something new, it doesn't just add it to a bag; it updates this graph.
Lalam: Because if a skill gets used successfully more often, or if its related skills get used together often, the connection between them gets stronger in this structure.
Tom: And they show that when you use this graph to pick skills for a task, you get an ordered list of skills that actually makes sense for the whole process.
Jane: It moves away from the agent just guessing what to do next and giving it a guided path based on its existing knowledge map.
Lu: The real power here is how they evolve that graph during training, so the agent keeps getting smarter by refining its own way of thinking skills.
Meng: From an engineering view, that means the learning isn't just about the policy changing; it’s also about restructuring the knowledge base itself based on what worked.
Lalam: It gives agents a way to learn from their successes and failures in a much more organized, traceable way than just trial and error.
Tom: And the results show it works really well across different tasks, proving this structured approach is actually effective for complex problem-solving.
More episodes
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck