Graph-of-Skills: Dependency-Aware Structural Retrieval for Massive Agent Skills
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Graph-of-Skills: Dependency-Aware Structural Retrieval for Massive Agent Skills".
Jane: The paper was written by Author information is not present in the provided excerpt. from.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Summary: Tom: Okay, so we've talked about the structure, but what is the paper actually saying it *does*? We need to unpack the summary section of "Graph-of-Skills: Dependency-Aware Structural Retrieval for Massive Agent Skills."
Jane: If I understand correctly, the paper summarizes that traditional methods struggle when agent skills become overwhelmingly numerous and interconnected, leading to retrieval failure.
Lu: It’s about moving past simple keyword matching toward a deep understanding of *why* a skill is needed in context, which is what structural retrieval aims for.
Meng: The summary seems to pinpoint that the challenge isn't having enough skills, but managing the *interplay* between them when you scale up to millions of potential actions.
Lalam: From a societal perspective, if we can manage skill interaction this well, it means we could build agents capable of tackling global problems requiring highly diverse human expertise.
Tom: Right, so the core finding seems to be that you need a specialized mechanism—the structural retrieval part—to guide the agent through this massive space effectively.
Jane: It’s like giving a very smart student a comprehensive study guide that doesn't just list topics, but shows which chapter builds on another chapter in sequence.
Lu: And I think the implication here is that agent reasoning moves from being purely reactive to being proactively structural; it anticipates the necessary components.
Meng: The practical implication I see is for complex workflow automation—imagine a supply chain where every single step, from ordering materials to final delivery verification, is managed by interconnected agents.
Lalam: If we can model that complexity structurally, we could build better support systems for human workers in high-stakes environments like disaster response or advanced manufacturing.
Tom: So, the summary suggests that this structural approach solves the scalability bottleneck inherent in current massive agent designs.
Jane: It sounds like they've built a way for agents to efficiently navigate a knowledge space that is simply too big for older retrieval methods to handle accurately.
Lu: I’m particularly excited about how this framework inherently handles ambiguity by mapping out all valid structural paths, rather than just selecting the closest match.
Meng: From an implementation standpoint, the scalability of this graph structure must be addressed with memory efficiency; storing and traversing millions of dependencies is a massive computational lift.
Lalam: Ultimately, mastering this level of structured knowledge retrieval means we are moving closer to AI that doesn't just know facts, but understands the *architecture* required to solve problems.
Improvements: Tom: We’ve covered the structure and the summary; now let's look at what improvements "Graph-of-Skills: Dependency-Aware Structural Retrieval for Massive Agent Skills" suggests. Jane, what’s the big leap here?
Jane: It seems like they aren't just proposing a better graph; they are suggesting *how* to optimize the generation and utilization of that graph itself.
Lu: I think the paper moves toward making these graphs adaptive, meaning they don't have to be fully pre-built for every single potential scenario, which is computationally impossible today.
Meng: That adaptation aspect is where the engineering challenge really lies; how do you incrementally build and validate dependencies without requiring constant, massive retraining cycles?
Lalam: If we can make the skill graph adaptive and self-correcting in real-time, it suggests a level of emergent intelligence far beyond what we currently see in narrow AI applications.
Tom: So, the improvement isn't just retrieval; it's about making the *knowledge base* itself more dynamic and less brittle when faced with novel situations.
Jane: Exactly! It’s moving from a static map that you print out to a living, breathing subway map that updates as new lines are built or old ones close down.
Lu: My thought is this adaptation mechanism could allow agents to model abstract causal relationships—understanding *why* something might fail structurally, not just *that* it failed.
Meng: If we look at the engineering side of improvement, the efficiency of updating edge weights or adding new nodes based on observed failure modes would be critical for real-world deployment.
Lalam: The implication for human culture is that if our systems can learn dependency improvements autonomously, they become true collaborators rather than just complex tools needing constant human oversight.
Tom: So, the key improvement seems to be making the system robust enough to self-optimize its understanding of its own capabilities as it operates.
Jane: It's about closing the loop: use the skills, see what works or fails structurally, and then automatically update the dependency graph for next time.
Lu: I’m picturing a system that learns that Skill X is always better placed *after* Skill Y, even if some initial
Paper discussion segment 3: Tom: So just to recap what we discussed earlier: this paper moves beyond simply listing agent skills and starts building a true dependency map of those skills.
Jane: Exactly, Tom; instead of treating every capability as an isolated tool, they're showing us how those capabilities connect in meaningful chains.
Lu: And that shift from a flat list to a structured graph is huge because it allows the system to reason about *why* one skill must follow another for success.
Meng: That structural knowledge is what translates directly into massive efficiency gains; instead of trying every possible combination, the agent can prune unlikely paths immediately.
Lalam: It suggests that AI won't just be smart; it'll be fundamentally structured, allowing it to anticipate needs based on learned relationships between tasks.
Tom: I mean, if the system knows that detecting a flood *requires* downloading USGS data before running a risk model, that saves so much computational overhead!
Jane: It’s like teaching an apprentice not just what tools exist in the workshop, but which tool to grab first when they hear a specific alarm.
Lu: Precisely! You're moving from reactive problem-solving to proactive, guided workflow execution based on known dependencies.
Meng: From an implementation standpoint, this means we can design far more reliable agents that don't get stuck in infinite loops or simply fail because they missed a prerequisite step.
Lalam: Beyond the technical efficiency, think about how this improves human interaction; if an AI understands the structural needs of a complex request—say, booking a trip—it doesn't just throw back links; it guides you through the necessary sequence of choices.
Jane: That’s really powerful because it makes the AI feel less like a search engine and more like an actual helpful consultant who knows the whole process inside and out.
Tom: So, if we could model every single human workflow—from filing taxes to performing surgery—into these dependency graphs, the potential impact is staggering.
Lu: It opens up entire new dimensions for complex scientific modeling that currently requires immense amounts of human domain expertise to structure properly.
Meng: And it gives us a much clearer path toward making truly autonomous systems that operate in unpredictable, real-world environments without constant human oversight.
Lalam: This kind of structural intelligence elevates AI from being merely powerful to being genuinely wise, shaping a future where technology anticipates and facilitates complex human flourishing.
Jane: I wonder how this graph structure could be applied to something incredibly messy, like cross-cultural communication or understanding evolving social norms?
Conclusion: Tom: So, wrapping up our deep dive into "Graph-of-Skills: Dependency-Aware Structural Retrieval for Massive Agent Skills," it really seems like we're looking at a fundamental shift in how AI agents can actually use their knowledge.
Jane: Exactly, Tom. What I think is most exciting about this paper is that it moves beyond just giving agents a massive list of tools; it gives them the *map* of how those tools relate to each other.
Meng: That structural understanding, that dependency awareness, is key because simply having a big database of functions doesn't solve the problem if the agent tries to use two things in an order that makes no sense.
Lu: Right, Meng hit on something vital there; it's not just about retrieval, it's about *syntactic* and *semantic* correctness across complex workflows, which is a huge leap for general AI reasoning.
Lalam: And from the perspective of cultural impact, this means we could build systems that genuinely mimic collaborative human problem-solving rather than just executing isolated tasks.
Tom: Speaking of collaboration, Jane mentioned the 'map'—it really implies that future AI agents won't just be single-purpose tools; they’ll be entire, interconnected teams within a digital system.
Jane: It takes the guesswork out of massive agent deployment, letting developers focus on the high-level goal rather than painstakingly chaining every single step together.
Meng: From an implementation standpoint, I wonder how scalable this graph representation is when you move into millions of skills and agents? Is there computational overhead we need to worry about?
Lu: Well, if the underlying structure is optimized—if the graph traversal algorithms are smart—it should be manageable, because it's leveraging inherent relationships rather than brute-force searching.
Lalam: And thinking about how this improves culture, a more reliable and robust AI foundation like this allows humans to trust these systems with much higher stakes, letting us tackle grander societal challenges.
Tom: It’s definitely a monumental step forward in agentic architecture; I feel like we've only scratched the surface of what "Graph-of-Skills" means for the next decade of AI.
Jane: It makes you think about all the applications, doesn't it? From medicine to logistics, anything complex needs this kind of structured reasoning.
Lu: I just can’t imagine advanced AI systems operating without this level of formalized dependency mapping; it’s almost a prerequisite for true general intelligence.
Meng: It gives me hope that we can start building these real-world prototypes much faster now that the theoretical framework is so solid.
Lalam: Overall, "Graph-of-Skills: Dependency-Aware Structural Retrieval for Massive Agent Skills" really points toward a future where AI augments human potential in ways we could only dream of before.
Tom: Well, that's our time up for this paper; I think my brain is buzzing with ideas about what comes next.
Jane: We certainly are! Next time, we’ll be looking at something entirely different, but the theme of advanced reasoning seems to be taking over the AI world.
cs.AI
Submitted: 2026-04-07
Updated: 2026-09-11
Comments: Accepted to EMNLP 2026 (Main Conference). 26 pages (10 main + 16 appendix). Core contribution by Dawei Liu and Zongxia Li. Code: https://github.com/davidliuk/graph-of-skills
Code: https://github.com/davidliuk/graph-of-skills
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 82/100
The gist: The paper introduces "Graph-of-Skills" (GoS), a novel retrieval mechanism designed to improve agent performance in complex, multi-step tasks by understanding the structural dependencies between
Key concepts
- Structural Retrieval
- This method moves beyond simple keyword matching by understanding *why* a skill is needed in context. It guides an agent through a massive space of potential actions by mapping out valid structural paths, rather than just selecting the closest match.
- Dependency-Awareness
- This refers to the system's ability to recognize that skills are not isolated tools. Instead, it understands which capabilities must follow others in a meaningful sequence (e.g., detecting a flood requires downloading USGS data first).
- Graph-of-Skills
- A specialized mechanism that maps all potential agent skills and their necessary relationships into a structured graph. This allows the system to reason about complex workflows and efficiently prune unlikely paths.
- Adaptive Graph
- An improvement suggesting that the skill knowledge base does not need to be fully pre-built for every scenario. Instead, it can incrementally build and validate dependencies in real-time as new situations arise.
Terminology
Summary
The paper introduces Graph-of-Skills
(GoS), a novel retrieval mechanism designed to improve agent performance in complex, multi-step tasks by understanding the structural dependencies between available skills. GoS addresses the limitations of traditional skill retrieval systems by moving beyond mere topical overlap; its core value lies in presenting agents with a context that is already close to the executable decomposition of the task,
thereby significantly reducing search friction and making intended execution paths explicit earlier.
Structural Advantage Over Competitors
The main pattern observed across qualitative cases is not simply that GoS retrieves skills with better topical overlap. Instead, GoS more often exposes a bundle that is already close to the executable decomposition of the task.
This structural advantage is evident when comparing retrieval methods:
-
Vanilla Skills: While capable of finding relevant tools, they often underperform the tighter GoS bundle, as seen in the pedestrian-traffic-counting example.
-
Vector Skills: These can succeed when recovering a right family of skills, but their success is
most convincing when the retrieved bundle becomes qualitatively similar to what GoS surfaces directly.
In contrast, Vector Skills' successful episodes do not always arise from a bundle that is as semantically crisp or structurally interpretable as the one surfaced by GoS.
Mechanism of Value: Reducing Search Friction
The primary benefit of GoS is its ability to convert structural knowledge into immediate operational utility. The discussion highlights that the main value of GoS is not higher reward, but a shorter path from retrieval to execution.
This efficiency is demonstrated in several scenarios:
-
In the flood-risk-analysis example, when the correct chain was available to multiple methods, GoS mainly reduced search friction and made the intended execution path explicit earlier.
-
The dialogue-parser and dapt-intrusion-detection cases specifically show how GoS can convert that structural advantage into
clearer downstream wins.
Case Studies Illustrating Structural Improvement
The qualitative case studies provide concrete evidence of GoS's superior context assembly capabilities:
-
Economic Detrending and Correlation: Here, GoS surfaced the latent preprocessing step—timeseries-detrending—and converted the task into a full pass. This shows that surfacing the right latent preprocessing step materially changes the result, unlike other methods which assembled less coherent bundles.
-
3D Scan Calculation: This task serves as a control where all conditions can succeed by recovering the same latent geometry bottleneck. GoS specifically exposed
mesh-analysis together with adjacent geometric helpers, directly matching the preprocessing structure of the task.
-
Dependency-Aware Bundling: The table evidence shows that in tasks like energy-market-pricing, GoS provided a bundle containing
dc-power-flow,power-flow-data, andlocational marginal prices, which passed the verifier, whereas other methods showed less cohesive or more noisy contexts.
Limitations and Future Directions
The analysis also identifies boundary conditions where even GoS falls short of perfect execution. For instance, the earthquake-phase-association case shows a real boundary condition in which GoS still falls short of an execution-complete bundle.
Furthermore, the comparison with adaptive-cruise-control and energy-market-pricing indicates that even when retrieval is broadly adequate, trajectory efficiency and verifier alignment remain separate bottlenecks.
Taken together, these cases support the core claim that structural retrieval helps by presenting agents with a more execution-ready context.
Improvements for AI systems
1. Transition from Semantic-Only to Structural-Graph Retrieval
-
Improvement: Replace flat vector-based retrieval with a directed, multi-relational graph retrieval layer.
-
Capabilities: The AI will no longer just find
topically similar
tools; it will retrieveexecution-complete bundles.
For complex technical tasks (e.g., seismic analysis or financial modeling), the system will automatically pull the high-level solver along with the specific low-level parsers, data converters, and setup utilities required to actually run that solver, eliminating theprerequisite gap
that causes current agents to fail.
2. Implementation of Reverse-Aware Dependency Diffusion
-
Improvement: Integrate a Personalized PageRank (PPR) mechanism that utilizes reverse-traversal weights on dependency edges.
-
Capabilities: When a task query matches a high-level tool, the system can
back-propagate
through the graph to recover upstream prerequisites. This allows the AI to autonomously assemble a complete functional pipeline (from data ingestion to final output) even when the user's query only mentions the end goal, ensuring the retrieved context is structurally sufficient for execution.
3. Deterministic I/O-Based Dependency Induction
-
Improvement: Deploy an offline indexing pipeline that uses deterministic schema-matching (comparing output fields of one skill to input fields of another) to construct dependency edges.
-
Capabilities: The AI can scale to massive libraries (2,000+ skills) with a highly reliable structural backbone. This reduces reliance on expensive and hallucination-prone LLM-generated relationships, providing a mathematically grounded map of how tools actually connect in a real-world execution environment.
4. Hybrid Semantic-Lexical Seeding and Budgeted Hydration
-
Improvement: Implement a dual-signal seeding process (combining dense embeddings with lexical keyword matching) followed by a reranking and hydration step constrained by a strict context budget.
-
Capabilities: The AI can significantly reduce inference latency and token costs (achieving up to 56% reduction) while simultaneously increasing task success rates. It prevents
context overload
and thelost in the middle
phenomenon by providing a compact, highly relevant, and execution-ready payload rather than an unstructured, noisy list of tools.
Abstract
As LLM agents act across personal applications, web browsers, and other interfaces, their reusable skill libraries can scale to thousands of skills. This scale introduces two challenges. First, loading the full library saturates the context window, driving up token costs, hallucination, and latency. Second, semantic retrieval surfaces topically relevant skills but can miss upstream and downstream prerequisite skills, creating a prerequisite gap that leaves the retrieved bundle insufficient for execution. We present Graph-of-Skills (GoS), an inference-time structural retrieval layer for large skill libraries. GoS constructs an executable skill graph offline from skill packages, then retrieves a bounded, dependency-aware bundle through hybrid semantic-lexical seeding, reverse-aware Personalized PageRank, and context-budgeted hydration. Across SkillsBench and ALFWorld, with three model families (Claude Sonnet 4.5, MiniMax M2.7, and GPT-5.2 Codex), GoS attains the highest average reward in all six model-benchmark blocks, at a fraction of the token cost of loading the full library. On SkillsBench with GPT-5.2 Codex it raises average reward by 7.0 absolute points over full skill loading, a 25.6% relative gain, while cutting total tokens by 56.7%. Ablations isolate the mechanism: replacing reverse traversal with forward propagation costs 9.1 reward points, a larger loss than removing the graph altogether. The gain thus comes from traversing dependencies backwards, not from graph diffusion as such. A budget-matched retrieval study holding seeding, reranking, hydration, and context budget fixed reproduces the same ordering, with dependency-pair co-recovery falling from 0.654 to 0.362. Code is available at https://github.com/davidliuk/graph-of-skills
Sources
- HippoRAG: Neurobiologically Inspired Long-Term Memory for Large Language Models
- Organizing, Orchestrating, and Benchmarking Agent Skills at Ecosystem Scale
- API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
- MM-Zero: Self-Evolving Multi-Model Vision Language Models From Zero Data
- From Local to Global: A Graph RAG Approach to Query-Focused Summarization
- Self-Rewarding Vision-Language Model via Reasoning Decomposition
- SkillNet: Create, Evaluate, and Connect AI Skills
- ToolNet: Connecting Large Language Models with Massive Tools via Tool Graph
- ControlLLM: Augment Language Models with Tools by Searching on Graphs
- Retrieval Models Aren't Tool-Savvy: Benchmarking Tool Retrieval for Large Language Models
- Reflexion: Language Agents with Verbal Reinforcement Learning
- ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
- OpenAI GPT-5 System Card
- Dynamic Tool Dependency Retrieval for Lightweight Function Calling
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- Gorilla: Large Language Model Connected with Massive APIs
- On the Tool Manipulation Capability of Open-source Large Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection