Skill Following: Evaluating Actual Skill Use in Retrieval-Enabled LLM Agents

arXiv:2609.00549 · cs.CL · Submitted 2026-09-01 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Skill Following: Evaluating Actual Skill Use in Retrieval-Enabled LLM Agents".

Jane: The paper was written by the authors from.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Summary: Tom: Okay, so we covered the basic idea of "Skill Following" in the last segment, which is basically making sure agents use tools correctly. Now, let's talk about what the paper actually summarizes about this problem—the core findings.

Jane: The paper really dives into how current evaluation methods often fall short because they don't capture the complexity of *how* a skill is used, only that it was called at all.

Lu: They seem to introduce several benchmarks that specifically test for procedural correctness within tool use, which is far more rigorous than just passing an API call.

Meng: I was interested in their discussion of parameter misuse. They show examples where the LLM correctly identifies the skill but provides garbage data or calls it with the wrong type of input.

Tom: So it’s not just about *what* the AI knows, but how well it translates that knowledge into structured, actionable inputs for an external system?

Jane: That's right. They summarize that relying on simple output checks isn't enough; you have to check the internal reasoning path as well.

Lalam: From a systemic perspective, this means we need to build models that not only generate text but also generate structured, verifiable execution plans *before* calling the skill.

Meng: And they seem to point out that many existing frameworks assume perfect adherence, which is unrealistic for the variability of real-world data and prompts.

Lu: It forces us to view the agent's workflow as a series of constrained decision points, each one needing its own dedicated evaluation metric for adherence.

Tom: So we're moving from a single pass/fail test to a multi-stage gauntlet where every step is inspected?

Jane: Exactly. They show that simply retrieving relevant documents isn't the end; the subsequent steps of interpreting and acting on that information are where most failures occur, and this paper tackles those specific failure points.

Lalam: The implications here for user experience are massive. If agents can reliably follow complex, multi-step instructions using various skills, they become true digital co-workers rather than just clever chatbots.

Meng: I appreciate that they grounded their summary in measurable failure modes. It gives us engineers concrete targets to hit when we try to stabilize these systems for production use.

Lu: They are effectively defining a new standard of competence for

Paper discussion segment 2: Tom: So, if I’m wrapping up our thoughts on this paper's core finding, it boils down to how often these AI agents actually *use* the tools they find versus just pretending they know how.

Jane: Right, that’s the tricky part for listeners to grasp; it’s not enough for an agent model to just *know* a function exists in its library.

Meng: Exactly, because you can write a perfect prompt telling it, "Use Tool X," and the model might hallucinate using it without actually passing the right arguments or following the procedure steps correctly.

Tom: You mean they might just mention the tool as if they used it successfully, but fail at the execution?

Jane: Think of it like giving someone a cookbook full of recipes; knowing you have a recipe for bread doesn't mean you know how to knead or proof the dough, does it?

Lu: This is where I get wildly excited because if we can rigorously measure the gap between *knowing* and *doing*, we unlock entirely new levels of complex reasoning systems. We could build AI that truly reasons through multi-step scientific discovery, not just suggesting keywords for a Google search.

Meng: But Lu, bridging that gap between suggesting and doing requires massive fidelity in the tool definition itself; are we talking about simple API wrappers, or do we need to embed entire execution environments into the agent's workflow? That’s a huge engineering lift.

Lalam: Meng raises a vital point about fidelity, but from a cultural standpoint, this means that if we perfect *actual* skill use, AI won't just be an answer engine; it becomes an apprentice collaborator that can reliably handle complex procedures across trades—from medicine to electrical engineering—which elevates human capability across the board.

Jane: So, the paper shows us the *measurement* of competence, which is a huge step toward trusting these systems with real-world tasks beyond simple chatbots.

Tom: It’s moving us past 'can it talk about X?' to 'can it reliably *do* X?' which changes everything for deployment.

Lu: And imagine applying this methodology across diverse domains right now; we could instantly benchmark the reliability of any specialized AI agent, accelerating scientific breakthroughs exponentially.

Meng: Before we get there, though, how do we handle conflicting or outdated skills in a massive corporate knowledge base? That’s where the maintenance nightmare begins.

Lalam: Ultimately, mastering this skill-following evaluation framework means that our future relationship with AI shifts from one of mere assistance to one of genuine augmentation—empowering us to solve problems that were previously considered too intricate for a single human team. Now, I wonder what happens when we try applying these rigorous evaluation standards to ethical decision-making skills?

Paper discussion segment 3: Tom: So, if I’m summarizing what this whole "Skill Following" paper really drives home, it’s that just *having* a skill doesn't mean an AI agent actually knows how or when to use it correctly in a complex task.

Jane: Exactly, Tom; it shifts the focus away from just retrieving a fancy tool and towards verifying that the model actually executes the steps properly within its workflow.

Lu: That’s right! It suggests we need layers of validation, almost like having an expert peer reviewing the agent's entire plan before it runs any code, which opens up huge possibilities for automated scientific discovery.

Meng: But Lu, if you add that many review layers—the planning, the retrieval check, and then the secondary validation—aren't we just building complexity until the whole thing becomes too slow or brittle for real-time use?

Lalam: You’re hitting on a critical point of friction there, Meng; but I think thinking about it as pure speed misses something important about trust. If we can build reliable scaffolding that ensures correctness, that fundamentally changes how humans trust and collaborate with AI systems.

Tom: Speaking of trust, Jane, what does this mean for the average developer who isn't deep in agentic frameworks?

Jane: Well, think of it like writing a recipe; you can find a thousand great techniques online—like making sourdough or caramelizing onions—but if the recipe doesn't tell you *when* to add the salt or *how long* to let it rest, the whole dish falls apart.

Lu: It’s about operational rigor, Jane; we’re moving from 'here are all the tools' to 'here is how these tools must interact sequentially for this specific outcome.'

Meng: I agree with Lu that rigor is key, but practically speaking, we need standardized interfaces for these skills so that any developer can plug in a validated function without needing a custom overhaul every time.

Lalam: And from a cultural standpoint, if we nail those standards, it means the barrier to entry for building genuinely capable AI applications drops dramatically; it elevates the conversation from 'can AI *do* this?' to 'what amazing things can we *ask* AI to do?'

Conclusion: Tom: So, if I’m wrapping up this amazing discussion, it really boils down to how much we can trust these complex agents when they try to use external tools or skills.

Jane: Exactly, Tom. It’s not just about *giving* the agent a library of functions; it's about proving that the agent actually knows *how* and *when* to use them correctly in a real-world workflow.

Tom: Right, because we’ve seen firsthand how easily things can go sideways if the skill usage isn't rigorous or if the underlying logic is flawed.

Lu: What struck me most was how this work frames skill following as a measurable capability, not just a theoretical hope. It opens up such exciting avenues for complex reasoning systems that need to interact with specialized domains.

Meng: I agree with Lu; from an engineering standpoint, being able to quantitatively measure reliable skill use is huge because it gives us the benchmarks we need before deploying anything critical in production.

Lalam: And what I see beyond the metrics is how this advances the entire relationship between humans and AI—it suggests a future where our intelligence isn't just simulated, but augmented by verifiable, precise capabilities.

Jane: That’s a really insightful point, Lalam; it shifts the focus from *what* the AI knows to *what* it can reliably do with what it knows.

Tom: It makes us realize that building intelligent agents is going to be less about massive language models and more about robust orchestration layers that enforce best practices.

Lu: I think we're moving toward a system where these agents act less like encyclopedias and more like highly competent, but supervised, apprentices.

Meng: Supervised is the keyword there; we can’t afford any unexpected behavior when these systems are managing real-world processes.

Lalam: It fundamentally changes what 'assistance' means in the digital age; it means having a partner that follows instructions with near-perfect fidelity.

Jane: So, as we wrap up our discussion on "Skill Following: Evaluating Actual Skill Use in Retrieval-Enabled LLM Agents," the message is clear: reliability and verifiable skill usage are the next frontier for agent development.

Tom: It’s been a fantastic deep dive, team; thank you all so much for joining us today.

Jane: We're really looking forward to tackling another cutting-edge topic with all of you right after this quick break!

cs.CL

Submitted: 2026-09-01

Updated: 2026-09-01

Comments: Accepted to Findings of EMNLP 2026

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 82/100

The gist: The paper, "Skill Following: Evaluating Actual Skill Use in Retrieval-Enabled LLM Agents," addresses a critical gap in current AI evaluation by moving beyond mere task correctness to assess whether

Key concepts

Skill Following
This concept measures an AI agent's ability to reliably use external tools or skills in a complex task. It moves beyond simply knowing a function exists, requiring the agent to correctly execute all necessary steps within its workflow.
Procedural Correctness
This is a rigorous evaluation standard that tests *how* an AI uses a tool, rather than just checking if the tool was called. It ensures the agent follows the correct sequence of steps and provides structured, verifiable inputs for external systems.
Retrieval-Enabled LLM Agents
These are advanced AI models that use external information retrieval to inform their responses. The discussion emphasizes that simply retrieving documents is insufficient; agents must reliably interpret and act on that information using specific skills.

Terminology

Summary

The paper, Skill Following: Evaluating Actual Skill Use in Retrieval-Enabled LLM Agents, addresses a critical gap in current AI evaluation by moving beyond mere task correctness to assess whether Large Language Model (LLM) agents genuinely utilize retrieved knowledge. The authors argue that simply achieving a correct answer is insufficient; the core measure of capability must be the agent's ability to follow and integrate specific, reusable procedural skills from an external library. This framework is vital because it provides a granular understanding of how an agent solves problems, rather than just if it can solve them.

Skill Pool Definition

The evaluation relies on meticulously curated pools of procedural skills for specialized domains. For coding tasks, the pool comprises nine distinct skills detailed in Table 20, designed to cover common algorithmic patterns. These include foundational techniques such as binary-search-boundary for monotonic search spaces, counting-hash for frequency analysis, and dfs-iterative for cycle-safe graph traversal. Furthermore, the pool incorporates specialized skills like sliding-window-longest, which finds contiguous windows satisfying a monotonic predicate, and two-pointer-pair-sum, used to find patterns in sorted sequences.

For mathematical reasoning, the system utilizes a pool of eight procedural skills (Table 21) tailored for cross-domain replication. These skills guide agents through structured mathematical thought processes. Examples include algebraic-substitution, which involves simplifying expressions by substituting known values; case-analysis, which requires splitting problems into mutually exclusive branches (e.g., sign or parity); and modular-arithmetic, which governs work with remainders and cyclic patterns.

Agent Tool Integration and Prompting

The interaction between the LLM agent and the skill library is formalized through a structured tool schema, specifically the SEARCH SKILLS TOOL. This tool allows agents to search for reusable problem - solving artifacts using a natural language query. The schema permits filtering by artifact type—utility, procedure, or guard—enabling agents to target specific kinds of assistance, such as defensive checks.

The agent prompts enforce this structured interaction. For instance, the SE Agent Prompt explicitly instructs the model: "You have a search skills tool that searches a skill library. Use it if you think it could help solve the task." This contrasts with the SD Agent Prompt, which requires direct implementation without explicit tool calling.

Evaluation Metrics and Annotation Fidelity

The evaluation framework measures success using quantitative metrics derived from agent performance across various interventions (Table 19). Key metrics include Coverage, defined as the percentage of tasks for which at least one skill was returned, and Retrieval-Augmented Evaluation (RAE) scores, reported in percentage points. The authors caution that the complexity of the evaluation design means that The merged-1 intervention changes both skill granularity and retrieval coverage, so it is not interpreted as a pure granularity effect.

Crucially, the system employs a specialized annotator prompt (GPT-5.5 Annotator Prompt) to judge quality independently of task success. This annotator's sole function is to annotate whether an LLM answer substantively follows retrieved skill content, emphasizing that the annotation process must not infer from task correctness.

Model Panel Diversity

To ensure robustness, the experiments are conducted across a diverse panel of state-of-the-art models (Table 22). This panel includes major families such as Anthropic Claude, Google Gemini, Meta Llama, Mistral AI, OpenAI GPT, and Qwen. The model access type is noted to describe diversity only; all models were accessed through the same OpenAI-compatible experiment interface.

Improvements for AI systems

Based on this detailed blueprint of procedural skill decomposition, structured tool usage, and specialized annotation methods, the core improvement is shifting AI reasoning from retrieval-augmented generation (RAG) to Retrieval-Augmented Procedural Planning (RAPP).

The existing system components (the skill pool, the search skills tool schema, and the annotator prompt) are powerful but operate in isolation. My improvements focus on integrating these components into a mandatory, verifiable planning pipeline.

Here are the specific improvements I would implement and what the resulting AI system can achieve:


The Issue: The current skill pools (Table 20 and Table 21) treat skills as independent components. In reality, solving a complex problem requires a sequence of skills where the output of one skill feeds directly into the input constraints of the next.

The Improvement: I would build and integrate a Skill Dependency Graph (SDG) that maps high-level problem patterns (e.g., Finding Optimal Subset) to required sequences of granular skills, including their necessary input/output data types and transformation functions.

  • Example: A problem requiring both counting hash and sliding window longest is not solved by calling two separate tools; it requires the SDG to dictate: Input to (Apply counting-hash on window boundaries) to Output A to (Pass A to sliding-window-longest) to Final Result.

What the Improved System Can Do:

The system will move beyond mere suggestion. When presented with a complex task, it will first output a mandatory, multi-step execution plan derived from the SDG before writing any code. This plan acts as an auditable proof-of-concept for the solution path, ensuring logical coherence and preventing premature leaps in reasoning that lead to costly errors.

  1. Decomposition Phase (Mandatory): The agent must use search skills first, forcing the output of a set of required skills and their dependencies (guided by the SDG). It must justify why each skill is necessary.

  2. Execution Phase: The agent calls the retrieved utility functions/templates in sequence, modifying its internal state after each step.

  3. Verification Phase: The agent uses the guard checklists retrieved via search skills to perform self-correction checks on its own intermediate outputs before generating the final answer.

The Skill Trace must be a JSON object listing every skill used, the exact lines of code/logic derived from that skill, and a quantified measure of contribution (e.g., This skill solved 70% of the boundary condition checks).

This level of granular failure analysis is invaluable for high-stakes research validation.

Abstract

Large Language Model (LLM) agents increasingly rely on external skills, yet standard evaluations obscure whether retrieving these skills actually helps. Aggregate metrics often compare retrieved versus non-retrieved tasks, introducing severe selection bias and failing to isolate the true effect of skill use. To measure this actual-use capability-which we formalize as Skill Following (SF)-we introduce the Retrieval-Invoked Actual-Use Effect (RAE). RAE computes the same-task outcome difference between matched skill-enabled and skill-disabled executions, conditioned exclusively on tasks where the agent actively retrieved a skill. Evaluating 17 LLMs across coding and mathematical domains, we uncover a stark evaluation paradox: models frequently show positive aggregate retrieval lift but negative RAE. On MBPP+, multiple models that appear to benefit system-wide actually harm their own performance on the exact tasks where retrieval occurred. These findings demonstrate that aggregate averages can create a misleading illusion of tool-use proficiency, whereas RAE directly measures whether the retrieval-to-answer pipeline genuinely rescues more outcomes than it harms.

Sources

Related papers