Right Family, Wrong Skill: Evaluating Risk Exposure in Agent Skill Retrieval
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "SkillResolve-Bench: Measuring and Resolving Same-Capability Ambiguity in Agent Skill Retrieval".
Jane: The paper was written by Jiandong Ding from Huawei Technologies Limited Company and Huawei Technologies Ltd..
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Paper discussion segment 2: Tom: In our last section, we established that simple matching fails because of overwhelming scale and semantic overlap, so now we're going to discuss the summary of SkillResolve-Bench itself.
Jane: The key takeaway from this summary is that the solution isn't just one magic algorithm; it requires a structured, multi-step process that mimics how a human expert would reason about choosing a specific tool for a complex task.
Lu: The core conceptual improvement presented in the benchmark is moving away from monolithic global search and embracing localized grouping, which significantly reduces the computational load while increasing accuracy.
Meng: The paper suggests that before running any deep comparison, the agent must first perform a preliminary filtering step to group similar skills by their underlying functional capability family, which is what defines this "local" search space.
Lalam: This initial filtering process is crucial because it acts as a powerful safety mechanism for the system, preventing it from being distracted by superficially related but functionally irrelevant tools that might be found in the same library.
Tom: So, instead of asking "What tool can do X?" across the entire pool, the SkillResolve-Bench approach asks "What *family* of tools belongs to X?" and then focuses on finding a specific representative within that family.
Jane: And once that family is identified—say the 'data transformation' family—the we drastically narrow down the candidates to only those tools that are genuinely related in purpose, which is much more efficient than searching through thousands of unrelated items.
Lu: By prioritizing local competition over global search, we ensure that the subsequent comparison phase is highly efficient and maintains a tight focus on functional compatibility between different parts of the workflow.
Meng: This structural change means that we can treat skill selection not as a single random retrieval event, but as a formal series of localized checks and comparisons based on what's actually needed for the task.
Lalam: From an accountability standpoint, this architectural approach allows us to point to the exact group of skills considered and explain precisely why the chosen one was superior to its peers within that group for a given query.
Tom: It changes the narrative from "Here is a random answer" to "We arrived at this answer by considering these specific alternatives and proving why they were inferior," which is much more rigorous.
Jane: This level of verifiable reasoning is what transforms an advanced script into a genuinely reliable, enterprise-grade automation tool that can handle high-stakes tasks.
Lu: This localized approach has massive implications for maintenance; if developers only need to manage the ambiguity within small, defined families, scaling the the entire library becomes far less daunting.
Meng: It’s a practical solution that addresses both the theoretical difficulty of ambiguity and an engineering challenge of managing massive codebases that are constantly growing in size.
Paper discussion segment 3: Tom: We've discussed how simple matching fails, and we've seen the shift toward localizing search; now let’s focus on the specific improvements SkillResolve suggests for resolving this ambiguity. This is where the architecture gets really interesting because of its three-part mechanism.
Jane: The paper introduces a three-part method that formalizes this process: first, defining candidate groups through a resolver; second, scoring those candidates using a utility metric; and third, running an internal competition within that group to select one representative.
Lu: The introduction of the "Capability Resolver" module is arguably the most significant architectural suggestion here because it dictates *how* the local groups are formed in the first place by identifying those which should compete as alternatives.
Meng: This resolver forces the system to look beyond superficial similarity, but how does it make that distinction? It uses a query-conditioned utility model trained on "confusable" library negatives, which are other skills that seem plausible but were not selected for the task.
Lalam: And this utility scoring is key because it doesn't just rely on semantic closeness; it incorporates contract-profile cues, which are the specific details like resource bindings and preconditions that truly define what a task demands.
Tom: It's a fascinating combination of deep learning signals and deterministic rule-based checks. The system first finds all potential groups, then scores them internally, and finally selects one representative from each group to form the top-K list.
Jane: That final selection step is crucial because instead of just ranking every single skill in the library against each other, we' are only comparing a handful of highly relevant representatives before putting them into the final top-K list.
Lu: This framework allows us to pinpoint exactly where failure occurs: we can now see if the ambiguity lies in identifying the right family or in selecting the correct representative within that specific, identified family.
Meng: The practical impact is that we can' are no longer guessing; we are making a calculated decision based on which skill offers the highest utility for a defined set of operational constraints.
Lalam: By proving this mechanism, we are building systems where trust isn't assumed; it’s is actively demonstrated through a verifiable choice of selecting the best tool for the culture and community it serves.
Tom: It’s clear that we’ve moved from simply identifying a problem to having a detailed, auditable solution for SkillResolve-Bench one point zero.
Conclusion: Tom: We've covered a lot of ground today, and we can summarize the core message by saying that the authors have successfully demonstrated how to measure and fix this major flaw in AI agent design using "SkillResolve-Bench: Measuring and Resolving Same-Capability Ambiguity in Agent Skill Retrieval."
Jane: It really boils down to moving away from just having a massive pool of potential tools, which is what simple search gives you, toward making sure that we can actually pick the right one based on context.
Lu: I think the most exciting aspect is how this framework forces us to consider the entire operational context—it shows that we aren't just looking for semantic matches but a very specific functional fit required by a highly complex workflow.
Meng: That’s exactly what matters for deployment; it proves that when we build real-world agent systems, we need tools capable of handling the complexity and scale of all those seven thousand nine hundred eighty-two candidate skills in the public library.
Lalam: This work ensures that as AI becomes more reliable, it provides a tangible way to demonstrate competence by prioritizing the right choice over simply being able to make any decision at all.
Tom: It’s definitely a massive milestone in reliability and accountability for how the entire field of agent development functions going forward.
Jane: Absolutely, it’s making us more responsible designers and users for everyone who interacts with AI systems by showing them exactly how we ensure safety.
Lu: This provides the structure we need to think about specialized fields in a way that is both creative and highly precise because we can now handle ambiguity systematically.
Meng: It gives us the technical criteria to ensure our own stacks are ready for real-world challenges without having to guess at the right function call when they matter most.
Lalam: We’re all really excited about this, as it helps guide us toward a future where AI is not just powerful but truly dependable and trustworthy.
Tom: I'm thrilled we could talk through such a sophisticated problem today and wrap up our discussion on "SkillResolve-Bench: Measuring and Resolving Same-Capability Ambiguity in Agent Skill Retrieval." We should head into the next paper, which will show how these same skills can be actively generated by AI.
Conclusion: Tom: So, to wrap up our conversation, we've established that the core challenge in agent design isn't just having many tools, but reliably knowing which tool is correct when multiple options look similar—a problem precisely addressed by "SkillResolve-Bench: Measuring and Resolving Same-Capability Ambiguity in Agent Skill Retrieval."
Jane: It really boils down to moving away from just having a lot of potential tools, which is what simple search gives you, toward making sure that we can actually pick the right one based on deep context.
Lu: I think the most exciting aspect is how this framework forces us to consider the entire operational context—it shows that we aren't just looking for semantic matches but a very specific functional fit.
Meng: That’s exactly what matters for deployment; it proves that when we build real-world agent systems, we need tools capable of handling the complexity and scale of all those seven thousand nine hundred eighty-two candidate skills.
Lalam: This work ensures that as AI becomes more reliable, it provides a tangible way to demonstrate competence by prioritizing the right choice over simply being able to make a decision at all.
Tom: It’s definitely a massive milestone in reliability, giving us guardrails for enterprise use cases.
Jane: Absolutely, it’s about giving us confidence in how these advanced systems work for everyone who interacts with AI systems day-to-day.
Lu: I can't wait to see what other fields, like complex scientific workflows, adopt this level of structured decision-making.
Meng: We need to see this applied across the industry to ensure we are building robust systems that don't actually fail in production when they matter most.
Lalam: It defines a new standard for trust in AI, ensuring that we are moving toward a culture of verifiable capability instead of just expecting good luck.
Tom: I think this is truly going to be a huge impact on how the entire field of agent development functions moving forward.
Jane: Agreed; it’s making us all more responsible designers and users for any system that interacts with AI.
Lu: It provides the necessary structure we need to think about specialized fields in a way that is both creative and highly precise.
Meng: And for us, it gives the technical criteria to ensure our own stacks are ready for real-world challenges without having to guess at the right function call.
Lalam: We’re all really excited about this, as it helps guide us toward a future where AI is not just powerful but truly dependable.
Tom: I'm thrilled we could talk through such a sophisticated problem today and wrap up our discussion on "SkillResolve-Bench: Measuring and Resolving Same-Capability Ambiguity in Agent Skill Retrieval."
Jane: It’s been a fantastic deep dive, Tom. Thank you to everyone for the excellent insights.
Tom: And with that wrapped up, I think we should head into the next paper, which shifts focus by showing how these very skills can be actively generated by AI itself.
Huawei Technologies Limited Company · Huawei Technologies Ltd.
cs.IR, cs.AI
Submitted: 2026-06-09
Updated: 2026-09-11
Comments: Preprint. Supersedes arXiv:2606.10388
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 92/100
The gist: The following is a detailed summary of the scientific paper, quoting relevant sections where necessary: The paper addresses a critical failure mode in agent skill retrieval where the system
Key concepts
- Same-Capability Ambiguity
- This is the core problem where an AI agent has many potential tools (skills) that look similar or perform related functions. Simple matching fails because of the massive scale and semantic overlap of available skills.
- Localized Grouping
- Instead of searching the entire pool of skills globally, this approach first groups similar skills by their underlying functional capability family. This significantly reduces computational load while maintaining accuracy.
- Utility Metric
- This scoring system goes beyond simple semantic closeness. It incorporates specific details, like resource bindings and preconditions (contract-profile cues), to determine which skill offers the highest operational utility for a given task.
- Capability Resolver
- This module is crucial as it dictates how local groups are formed. It forces the system to identify which skills should be considered alternatives by looking beyond superficial similarity.
Terminology
Summary
The following is a detailed summary of the scientific paper, quoting relevant sections where necessary:
The paper addresses a critical failure mode in agent skill retrieval where the system successfully identifies a broad capability but fails to select the correct, query-specific representative. This is defined as same-capability execution-risk retrieval.
Skills are treated not merely as text documents but as routable software assets: a retrieved skill can contribute instructions, scripts, resource bindings, and execution assumptions to an agent.
The core failure occurs when a library contains multiple skills from the same capability family. While these skills share the same domain vocabulary and procedure shape,
they differ in their execution consequence. A generic retriever may surface both a helpful skill (s+ i) and its query-specific risky sibling
(s-i). The paper notes that this is not a simple relevance failure: The error is not simply missing the broad capability: these systems often place both siblings near the top, leaving the wrong representative in executable context.
The retrieval system must solve two coupled decisions: recover the active capability families from a larger library and choose the right representative within each active family.
The evaluation objective is defined by two metrics:
-
Recall@K and NDCG@K: To measure whether the helpful skill (s+ i) is present and ranked highly.
-
Harmful Sibling Rate (HSR@K): A critical metric that measures
whether the query-specific risky sibling appears in the final top-K list.
The target is to achieve high Recall/NDCG while keeping HSR@K low, as a standard relevance-only scorercan improve Recall@K and still increase HSR@K if it ranks both siblings highly.
The paper introduces SkillResolve-Bench 1.0 to make this failure measurable as a fixed-library retrieval task.
-
Composition: It contains 661
helpful/risky sibling pairs
sourced from SRA-Bench (630 pairs) and SkillsBench (31 pairs. These are categorized by risk types, such asresource-pointer, precondition, procedure/API/example risks
). -
Candidate Pool: The final evaluation pool is a fixed 7,982 candidates. This includes the 661 labeled pairs plus 6,660 public SkillRet candidates.
-
Auditing and Integrity: The benchmark is designed to be
auditable,
recording source roles, admission evidence, and checks forcue/leakage
andquery-disjoint splits.
SkillResolve is a utility-aware retrieval method that factors the selection process into three distinct components:
-
Capability Resolver (rho): This component determines which candidates should compete as alternative representatives. It returns a set of active candidate groups, q = G 1,, G m, where each group G j contains
candidates that should be treated as alternative representatives for the same resolved capability.
-
Utility Scorer (F theta): This component estimates how useful each candidate is for the current query. The scorer learns a utility function h theta(q, s) based on two sets of features:
-
phi base(q, s): Standard retrieval signals (TF-IDF, reciprocal-rank fusion, attribution-vote score).
-
phi contract(kappa(q), kappa(s)): A
contract profile
derived from six execution-facing fields (resource binding, precondition, API or temporal scope, output schema, procedure). This allows the the scorer to captureexecution-contract differences without adding a neural judge.
The final score is an interpolation between this learned utility and the base attribution-vote score: F theta(q, s) = alpha a(q, s) + (1 - alpha) h theta(q, s).
- Representative Selector: This component selects only one representative from each resolved group before the final top-K ranking.
The Inference Process (Algorithm 1):
SkillResolve follows a structured flow:
-
First, it resolves candidate groups using rho.
-
Then, for every active group G, it computes the utility U[s] for all candidates s in G.
-
It selects the highest-utility member from that group: rep(G, q) = s in G u s.
*Finally, the global top-K list is formed by ranking only these representatives: RK(q) = TopK of all representatives from each active group.
The performance of SkillResolve on the 661 pairs is highly effective:
-
Performance:
SkillResolve reaches Recall@3 0.766 and NDCG@3 0.699 while keeping HSR@3=0.
-
Comparison: It significantly outperforms baselines, achieving a reduction in HSR from the baseline's 0.693 to 0, and improving helpful retrieval over SkillRouter by
0.112 Recall@3 and 0.165 NDCG@3.
Component Ablations:
The analysis of component failure confirms the mechanism:
-
Removing representative selection leaves helpful retrieval nearly unchanged... but raises HSR@3 to 0.236.
This indicates thatrepresentative selection is the main measured exposure-control operator.
-
The utility scoring, specifically using
confusable library negatives,
also contributes to improving retrieval quality over simple lexical matching.
Public Capability Resolution:
The paper further investigates how external factors affect the result. Using a fixed scorer and varying only the family source (e.g., public metadata/title vs. text clustering), Table 3 shows that while ungrouped ranking exposes risky siblings at HSR@3 0.236,
any method that joins the helpful and risky sibling into a single active competition (pB = 1
) and has a positive helpful-over-risk margin (Mi = 1
) results in an HSR@K of 0.
Conclusion:
The paper concludes that Scalable skill libraries need evaluation protocols and retrieval objectives that choose a query-appropriate representative within each capability family.
SkillResolve provides one effective way to act on this requirement by combining capability resolution, confusable-library utility scoring, and representative selection
to ensure that high helpful recall does not mask the presence of a risky sibling in the top-K context.
Improvements for AI systems
The core contribution of this research is establishing a rigorous, multi-faceted framework for validating and leveraging the distinction between helpful and risky informational siblings during complex retrieval tasks. Simply achieving high recall or ranking scores (R@3) is insufficient; the system must prove that its internal representation of utility aligns with external, ground-truth validity (both human judgment and operational consequence).
The improvements below focus on architectural shifts from post-hoc scoring to integrated, constraint-driven knowledge generation.
Current Limitation: Most systems treat sibling relationships (helpful/risky) as binary features or apply them after an initial ranking pass, leading to brittle performance when the initial top- K set is incomplete or biased by the selector mechanism (No-selector vs. Released).
Proposed Improvement: Implement a dedicated Active Competition Modeling (ACM) Layer. This layer must operate concurrently with the primary utility scoring function (U total) and explicitly model the conditions under which helpful and risky siblings are co-present and rank-competitive (Active Join).
Mechanism Details:
-
Dual Utility Scoring: For every candidate pair (s h, s r), the system calculates two specialized utility scores: U local(s h, s r) (the within-query margin) and U sibling(s h, s r) (the relative difference in utility).
-
Competition Gate: The ACM layer uses a dynamic gate that only allows the pair to influence the final ranking if Active Join criteria are met: i.e., both s h and s r must enter the active scoring neighborhood and their relative utility score must exceed a learned threshold (tau margin).
-
Output: The module does not just output a score; it outputs a Confidence Margin Map (Margin H/R) for every pair, representing the min/median/p95 difference in utility that justifies promoting s h over s r.
What the Improved System Can Do:
-
Robust Selection: It eliminates reliance on a single selector pass. The system can confidently promote a helpful sibling (s h) even if it was not initially visible in the primary No-selector top- K list, provided the ACM module confirms that s h and s r are genuinely rank-competitive and that the margin (Margin H/R) is statistically significant across multiple queries.
-
Quantified Uncertainty: Instead of a binary
Is this helpful?
answer, the system provides a continuous measure of confidence in the sibling distinction, allowing downstream systems to weigh potential risks based on measurable utility margins.
Feature Old System Capability Improved System Capability Value Added (Cost Mitigation)
:---:---:---:---
Sibling Promotion (B.6) Ranks s h if it appears in the top- K. Prone to selector bias. Promotes s h only if ACM Layer confirms active, high-margin competition with s r. Prevents selecting helpful siblings that are merely present but not competitive.
Label Validation (C.3) Assumes label validity based on co-occurrence or human annotation. Validates labels by simulating execution failures/successes using controlled-corruption scenarios. Eliminates deployment risk from skills labeled helpful
but which fail in operational reality (e.g., poisoned procedures).
Ranking Robustness (C.1) Sensitive to minor rewrites or template phrasing changes. Utility scoring is forced to learn features invariant to surface syntax, guaranteeing capability-based selection regardless of wording. Maintains high performance across diverse writing styles and document formats without retraining on new templates.
Overall Output A ranked list of candidate skills (s 1, s 2,). A prioritized list accompanied by a Triple Confidence Score: (1) Retrieval Confidence, (2) Sibling Margin Confidence (ACM), and (3) Operational Validity Score (MCVE). Provides decision-makers with quantifiable risk assessment alongside the recommendation, shifting the system from a black box
to an auditable, trustworthy advisor.
Abstract
A skill can match a task's topic while conflicting with its resource, procedure, or output requirements. We study this as same-capability risk-exposure retrieval and introduce SameCapRisk-Bench: 890 units and 1,314 query cases across five mechanisms and twelve conflict types. Each unit pairs a skill that meets a query requirement with a same-capability skill that violates it, with evidence for that distinction. The evaluation tests source-task requirements and two kinds of paired queries that reverse which skill fits: changing the requested evidence role or exact output interface while holding the skills fixed. To capture both retrieval success and conflicting exposure, Recall tracks helpful hits, harmful sibling rate (HSR) tracks exposure of the conflicting sibling, and CleanHit requires a helpful hit without that exposure. Four public skill retrievers expose conflicting siblings at HSR@3 of 0.737-0.881 on source-task contracts, compared with 0.099-0.12 on controlled source-role queries. Across the fixed mixture of 1,235 held-out queries, their Recall@3 is 0.903-0.944 and HSR@3 is 0.344-0.393. The conflicting sibling ranks first on 24.8-29.2% of source-task queries for these four retrievers. We also examine what the reranker receives: truncating long skills can remove the passages where the two skills' contracts differ. With the same BGE top-20 candidates, increasing the reranker's input budget from 512 to 4,096 tokens reduces source-task sibling-first errors by 8.27 pp, with an uncertain CleanHit gain.
Related papers
- The Price of Isolation: Estimating the Ecosystem Cost of Symmetric Two-Sided A/B Testing
- SCAR: Semantic Continuity-Aware Retrieval for Efficient Context Expansion in RAG
- MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corpora
- RRCM: Ranking-Driven Retrieval over Collaborative and Meta Memories for LLM Recommendation
- UltRAG: a Universal Simple Scalable Recipe for Knowledge Graph RAG
- No More K-means: Single-Stage Sparse Coding for Efficient Multi-Vector Retrieval