Skill-Use: Can LLMs Actually Use Skills in Agentic Harnesses?
Jinyi Han, Yuanjian Xu, Ying Liao, Xinyi Wang, Zishang Jiang, Zixiang Di, Fanyang Lu, Zhichao Hu, Yanghua Xiao
cs.CL
Submitted: 2026-08-05
Code: https://github.com/JinyiHan99/Skill-Use-Bench
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language model (LLM) agents increasingly rely on skills, structured documents that specify when to act, which procedure to follow, and which tools are allowed.
Terminology
Abstract
Large language model (LLM) agents increasingly rely on skills, structured documents that specify when to act, which procedure to follow, and which tools are allowed. Existing evaluations mostly judge the quality of a skill or its contribution to task success, leaving unexamined whether an agent can recognize a relevant skill and apply it on its own. We introduce Skill-Use, a benchmark that evaluates skill use under progressive disclosure, where an agent sees only a skill's name and short description and must retrieve the full procedure before following it. Skill-Use separates three facets of skill use. Trigger measures whether the agent invokes the relevant skill, Compliance measures how faithfully it follows the prescribed procedure, and Boundary measures whether it avoids forbidden operations. A Skill-Use (SU) score combines the three and credits execution only after the skill is triggered. Skill-Use pairs 79 real skills with 177 executable tasks across nine domains, each grounded in real files, run in an isolated Docker sandbox, and scored by a trajectory-based rubric. Evaluating eight LLMs under two agent harnesses, we find that reliable skill use remains out of reach, as the strongest configuration reaches an SU of only 0.613. Triggering and procedural compliance fail as independent bottlenecks, and both scores and model rankings shift with the harness, so skill use behaves as a capability conditioned on the harness rather than a fixed property of the model.
Sources
- Large Language Model Agent: A Survey on Methodology, Applications and Challenges
- Skill-R1: Agent Skill Evolution via Reinforcement Learning
- Agent Skills: A Data-Driven Analysis of Claude Skills for Extending Large Language Model Functionality
- SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
- SkillLearnBench: Benchmarking Continual Learning Methods for Agent Skill Generation on Real-World Tasks
- Organizing, Orchestrating, and Benchmarking Agent Skills at Ecosystem Scale
- SkillNet: Create, Evaluate, and Connect AI Skills
- AutoSkill: Experience-Driven Lifelong Learning via Skill Self-Evolution
- Memento-Skills: Let Agents Design Agents
- SLBench: Evaluating How LLM Agents Follow Logical Relations in Skills
- Instruction-Following Evaluation for Large Language Models
- Stop Comparing LLM Agents Without Disclosing the Harness
- Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows
- Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following
- SOPBench: Evaluating Language Agents at Following Standard Operating Procedures and Constraints
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering