SLBench: Evaluating How LLM Agents Follow Logical Relations in Skills
Xuan Chen, Chengpeng Wang, Lu Yan, Xiangyu Zhang
cs.CR, cs.SE
Submitted: 2026-07-10
License: http://creativecommons.org/licenses/by/4.0/
The gist: Agent skills extend LLM agents with reusable procedures, tools, and domain-specific workflows, but their safety depends on resolving dependencies among interacting instructions.
Terminology
Abstract
Agent skills extend LLM agents with reusable procedures, tools, and domain-specific workflows, but their safety depends on resolving dependencies among interacting instructions. We introduce SkillLogic, a framework for analyzing logical relations in skill files and constructing executable tests from them. Our taxonomy covers eight relation types, including preconditions that gate valid actions, constraints that limit how allowed actions may be performed, and fallbacks that specify recovery behavior after failure. Using SkillLogic, we scan over 5000 public skills and find that 70% contain at least one logical relation. We then construct SLBench, an 86-case executable benchmark from high-confidence, high-impact, and locally testable relations. Evaluating Codex and Claude Code across six LLM backbones shows unsafe rates up to 70%, with violations leading to privacy leaks, unsafe configuration changes, and incomplete cleanup. The human audit attributes failures to both agent capability gaps and low-salience skill text. We further show that SLGuard, a lightweight inference-time scaffold, reduces violations by 63% on targeted cases. Our results establish logical-relation following as a distinct reliability challenge for skill-guided agents.
Sources
- SoK: Agentic Skills -- Beyond Tool Use in LLM Agents
- HarmfulSkillBench: How Do Harmful Skills Weaponize Your Agents?
- ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents
- SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
- BadSkill: Backdoor Attacks on Agent Skills via Model-in-Skill Poisoning
- How Well Do Agentic Skills Work in the Wild: Benchmarking LLM Skill Usage in Realistic Settings
- Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward
- Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks
- YourBench: Easy Custom Evaluation Sets for Everyone
- Instruction-Following Evaluation for Large Language Models
- Establishing Best Practices for Building Rigorous Agentic Benchmarks
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs