Auto-Policy, not Auto-Skill: Compiled Agent Skills for the Physical World
cs.AI, cs.CR
Submitted: 2026-08-25
Updated: 2026-08-25
Comments: Presented at the 1st Workshop on Agent Skills (Agent Skills '26), ACM CAIS 2026, San Jose, May 26, 2026
Code: https://github.com/midudev/autoskills
License: http://creativecommons.org/licenses/by/4.0/
The gist: Self-evolving Skill harnesses (AutoSkills, Hermes Agent) generate more advisory orchestration automatically; their reported gains are efficiency, not safety.
Terminology
Abstract
Self-evolving Skill harnesses (AutoSkills, Hermes Agent) generate more advisory orchestration automatically; their reported gains are efficiency, not safety. This misses the actual gap: a Skill describes how an agent should behave; a Policy decides which behavior is allowed to become an action. Today's format covers the first with markdown and scripts; the second is left to the model. Generating more Skills scales the gap, not the safety, especially when a wrong invocation can unlock a door or move money. Two adjacent attacks are documented: malicious skills compromising cloud software, and jailbroken LLM-controlled robots causing physical harm. Their intersection, malicious agent skills causing physical harm, follows directly but has not been reported. We name this class Borrowed Authority: Skills format gives the receiving agent no typed way to reject an inter-agent permission claim, so a malicious or misused Skill can drive actuation by attaching one. We propose Edge Skillguard, a typed authority layer that lives inside the Skill artifact rather than between tools as workflow engines do, with guards over world state and sensor evidence. On a live edge control-plane testbed, the guards reject 60/60 borrowed-authority requests across five attack variants without blocking benign requests, and the result holds at 5x scale and across hosts over a Tailscale mesh. These results suggest that high-risk Skills should co-package typed invocation policy with procedural knowledge, so that physical actions depend on machine-checkable evidence rather than peer-agent claims.
Sources
- Optimizing Agentic Workflows using Meta-tools
- LLMind 2.0: Distributed IoT Automation with Natural Language M2M Communication and Lightweight LLM Agents
- SkillSieve: A Hierarchical Triage Framework for Detecting Malicious AI Agent Skills
- SkillJect: Effectively Automating Skill-Based Prompt Injection for Skill-Enabled Agents
- Proactive Rejection and Grounded Execution: A Dual-Stage Intent Analysis Paradigm for Safe and Efficient AIoT Smart Homes
- SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
- "Do Not Mention This to the User": Detecting and Understanding Malicious Agent Skills in the Wild
- Blueprint First, Model Second: A Framework for Deterministic LLM Workflow
- SAGE: Smart home Agent with Grounded Execution
- Leveraging LLMs for Efficient and Personalized Smart Home Automation
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection