AgentSpec: Customizable Runtime Enforcement for Safe and Reliable LLM Agents
cs.AI, cs.CL
Submitted: 2025-03-24
Updated: 2025-07-31
Comments: Accepted by the 48th IEEE/ACM International Conference on Software Engineering (ICSE 2026)
Journal ref: Proc. ICSE'26, pages 2938-2950. ACM, 2026
Code: https://github.com/haoyuwang99/AgentSpec
Project page: https://microsoft.github.io/autogen/stable//index.html
License: http://creativecommons.org/licenses/by/4.0/
The gist: Agents built on LLMs are increasingly deployed across diverse domains, automating complex decision-making and task execution.
Terminology
Abstract
Agents built on LLMs are increasingly deployed across diverse domains, automating complex decision-making and task execution. However, their autonomy introduces safety risks, including security vulnerabilities, legal violations, and unintended harmful actions. Existing mitigation methods, such as model-based safeguards and early enforcement strategies, fall short in robustness, interpretability, and adaptability. To address these challenges, we propose AgentSpec, a lightweight domain-specific language for specifying and enforcing runtime constraints on LLM agents. With AgentSpec, users define structured rules that incorporate triggers, predicates, and enforcement mechanisms, ensuring agents operate within predefined safety boundaries. We implement AgentSpec across multiple domains, including code execution, embodied agents, and autonomous driving, demonstrating its adaptability and effectiveness. Our evaluation shows that AgentSpec successfully prevents unsafe executions in over 90% of code agent cases, eliminates all hazardous actions in embodied agent tasks, and enforces 100% compliance by autonomous vehicles (AVs). Despite its strong safety guarantees, AgentSpec remains computationally lightweight, with overheads in milliseconds. By combining interpretability, modularity, and efficiency, AgentSpec provides a practical and scalable solution for enforcing LLM agent safety across diverse applications. We also automate the generation of rules using LLMs and assess their effectiveness. Our evaluation shows that the rules generated by OpenAI o1 achieve a precision of 95.56% and recall of 70.96% for embodied agents, successfully identify 87.26% of the risky code, and prevent AVs from breaking laws in 5 out of 8 scenarios.
Sources
- Safeguarding Large Language Models: A Survey
- LLM Multi-Agent Systems: Challenges and Open Problems
- Detecting Standard Violation Errors in Smart Contracts
- A Language Agent for Autonomous Driving
- From Prompt Injections to SQL Injection Attacks: How Protected is Your LLM-Integrated Web Application?
- Risks of AI Scientists: Prioritizing Safeguarding Over Autonomy
- $\mu$Drive: User-Controlled Autonomous Driving
- GuardAgent: Safeguard LLM Agents by a Guard Agent via Knowledge-Enabled Reasoning
- SafeAgentBench: A Benchmark for Safe Task Planning of Embodied LLM Agents
- Breaking Agents: Compromising Autonomous LLM Agents Through Malfunction Amplification
- Adversarial Representation Engineering: A General Model Editing Framework for Large Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection