StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing
cs.AI, cs.CR
Submitted: 2026-08-25
Updated: 2026-08-25
Comments: Accepted by EMNLP 2026. Project page: https://zheng977.github.io/StepGuard/
Code: https://github.com/zheng977/StepGuard
Project page: https://zheng977.github.io/StepGuard
License: http://creativecommons.org/licenses/by/4.0/
The gist: LLM-based agents can interact with external environments through tool invocation, but this capability also introduces security risks such as file modification, information leakage, and unauthorized
Terminology
Abstract
LLM-based agents can interact with external environments through tool invocation, but this capability also introduces security risks such as file modification, information leakage, and unauthorized actions. Existing guardrails often evaluate completed trajectories, leaving pre-execution monitoring of step-level actions underexplored. We propose StepGuard, a step-level guard model that can audit completed agent trajectories and check tool actions before they are executed. To train StepGuard, we introduce StepGen, an automatic data engine that generates safe and unsafe trajectories with the same context but different actions at the risky step. To further reduce over-defense and under-defense, we propose Balance-GRPO, which dynamically balances learning between safe and unsafe actions based on their observed accuracy. Experiments show that StepGuard achieves the highest average accuracy among open-weight guard models, with performance comparable to GPT-5.4. When used to guard agents on AgentDojo and AgentDyn, StepGuard reduces mean attack success rate by 77.3% relative to the no-guard setting, while mean utility drops by only 2.8 percentage points.
Sources
- Constitutional AI: Harmlessness from AI Feedback
- Meta-SecAlign: Training LLMs against Prompt Injection for Robust Agents
- Geometrically-Constrained Agent for Spatial Reasoning
- OR-Bench: An Over-Refusal Benchmark for Large Language Models
- Safe RLHF: Safe Reinforcement Learning from Human Feedback
- AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
- A Safety and Security Framework for Real-World Agentic Systems
- Are Your Agents Upward Deceivers?
- Building a Foundational Guardrail for General Agentic Systems via Synthetic Data
- SciHazard: A Benchmark for Measuring Scientific Safety Risks with Decomposed Harm Scoring
- AgentDyn: Are Your Agent Security Defenses Deployable in Real-World Dynamic Environments?
- Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents
- ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis
- AgentDoG 1.5: A Lightweight and Scalable Alignment Framework for AI Agent Safety and Security
- AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security
- SafeHarbor: Defining Precise Decision Boundaries via Hierarchical Memory-Augmented Guardrail for LLM Agent Safety
- The Llama 3 Herd of Models
- AgentAuditor: Human-Level Safety and Security Evaluation for LLM Agents
- Large Language Model Agent: A Survey on Methodology, Applications and Challenges
- AGrail: A Lifelong Agent Guardrail with Effective and Adaptive Safety Detection
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection