StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing

arXiv:2608.24777 · cs.AI, cs.CR · Submitted 2026-08-25 · Read on arXiv

cs.AI, cs.CR

Submitted: 2026-08-25

Updated: 2026-08-25

Comments: Accepted by EMNLP 2026. Project page: https://zheng977.github.io/StepGuard/

Code: https://github.com/zheng977/StepGuard

Project page: https://zheng977.github.io/StepGuard

License: http://creativecommons.org/licenses/by/4.0/

The gist: LLM-based agents can interact with external environments through tool invocation, but this capability also introduces security risks such as file modification, information leakage, and unauthorized

Terminology

Abstract

LLM-based agents can interact with external environments through tool invocation, but this capability also introduces security risks such as file modification, information leakage, and unauthorized actions. Existing guardrails often evaluate completed trajectories, leaving pre-execution monitoring of step-level actions underexplored. We propose StepGuard, a step-level guard model that can audit completed agent trajectories and check tool actions before they are executed. To train StepGuard, we introduce StepGen, an automatic data engine that generates safe and unsafe trajectories with the same context but different actions at the risky step. To further reduce over-defense and under-defense, we propose Balance-GRPO, which dynamically balances learning between safe and unsafe actions based on their observed accuracy. Experiments show that StepGuard achieves the highest average accuracy among open-weight guard models, with performance comparable to GPT-5.4. When used to guard agents on AgentDojo and AgentDyn, StepGuard reduces mean attack success rate by 77.3% relative to the no-guard setting, while mean utility drops by only 2.8 percentage points.

Sources

Related papers