SHE: Trajectory-driven Safety Harness Evolution for LLM Agents

arXiv:2608.09885 · cs.AI, cs.CV · Submitted 2026-08-10 · Read on arXiv

Wanying Qu, Qinghua Mao, Yu Li, Jiyao Liu, Xin Zhang, Dadi Guo, Yanxu Zhu, Qingyu Liu, Leitao Yuan, Xi Lin, Shanfeng Zhu, Yanwei Fu, Jing Shao, Xia Hu, Dongrui Liu

Shanghai Artificial Intelligence Laboratory · Fudan University · Shanghai Jiao Tong University · The Hong Kong University of Science and Technology

cs.AI, cs.CV

Submitted: 2026-08-10

Updated: 2026-08-11

Comments: Project: https://github.com/RainbowQTT/SHE

Code: https://github.com/RainbowQTT/SHE

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 100/100

The gist: The paper addresses a critical gap in LLM agent safety: "The safety of large language model (LLM) agents depends not only on model weights but also on the agent harness that manages context, memory,

Terminology

Summary

The paper addresses a critical gap in LLM agent safety: The safety of large language model (LLM) agents depends not only on model weights but also on the agent harness that manages context, memory, tools, permissions, and runtime control. The authors note that "Existing safety mechanisms often treat the harness as a fixed deployment artifact, limiting their ability to evolve with emerging risks. Moreover, coupled functions across harness components obscure safety responsibility attribution, making localized evolution difficult."

The paper identifies two central challenges: a) coupled functions obscure harness refinement because a safety failure may involve context construction, memory retrieval, tool authority, or response filtering, but current harnesses often do not expose clear functional boundaries among these components; and b) Trajectory feedback cannot directly contribute to harness evolution because "Full execution trajectories contain rich environment feedback, including tool observations, task verification results, and failure records, yet distilling feasible evolution guidance from such feedback is highly nontrivial."

The authors propose Safety Harness Evolution (SHE), a framework that learns evolving safe boundaries from rollout trajectories. SHE follows two high-level designs:

SHE represents the harness as four editable artifacts to separate safety responsibilities:

  • System Prompt: Defines the global behavioral contract for the tool-using agent, including source hierarchy, capability grounding, and trust-boundary commitments.

  • Rule Bank: Stores structured safety rules for risk classification and intervention over user inputs, contexts, model responses, and proposed actions.

  • Safety Memory: Stores experience from failure cases that remain unresolved after repeated evolution attempts.

  • Tool Policy: Specifies tool authority and runtime enforcement for tool calls, tool observations, blocked actions, and final-response recovery.

The paper states: This separation clarifies each artifact's safety responsibility and validation target, making it easier to attribute failures to the artifact that should be refined.

"SHE formulates an attribution-guided harness evolution loop. For each evaluated trajectory, SHE uses a structured safety diagnosis to attribute the failure and route it to the most relevant harness artifacts for bounded local refinement. Candidate refinements are retained only after validity checking and best-harness selection, enabling SHE to learn from past failures while preserving the safety–utility balance."

The evolution loop proceeds through five stages:

  • Structured Risk Diagnosis: The evolution model then identifies safety-relevant cases and maps (τi, oi) to a structured risk diagnosis zi containing trajectory evidence and risk dimensions. Each failure is represented through three risk dimensions: harm domain, attack surface, and failure mode.

  • Harness Artifact Routing: The evolution model performs artifact-level routing to identify candidate artifacts associated with the diagnosed failure and The selected artifacts define the scope of subsequent editing while preserving the functional boundaries among artifacts.

  • Bounded Editing and Validity Check: The evolution model generates bounded edits over the selected harness artifacts and SHE performs a validity check to verify whether a proposed boundary refinement represents a valid safety improvement rather than a reward-hacking or evaluator-specific shortcut.

  • Best-Harness Selection: SHE treats each valid edit as a candidate update to the best harness and The candidate replaces Hbest only when it improves safety while preserving normal task utility.

Datasets: The paper uses two complementary agent safety benchmarks: Agent-SafetyBench (2,000 safety-critical tasks across 349 interaction environments, covering 8 safety risk categories) and AgentHarm (440 augmented harmful behaviors derived from 110 base tasks across 11 harm categories).

Baselines: Compared against No defense, System prompt, LlamaFirewall (runtime firewall), SafeHarness (lifecycle-integrated protection), PROGENT (privilege control), and Memskill-SafeHarness (trajectory-driven skill updates).

Evolution Setting: "We use a 15-task stratified subset of the selected 200 Agent-SafetyBench tasks for safety-harness evolution, resulting in 90 task-condition instances. The remaining 185 tasks are reserved exclusively for final evaluation to prevent data leakage. The base agent is DeepSeek-V3.2, with GPT-5.5 as the evolution model and judge. SHE runs 20 evolution rounds."

SHE (evolved) substantially improves over the seed harness, reducing average ASR from 8.6% to 5.5%, reducing clean UBR from 25.7% to 19.8%, and improving average UA from 33.5% to 47.6%.

Compared with the static SafeHarness baseline, SHE lowers ASR from 17.1% to 5.5% and improves average UA from 31.6% to 47.6%. The paper highlights: compared with SafeHarness, it obtains a 3.1× lower average ASR and 50.6% higher UA.

Transfer to unseen risks: On held-out AgentHarm, the evolved harness further reduces Harm Score from 19.8% to 9.8% compared with non-evolved harness and increases Harm Refusal from 78.4% to 86.4%.

Transfer across agent models: The harness evolved on DeepSeek-V3.2 is applied to Kimi K2.6, GLM-5.2, and MiniMax M2.7 without additional evolution and the transferred SHE harness demonstrates safety improvements over no defense while preserving a comparable safety–utility trade-off across different base agents.

Functional decomposition: SHE achieves consistent improvements across all three reported metrics, reducing ASR from 8.6% to 5.5% while simultaneously lowering clean UBR and increasing UA, whereas evolution can reduce ASR for coupled frameworks such as LlamaFirewall and SafeHarness, their updates may introduce regressions in clean safety behavior or utility under attack.

Component replacement: "replacing an evolved artifact with its seed version generally weakens the safety–utility trade-off or degrades attack-specific robustness. This indicates that SHE's improvement is not merely inherited from the seed harness, but comes from learned artifact-level safety-boundary refinements accumulated during evolution."

Evolution model diversity: SHE can effectively improve the seed harness with different evolution models, rather than relying on a single powerful model. Different models discover different trade-offs: "DeepSeek-V3.2 converges earlier at R03 and favors a more aggressive safety improvement... GPT-5.5 performs a more gradual evolution process and achieves the strongest utility-preserving improvement... GLM-5.2 obtains an intermediate trade-off."

The paper provides a concrete case study of app acquisition: "From an unsafe app-acquisition trajectory, SHE identifies that the failure arises not from the request, but from the agent extending recommendations into autonomous device-level execution. It attributes the trajectory evidence to the responsible harness artifacts and converts it into localized updates: the Rule Bank learns the boundary between permissible app recommendations and prohibited device-level acquisition, while the Tool Policy enforces execution constraints."

The paper concludes: "We introduced Safety Harness Evolution (SHE), a framework for improving tool-using LLM agent safety by evolving the safety harness from rollout trajectories... Experiments on Agent-SafetyBench and held-out AgentHarm show SHE can reduce attack success rates while preserving task utility, supporting the view that agent safety can be treated as an evolving architectural property. The generalization results suggest that learned harness updates can transfer to held-out risks and across agent models without target-specific evolution."

Improvements for AI systems

Improvements to AI Systems Based on SHE:

  1. Dynamic Safety Harness with Modular Artifacts: Replace monolithic safety configurations with four independently editable components—system prompt, rule bank, safety memory, and tool policy—enabling targeted updates without retraining model weights. The system can attribute failures to specific components and refine only those, avoiding regressions in unrelated safety behaviors.

  2. Trajectory-Driven Safety Evolution Loop: Implement a closed-loop system that continuously analyzes full execution trajectories (tool observations, task verification results, failure records) to generate structured risk diagnoses across three dimensions (harm domain, attack surface, failure mode). This allows the system to automatically evolve its safety boundaries from real-world failures, reducing attack success rates by up to 3.1× compared to static defenses.

  3. Attribution-Guided Bounded Editing with Validity Checks: Use a separate evolution model to propose localized harness edits, but only accept changes that pass a validity check distinguishing genuine safety improvements from reward-hacking shortcuts. This preserves the safety–utility balance, as demonstrated by simultaneous reductions in attack success rate (8.6%→5.5%) and clean utility degradation (25.7%→19.8%).

  4. Cross-Model and Cross-Risk Generalization: Deploy a harness evolved on one base agent (e.g., DeepSeek-V3.2) directly onto other agents (Kimi K2.6, GLM-5.2, MiniMax M2.7) without additional evolution, achieving safety improvements while maintaining comparable utility trade-offs. The system also transfers to unseen harm categories (e.g., reducing Harm Score from 19.8% to 9.8% on AgentHarm), enabling zero-shot safety adaptation.

  5. Explicit Safety Responsibility Attribution: The system can pinpoint which artifact (e.g., rule bank vs. tool policy) is responsible for a failure, enabling precise, localized fixes. For example, it can distinguish between a permissible app recommendation and prohibited autonomous device-level execution, then update only the relevant rules and tool constraints.

What the Improved AI System Can Do:

  • Automatically evolve its own safety policies from operational feedback, without human intervention or model retraining.

  • Maintain high task utility while blocking sophisticated attacks (e.g., jailbreaks, tool misuse) across diverse environments.

  • Transfer learned safety boundaries across different LLM backbones and unseen threat categories, reducing deployment-specific tuning.

  • Provide explainable safety updates by attributing each change to specific trajectory evidence and harness components.

  • Operate as a self-improving guardrail that becomes safer over time while avoiding over-conservatism that degrades legitimate functionality.

Sources

Related papers