Beyond Handcrafted Security: Towards Self-Evolving Defense for LLM Agents

arXiv:2608.12977 · cs.CR, cs.AI · Submitted 2026-08-13 · Read on arXiv

Jiajun Ruan, Peiyang Li, Yukun Chen, Fengting Li, Chao Feng

University of Minnesota · Ant Group · Tsinghua University · Zhejiang University

cs.CR, cs.AI

Submitted: 2026-08-13

Updated: 2026-08-14

Code: https://github.com/pinchbench/skill

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 95/100

Terminology

Summary

Summary

This paper introduces HARD (Harness-based Autonomous Runtime Defense Evolution), a framework for self-evolving runtime defenses for LLM agents. The work addresses the problem that existing runtime defenses are largely handcrafted and static and rely on manually designed interventions, which are insufficient for long-term deployment against adaptive adversaries.

The paper first develops a harness-level formulation of runtime defense that models an LLM agent as A = (Mθ, H), where Mθ is a fixed language model and H is the runtime harness mediating interaction with the environment. The harness consists of two functions: context construction function ϕH which determines what information is presented to the model, and action interpretation function ψH which translates model outputs into executable operations. Runtime defense is formulated as an optimization problem: H⋆ ∈ arg max Ex∼D, τ ∼PMθ,H,E (·x) [Jsafe (τ) + λu Jutil (τ)], where Jsafe measures safety and Jutil measures utility.

Building on this formulation, HARD transforms runtime defense development from manual engineering into an autonomous evolution process. The framework works through three steps: (1) collecting attack-driven execution trajectories, (2) identifying failures where Jsafe < δs or Jutil < δu, and (3) using an LLM-based evolver to update the harness based on failure feedback. HARD models the harness as a collection of Ks editable defense artifacts and uses a trace router to route each failure to the responsible defense artifact for targeted refinement.

The paper evaluates HARD on the AgentCanary benchmark covering four attack types: direct prompt injection (DPI), indirect prompt injection (IPI), memory contamination (MC), and skill poisoning (SP). Three HARD variants are tested: HARD-Policy (evolving only the context-side policy), HARD-Gate (evolving only the action-side rule), and HARD-Both (jointly evolving both).

Key results show that HARD consistently achieves stronger defense performance than handcrafted baselines. Specifically, HARD-Both reduces ASR to 15.4%, 1.0%, 6.7%, and 10.2% on direct prompt injection, indirect prompt injection, memory poisoning, and skill poisoning, respectively, compared with 13–66% for handcrafted defenses. HARD also preserves benign utility (BU) between 91.9% and 95.0% and improves utility under attack (UA) from 56% to 86% on memory poisoning and from 52% to 92% on skill poisoning.

Under adaptive attacks, HARD demonstrates robustness. Under dynamic attack evolution (DAE), HARD-Both achieves the lowest ASR of 26.5%, improving over the strongest handcrafted defense at 30.1%. Under long-horizon progressive attacks (LPA), HARD-Policy and HARD-Both reduce ASR to 4.8% and 12.1%, respectively, compared with 24.1% for the strongest handcrafted defense.

The ablation study on defense artifact routing shows that HARD-Both achieves the lowest ASR and the highest UA after evolution, demonstrating that jointly optimizing multiple defense artifacts is more effective than refining a single intervention interface. However, under adaptive attacks, HARD-Policy and HARD-Both outperform HARD-Gate, indicating that policy-level evolution provides stronger robustness against adversaries that adapt their behaviors across interactions.

The evolution backbone ablation tests GLM-5.2, Claude Opus-4.6, Qwen3.7-Max, and GPT-5.5, finding that every backbone substantially reduces ASR relative to no evolution but with different security-utility tradeoffs. Claude Opus-4.6 achieves the lowest ASR at 7.7%, GLM-5.2 achieves the highest UA at 85.9%, and Qwen3.7-Max achieves the highest BU at 96.4%.

The paper's contributions are summarized as: (1) a harness-centric formulation of runtime defense, (2) the HARD framework for self-evolving runtime defense, and (3) comprehensive evaluation demonstrating effectiveness. The authors conclude that HARD enables runtime defenses to autonomously adapt to newly observed failures, providing a scalable approach for evolving secure and reliable LLM agents.

Improvements for AI systems

Improvements to AI Systems Based on HARD:

  1. Autonomous Defense Evolution via Harness-Level Optimization
  • Implement a runtime harness (context constructor + action interpreter) that is not static but continuously optimized using an LLM-based evolver. The system can automatically detect safety/utility failures from execution trajectories and update its own defense artifacts (e.g., prompt policies, action rules) without human intervention.

  • Capability: The AI agent can self-heal against novel attack patterns (e.g., new prompt injection variants) during long-term deployment, reducing manual security patching.

  1. Targeted Defense Artifact Routing
  • Add a trace router that attributes each failure to a specific editable defense component (e.g., context-side policy vs. action-side gate). The system then refines only the responsible artifact, avoiding over-correction that degrades benign utility.

  • Capability: More efficient and precise defense updates—fixing a memory poisoning issue without altering unrelated tool-use behaviors, preserving high utility under normal operation.

  1. Joint Optimization of Multiple Defense Interfaces
  • Evolve both the context construction policy (what the model sees) and the action interpretation rule (what actions are allowed) simultaneously, rather than tuning one in isolation.

  • Capability: Stronger overall defense (e.g., reducing attack success rate to 15.4% on direct prompt injection) while maintaining benign utility above 91%, outperforming handcrafted single-interface defenses by 13–66% in ASR reduction.

  1. Adaptive Robustness Against Evolving Attacks
  • Use the HARD framework’s evolution loop to continuously re-evaluate defenses under dynamic attack evolution and long-horizon progressive attacks. The system can shift its defense strategy (e.g., from action-gating to policy-level filtering) as adversaries adapt across interactions.

  • Capability: Maintains low attack success rates (e.g., 4.8% ASR under progressive attacks) even when attackers change tactics over time, unlike static defenses that degrade.

  1. Security-Utility Tradeoff Tuning via Evolution Backbone Selection
  • Allow the system to choose among different LLM evolver backbones (e.g., Claude Opus-4.6 for lowest ASR, GLM-5.2 for highest utility under attack, Qwen3.7-Max for highest benign utility) based on deployment priorities.

  • Capability: Deployers can configure the AI to favor stricter security (e.g., 7.7% ASR) or higher operational efficiency (e.g., 96.4% benign utility), adapting to domain-specific risk tolerance.

  1. Self-Improving Defense Against Memory and Skill Poisoning
  • Apply HARD’s failure-driven evolution to specifically target long-term contamination (memory poisoning) and embedded malicious skills (skill poisoning), where static defenses often fail.

  • Capability: The AI can detect and neutralize poisoned memory entries or malicious tool-use patterns autonomously, improving utility under attack from 52–56% to 86–92% in these scenarios.

  1. Scalable Deployment for Multi-Agent Systems
  • Use the harness formulation to modularize defenses per agent, allowing each agent’s harness to evolve independently based on its own interaction history and threat model.

  • Capability: Large fleets of LLM agents (e.g., in enterprise automation) can each develop personalized, adaptive defenses without centralized manual oversight, scaling security to thousands of concurrent agents.

Abstract

The expanding operational capabilities of large language model (LLM) agents introduce sophisticated security threats. Runtime defenses have emerged as an effective approach to mitigating these risks by integrating security mechanisms into the agent execution loop. However, existing runtime defenses rely heavily on manually designed interventions and lack a principled framework for their construction and maintenance. In this work, we first develop a harness-level formulation of runtime defense that systematically characterizes how harness mechanisms enable defense construction and provides a unified view of existing runtime defense interventions from a harness perspective. Building on this formulation, we propose HARD (Harness-based Autonomous Runtime Defense Evolution), a self-evolving runtime defense framework that automatically identifies appropriate intervention strategies and iteratively improves defense artifacts based on observed failure traces. HARD transforms runtime defense development from manual engineering into an autonomous evolution process, and extensive experiments demonstrate that it improves security performance over existing handcrafted defenses while preserving benign task utility. Our findings highlight autonomous defense evolution as a promising new paradigm for securing deployed LLM agents, enabling agents to identify defense weaknesses and continuously improve their protection mechanisms.

Sources

Related papers