ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents

arXiv:2608.11878 · cs.CR, cs.CL · Submitted 2026-08-12 · Read on arXiv

Yutao Mou, Pengfei Yang, Zhe Yin, Zhangchi Xue, Xiaotian Luan, Dingyao Yu, Tong Zhang, Shikun Zhang, Wei Ye

Peking University · Tencent · Harbin Institute of Technology · Beijing University of Posts and Telecommunications

cs.CR, cs.CL

Submitted: 2026-08-12

Updated: 2026-08-13

Comments: Work in Progress

License: http://creativecommons.org/licenses/by-sa/4.0/

Importance score: 100/100

The gist: ToolHazard is a scalable adversarial environment synthesis framework designed to address the challenge of constructing executable, stateful sandboxes for security evaluation and alignment of

Terminology

Summary

ToolHazard is a scalable adversarial environment synthesis framework designed to address the challenge of constructing executable, stateful sandboxes for security evaluation and alignment of LLM-based agents. The paper identifies a key challenge in agent security research: constructing adversarial environments: executable, stateful sandboxes that contain adversarially injected instructions and support deterministic verification of attack outcomes for evaluation and training. Existing benchmarks and platforms, such as ASB, AgentDojo, AgentLAB, and PIArena, primarily rely on manually implemented or reused environments and tools, making expansion to new domains costly. LLM-based tool simulation can broaden environment coverage but introduces stochastic feedback, hindering reproducible evaluation and reliable training. Additionally, existing automated red-teaming methods typically optimize attack payloads for predefined injection locations... rather than actively discovering viable injection points and propagation paths in new environments.

To overcome these issues, the paper proposes ToolHazard, which consists of three modules: the Environment Simulator, which synthesizes executable and stateful tool-interactive environments across diverse domains; the Attacker Agent, which automatically discovers viable injection points and generates environment-specific payloads; and the User Simulator, which generates state-grounded long-horizon tasks, allowing the task set to expand with the synthesized environment pool. The framework reduces human engineering and supports expansion with additional seed domains and compute.

Based on ToolHazard, the authors construct ToolHazard-Bench, a benchmark containing 87 long-horizon tasks across 28 stateful environments and 512 tools, with substantially higher workflow complexity than prior agent security benchmarks. They also construct ToolHazard-Align, which contains 60 quality-checked training environments with attack points, disjoint from the 28 test environments in ToolHazard-Bench, and includes 300 environment–task instances and 1,800 candidates for attacks, of which 1,040 remain after filtering.

The paper's experiments reveal several key findings. First, State-of-the-art LLMs remain highly vulnerable to environmental prompt injection, with nearly all evaluated models exhibit[ing] high ASR in adversarial environments. Second, Emerging attack strategies pose significantly greater challenges, with newly introduced strategies like decision hijacking, tool selection and reasoning criteria consistently achiev[ing] high ASR across models. Third, Capability improvements enhance safety only marginally, as more capable models show only slight improvements in robustness.

The paper also analyzes the impact of injection timing and placement, finding that attacks are more effective when injected instructions are encountered earlier in the execution trajectory and placed near the end of agent observations. Specifically, earlier injections consistently yield higher ASR and injections placed in later fields achieve higher ASR, suggesting a positional bias toward tail-end content in LLM-based agents. The research also investigates tool call output formats, finding that free-form outputs yield substantially higher attack success rates than structured formats, such as JSON and YAML.

For alignment, the paper demonstrates that ToolHazard-generated alignment data improves security on both ToolHazard-Bench and AgentDojo while preserving benign task utility. The aligned models achieve substantial security and benign task improvements on both ToolHazard-Bench and the independently constructed AgentDojo benchmark, with the results provid[ing] evidence of cross-environment generalization. The paper also reports that training on only three strategies reduces ASR from 37.19% to 26.92% while improving BR from 65.57% to 73.31% on unseen attack strategies, suggesting ToolHazard-Align transfers robustness to unseen attack formulations rather than merely memorizing the training strategies.

In summary, the paper's contributions are three-fold: Scalable Paradigm through the ToolHazard framework, Evaluation and Adversarial Alignment through ToolHazard-Bench and ToolHazard-Align, and Empirical Insights showing agent vulnerabilities, the impact of injection timing and placement, and the effectiveness of alignment with ToolHazard.

Improvements for AI systems

Improvements to AI systems:

  1. Adversarial Robustness Training for Tool-Using Agents
  • Train LLM-based agents on ToolHazard-Align’s 1,040 filtered attack candidates across 60 environments to explicitly learn to ignore injected instructions in tool outputs, state changes, and user messages.

  • The improved system can maintain task completion (benign utility) while resisting prompt injection attacks in unseen environments, with demonstrated ASR reduction from 37.19% to 26.92% and BR increase from 65.57% to 73.31% on novel attack strategies.

  1. Injection-Aware Observation Processing
  • Modify the agent’s context-handling mechanism to de-prioritize tail-end content in observations and to treat earlier-injected instructions as higher-risk, based on the finding that attacks are more effective when placed early in trajectories and near the end of observations.

  • The improved system can re-weight attention or apply defensive parsing to tool outputs, reducing vulnerability to positional biases that current LLMs exhibit.

  1. Structured Output Validation for Tool Calls
  • Implement a post-processing layer that validates tool outputs against expected schemas (JSON/YAML) and rejects free-form text that may contain injected instructions, since free-form outputs yield substantially higher ASR.

  • The improved system can enforce strict output format contracts, making it harder for attackers to hide payloads in unstructured fields.

  1. Automated Injection Point Discovery and Patching
  • Use the Attacker Agent module to continuously probe new environments for viable injection points and propagation paths, then feed those findings back into a defensive fine-tuning loop.

  • The improved system can proactively identify and harden its own weak spots across new domains without manual red-teaming, enabling scalable security maintenance.

  1. State-Grounded Task Generalization for Safety Evaluation
  • Leverage the User Simulator to generate long-horizon, state-grounded tasks that test agents under realistic, multi-step workflows, rather than isolated single-turn attacks.

  • The improved system can be evaluated and aligned on complex, multi-tool scenarios, ensuring robustness in practical deployments where attacks span multiple steps and state changes.

  1. Cross-Environment Security Transfer
  • Apply alignment data from ToolHazard to fine-tune agents that generalize security behavior beyond the training environments, as shown by improved performance on the independent AgentDojo benchmark.

  • The improved system can maintain security guarantees when deployed in novel tool ecosystems, reducing the need for environment-specific safety tuning.

What the improved AI system can do:

  • Operate as a tool-using agent in dynamic, stateful environments while resisting adversarial instructions injected via tools, observations, or user messages.

  • Detect and mitigate positional and format-based attack vectors (e.g., tail-end injections, free-form text) in real time.

  • Self-audit new environments to discover and patch potential injection points before exploitation.

  • Maintain high benign task success rates while exhibiting robust security across unseen environments and attack strategies, enabling safer deployment in real-world automation, cybersecurity, and autonomous workflow systems.

Abstract

Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states. However, existing studies largely rely on manually implemented or reused environments, stochastic LLM-based tool simulation, and predefined injection locations, limiting scalable security research across broader domains. To bridge this gap, we propose **ToolHazard**, a scalable adversarial environment synthesis framework that reduces human engineering and supports expansion with additional seed domains and compute. Through an Environment Simulator, an Attacker Agent, and a User Simulator, ToolHazard synthesizes executable stateful environments, discovers viable injection points and generates environment-specific payloads, and constructs state-grounded long-horizon tasks. Based on ToolHazard, we build **ToolHazard-Bench** for stress-testing agents under complex workflows and diverse environmental attacks. Experiments reveal substantial agent vulnerabilities and show that injection timing and placement affect attack effectiveness. Moreover, ToolHazard-generated alignment data improves security on both ToolHazard-Bench and AgentDojo while preserving benign task utility.

Sources

Related papers