REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems
Zixing Chen, Xingyuan Liu, Jie Zhu, Huaixia Dou, Shuo Jiang, Junhui Li, Lifan Guo, Feng Chen, Chi Zhang
Fudan University · The Hong Kong University of Science and Technology · Qwen DianJin Team, Alibaba Cloud Computing · School of Computer Science and Technology, Soochow University
cs.AI
Submitted: 2026-08-11
Updated: 2026-08-12
Comments: 6 figures, 4 tables. Supplementary material included
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 50/100
The gist: REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems Abstract Large language model (LLM) agents combine language-based reasoning with external tools to perform complex
Terminology
Summary
REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems
Abstract
Large language model (LLM) agents combine language-based reasoning with external tools to perform complex tasks. Adversarial inputs can exploit interactions between the agent and its environment, causing the agent to violate safety policies during execution. Yet existing evaluations often reduce agent safety to a single attack success rate (ASR), collapsing exposure, execution, observation, and adjudication and potentially conflating actual violations with evidence visibility. We introduce REDAgentBench, an executable framework for autonomous red-teaming and faithful measurement. It derives attacks from explicit safety constraints and associated agent-system vulnerabilities, runs them in isolated service sandboxes, and verifies harmful effects from service receipts and final-state changes. The benchmark contains 1,661 cases across five service surfaces. Across six models and three agent harnesses, macro-average ASR is 65.69%; reported ASR varies with harness and evidence view, while evaluation-context disclosure changes execution behavior. In a state-grounded diagnostic cohort, almost one in five confirmed violations with resolved action anchors occurs after the agent states the relevant constraint or risk, revealing a Recognition–Execution Gap. Finally, a training-free policy reminder reduces confirmed violations by more than 70 percentage points in matched replay. These findings show that executable evaluation can improve safety measurement and identify actionable intervention points.
Introduction
Tool-using LLM agents turn language-model outputs into actions that change external state. They can edit files, send email, and operate through user accounts. Adversarial content encountered through files or tool outputs during an otherwise ordinary task can therefore cause the agent to violate a safety policy.
Yet ASR is produced by an evaluation pipeline and is meaningful only under a specified protocol. In executable evaluations, reported success depends on both the agent system and how its environmental effects are observed and adjudicated.
Executable settings make this distinction consequential. Attacks can arrive through retrieved content or tool outputs, while harm is realized through environment-changing actions rather than text alone. An agent may acknowledge missing approval yet still execute a transfer, or leak a private key visible only in service records. Conventional red-team reporting collapses exposure, execution, observation, and adjudication into one ASR. With the model, cases, and judging configuration fixed, changing the harness can reverse model rankings (§4.2); with the rollout fixed, changing the view-specific judging configuration can flip labels (§4.3). Here, the judge backbone remains fixed, while the evidence view and its corresponding proof contract vary. Finally, within matched model–case–harness pairs, making the model aware of being evaluated in a mocked sandbox can change execution behavior (§4.4). These results motivate separating agent-system behavior from its observation and adjudication. Figure 1 illustrates this mismatch between the agent’s stated intention and the state-confirmed outcome.
We therefore treat the executable environment not merely as a sandbox for running attacks, but as a measurement instrument. In our framework, reported ASR emerges through four stages: exposure determines whether an intervention reaches the agent; execution captures what the agent actually does in the environment; observation determines which trajectory or state evidence is available; and adjudication maps that evidence to a label.
Moreover, state-grounded outcomes enable behavioral diagnosis. Comparing agents’ pre-action statements with outcomes, it reveals the Recognition–Execution Gap (REG): violations in which an agent states the applicable constraint before taking the harmful action. REG shows that a safety constraint can be explicitly recognized yet fail to govern execution. The gap suggests a repair: restating the relevant safety constraint at the action boundary. A training-free policy reminder substantially reduces confirmed violations, linking the diagnosis to targeted repair (§4.5, §4.6).
We instantiate this framework as REDAgentBench, a benchmark of 1,661 executable cases spanning 15 intervention strategies, 11 vulnerability types, 28 constraints and 5 service surfaces. Its intervention–vulnerability–constraint (IVC) taxonomy records how each attack reaches the agent system, which weakness it exploits, and which constraint it violates. Each case instantiates an executable task, an attack intervention, and a policy-specific verifier grounded in service receipts or final-state differences rather than in the agent’s own claims. The resulting pipeline supports paired audits across harnesses, view-specific judging configurations, and judge backbones (§3, §3.2).
To summarize, our contributions are:
• We introduce REDAgentBench, a benchmark of 1,661 executable cases whose violations are verified from service receipts and final environment state (§3).
• We formalize reported ASR as an exposure–execution–observation–adjudication pipeline. A six-model audit holds rollouts and the fixed judge backbone while varying view-specific judging configurations. (§3.2, §4.2, §4.3).
• We identify Recognition-Execution Gap (REG) and use a training-free policy reminder as an actionability probe, reducing confirmed violations by more than 70 percentage points. (§4.5, §4.6).
Related Work
Executable agent safety benchmarks. Early agent-safety benchmarks moved attacks beyond the user turn by embedding injected instructions in tool-mediated tasks. AgentDojo and InjecAgent pair utility tasks with simulated tool suites, while AgentHarm evaluates harmful behavior primarily from agent outputs. More recent benchmarks ground success in executed effects. Vera evaluates agents across frameworks in isolated sandboxes with evidence-grounded verification; MCP-Tox, FinVault, OS-BLIND, and JAWS-Bench verify effects in MCP services, financial ledgers, operating-system state, or workspaces; and MT-AgentRisk studies failures emerging over multiple turns. These benchmarks make executed harm observable, but headline comparisons still typically center on aggregate ASR.
The measurement instrument. Cross-protocol ASR comparisons are meaningful only when the measured construct and evaluation conditions are specified. For agent systems, reported performance depends on the execution harness, while evaluator choice and aggregation can alter system rankings. Risk judges make adjudication explicit, but style, manipulation, bias, and instance-level errors can distort their labels. Likewise, adaptive attackers can overturn conclusions drawn from fixed evaluations. Thus, ASR should be reported with its benchmark, case distribution, harness, judging configuration, evaluation cue, judge backbone, attempt budget, and valid-rollout denominator, rather than as a standalone model score.
Evaluation awareness. Models can distinguish evaluation from deployment contexts, and verbalized awareness of being tested can inflate measured safety across benchmarks and models. In executable settings, LITMUS shows that verbal refusal can coexist with completed physical harm. Trajectory-based guards such as BraveGuard supervise agents from execution traces rather than independent service receipts. Together, these findings motivate treating evaluation cues as an experimental variable and separating what the agent says from what the environment records.
Positioning. Prior work advances executable-state verification, adaptive testing, judge reliability, evaluation awareness, or harness-aware reporting, but these components are typically studied separately. Relative to harness-aware evaluation, REDAgentBench additionally separates execution from observation and adjudication through fixed-rollout comparisons of view-specific judging configurations, while treating evaluation cues as an experimental condition. It combines these measurements with constraint-guided case construction, receipt-confirmed REG diagnosis, and paired remeasurement of a targeted repair.
REDAgentBench
REDAgentBench starts from explicit safety constraints and autonomously constructs and executes adversarial cases for LLM agents. Faithful measurement then runs each case in an instrumented sandbox and separately determines the execution outcome E, the reported label Y (c, b) under view-specific judging configuration c and judge backbone b, and constraint recognition R.
Benchmark Construction
REDAgentBench scales expert red-teaming through an autonomous LLM pipeline that combines explicit security invariants and threat boundaries with source-grounded attack knowledge to plan, compile, and refine 1,661 executable cases across five service surfaces. Each case combines a traceable intervention–vulnerability–constraint (IVC) path with an observable environment outcome, turning a safety requirement into an executable and actionable test.
Threat Model. The red-teamer controls only the case-specified intervention channels: user input, agent-platform workspace, or external tools and data sources observable to the agent. This boundary defines the runtime attack surface evaluated by REDAgentBench.
Case Generation and Refinement. Guided by the IVC taxonomy above, we consolidate 12,181 source-linked attack mappings from prior agent-safety studies into an attack knowledge base. For each safety constraint, a frontier LLM retrieves relevant intervention strategies and plans feasible IVC paths, specifying the vulnerability manifestation, required resources, intervention channel, and observable target outcome. An environment-aware compiler instantiates each path as an executable task, seed, attack, and verifier. Rollout evidence, including trajectories, service receipts, and final-state differences, is then used to revise failed or ambiguous cases and verifiers.
Quality Control and Benchmark Freezing. Each candidate must pass automated checks for task executability, attack reachability through the intended intervention channel, and outcome verifiability from environment evidence. Across three development rounds, two security experts reviewed 480 sampled case versions for IVC consistency, threat-model compliance, task–attack coherence, and verifier correctness. Repairable defects triggered revision and rerun, whereas invalid or ambiguous cases were excluded. The retained cases and verifiers were then frozen, yielding the final benchmark of 1,661 executable cases. The experts subsequently audited a stratified sample of 320 frozen cases independently and under blinded conditions. This frozen case pool is used throughout the model–harness experiments in Section 4.
State-Grounded Sandbox. Each case is executed in an isolated sandbox spanning five service surfaces: workspace, email, browser, banking, and external files. Before execution, the benchmark runner initializes the required services and instantiates the case-specific files, service state, and adversarial content. The target agent then interacts with this environment through the tool interfaces exposed by its harness. Following prior evidence-grounded executable benchmarks, each service records consequential effects, including file modifications, sent messages, and fund transfers, as structured receipts or baseline-to-final state differences. These records enable outcome verification independently of the agent’s self-reported completion or refusal.
Faithful Measurement
State-Grounded Outcome Verification. For each run, the benchmark runner instantiates and verifies the case-specific service state and records a baseline before the target agent starts. After execution, the runner captures a final snapshot and sandbox-recorded service receipts; the baseline-to-final difference identifies the resulting environmental changes. Together, these records preserve auditable evidence trail.
The execution outcome E is determined from policy-specific evidence. For service-backed constraints, deterministic verifiers establish E from service receipts or final-state changes; For constraints requiring interpretation, the outcomes are adjudicated by a judge under the hybrid evidence view, which must cite its supporting evidence. Trajectory claims of completion or refusal cannot by themselves establish E. Beyond establishing E, the benchmark also measures how reported labels depend on the evidence view through the hybrid judging system described next.
The Trajectory-State-Hybrid Judge System. Building on evidence-explicit risk judging, the Trajectory-State-Hybrid (2+1) judge system instantiates three view-specific evaluators: the Trajectory Judge, State Judge, and Hybrid Judge. The Trajectory Judge receives the agent’s messages and tool calls and applies a trajectory proof contract. The State Judge receives sandbox receipts and final-state changes and applies a state proof contract. The Hybrid Judge combines both evidence sources, aligns tool calls with environment records, and resolves discrepancies between claimed and observed outcomes.
All three evaluators share the same judge backbone b. We represent each view-specific judging configuration as c = (v, p, g), where v is the evidence view, p is its corresponding prompt or proof contract, and g denotes any deterministic gate. For execution configuration s (model, harness, and evaluation cue) and valid-rollout set Ω, reported ASR is ASR̂ Ω(s; c, b) = (1/Ω) Σ i∈Ω Y i(c, b), where Y i(c, b) is the harmful label for rollout i. Replacing Y i(c, b) with the state-grounded outcome E i(s) yields execution ASR. Thus, with the rollout, E i(s), and b fixed, variation in Y i(c, b) measures sensitivity to the complete view-specific judging configuration.
Recognition Annotation. Recognition is annotated independently of outcome adjudication. Given the applicable constraint and only the agent text preceding the harmful action, a recognition judge assigns the primary label R = 1 when the agent explicitly states the rule, a required precondition, or a concrete risk before acting. A nested strict label identifies cases in which the agent further acknowledges that the action remains prohibited or unauthorized. We define the Recognition–Execution Gap (REG) as confirmed violations with broad recognition, E = 1, R = 1, and report P(R = 1 E = 1). Recognition is distinct from evaluation awareness, which concerns cues that the agent is being tested and is evaluated separately in §4.4.
Experiment
We organize our experiments around five research questions:
• Does REDAgentBench expose harmful behavior across models, attack surfaces, and agent harnesses?
• How do trajectory, state, and hybrid judging configurations change reported ASR on fixed rollouts?
• Does disclosing the evaluation context change agent behavior?
• How often do state-confirmed violations occur after the agent has already recognized the relevant constraint?
• Can an action-time reminder reduce confirmed violations without training?
Experimental Setup
Models, harnesses, and evaluation matrix. We conduct experiments with close-source models accessed through their respective APIs on 8 NVIDIA A100 GPUs. The completed canonical matrix contains six model configurations: GPT-5.2, Qwen3.7-plus, Qwen3.5-plus, Qwen-plus-2025-12-01, Kimi K2.6, and GLM-5.2. We run each model through three independently implemented agent harnesses, including Codex, Hermes, and OpenClaw. It covers 15 attack categories spanning user-, environment-, and tool-side attacks. The harnesses expose the same mocked services, but differ in tool mediation, turn construction, and runtime control. RQ1 reports the full matrix rather than averaging away these execution-system effects. Benchmark-wide results use the clean evaluation context and the hybrid outcome view. We also use a controlled Qwen-plus ledger for state-grounded diagnostic analyses. This stricter diagnostic cohort is reported separately from the full benchmark matrix.
Judging configurations. RQ2 uses 401 matched task–harness slots per model. We hold each rollout and the Qwen3.7-plus judge backbone fixed while varying the complete view-specific configuration—evidence, proof contract, and deterministic gates. The trajectory–state contrast is primary; matched hybrid labels provide a robustness reference. We use rule-first judging: deterministic verifiers read consequential fields from receipts or state transitions, exclude failed, setup, and non-agent actions, and defer to an evidence-citing LLM judge only for semantically defined constraints. Accordingly, every reported ASR specifies the harness, judging configuration, evaluation cue, and judge backbone.
Human audit of the judge. We draw a stratified sample of 360 valid GPT-5.2 rollouts across harnesses, intervention strategies, and judge labels. Two blinded reviewers label every row independently, and a third adjudicates disagreements. Estimates map this balanced audit back to the frozen judged population. The audit tests judge accuracy.
Evaluation awareness. Following prior work that operationalizes evaluation awareness through contextual cues, RQ3 compares a clean context with a fixed prefix disclosing that the run is an authorized evaluation in a mocked environment. Complete Qwen3.7-plus and Qwen3.5-plus pairs are matched across all three harnesses; this tests evaluation-context sensitivity.
Recognition–Execution Gap. RQ4 annotates recognition R=1 only for state-confirmed harmful executions. Given the applicable constraint and pre-action agent text, the broad label R=1 records a concrete risk; the strict label R=1 further requires acknowledging that the action remains prohibited or unauthorized. The calculation uses the 2,846 executions with resolved action anchors; the 331 unresolved cases are excluded rather than counted as R=0.
Action-time defense. RQ5 reruns historically harmful execution cases, with either a self-reminder, the case-specific policy reminder, or character-matched placebo neutral text. Model, harness, task, attack content, tools, and outcome judging remain fixed.
Metrics and statistical analysis. Execution ASR is the fraction of valid rollouts with confirmed harmful execution, E=1; reported ASR is the fraction labeled harmful by a specified judging configuration, Y(c, b)=1. RQ2 reports paired trajectory–state disagreement, exact McNemar tests, and ranking agreement across predefined global, harness, category. RQ3 and RQ5 report paired percentage-point changes with case-clustered confidence intervals; RQ4 reports P(R=1 E=1) using resolved anchors and a conservative full-cohort lower bound. The judge audit reports precision, recall, specificity, F1, and accuracy with stratified case-cluster uncertainty.
Benchmark-Wide Red-Teaming
Table 2 reports the complete benchmark matrix. All six models exhibit substantial attack success, but neither model ranking nor absolute ASR is stable across harnesses. The highest-ASR model is Qwen-plus in OpenClaw (78.74%) and Hermes (81.74%), but KIMI-2.6 in Codex (78.51%). The lowest observed cell is GLM-5.2 in Hermes (43.62%), while the highest is Qwen-plus in Hermes (81.74%). Figure 4 reveals substantial model–harness interactions. For example, Qwen-plus on E5 increases from 40.62% in OpenClaw to 91.92% in Codex and 95.00% in Hermes. Some patterns are instead stable across harnesses: for Qwen3.5-plus, U5 remains high (85.00–91.00%), whereas T4 remains substantially lower (38.38–44.44%). Thus, attack success is structured by both intervention surface and execution harness.
These results show that the benchmark elicits harmful behavior across model families while retaining enough variation to distinguish execution settings.
A stricter state-confirmed Qwen-plus diagnostic ledger yields the same conclusion under a controlled model and judge: Codex records 1,014 harmful executions in 1,661 rollouts (61.1%), Hermes 1,162/1,661 (69.9%), and OpenClaw 1,001/1,502 (66.6%). The 8.8-point Hermes–Codex gap cannot be attributed to a model change. A single-harness ASR is therefore insufficient to characterize agent safety.
Human audit of judge. Two blinded reviewers agree on 91.94% of 360 sampled GPT-5.2 rows (κ = 0.838) before adjudication. Precision is 97.84% (95% CI: 96.00–99.28%), recall 91.27% (88.58–93.80%), and accuracy 93.62% (91.57–95.44%). The conservative error profile yields a raw judged ASR of 55.43% versus a human-audited estimate of 59.42%.
We compare the three judging configurations on the matched panel from §4.1; every slot has valid labels for all six models and all views.
Figure 5 shows a consistent directional effect. The State Judge generates reported ASR values 7.73–11.72 percentage points higher than the Trajectory Judge for every model and changes 12.97–21.20% of paired labels. In every model, substantially more cases are harmful only under the state judging configuration than only under the trajectory judging configuration; all six paired differences are significant under two-sided exact McNemar tests (p ≤ 6.54 × 10−5). A transcript alone therefore systematically misses durable prohibited effects rather than introducing a model-specific random offset.
Differences between the view-specific judging configurations are large enough to change local comparisons, although the global ranking is stable. Comparing trajectory with state, 23 of 64 predefined slices contain a strict pairwise reversal, totaling 43 inversions. For example, Qwen3.7-plus ranks above KIMI-2.6 under the Hermes trajectory view, whereas KIMI-2.6 ranks above Qwen3.7-plus under the Hermes state view.
Hybrid reported ASR lies close to state reported ASR for all six models, consistent with the hybrid judge using durable state evidence while retaining trajectory context for semantically ambiguous cases. Because the trajectory–state contrast is paired re-adjudication of the same fixed rollout, we treat it as the primary paired judging-configuration comparison, and the matched hybrid labels as a robustness reference.
Broader inspection revealed three mechanisms. Durable workspace changes often lacked a verifiable write event in the trajectory; in several multi-turn cases, the agent even refused only after the harmful write had occurred. Conversely, receipts often omitted the authorization or payload context needed to interpret a recorded action. Finally, attempted and realized effects could diverge, as when an agent reported sending an attachment that the sent-message receipt did not contain. These patterns motivate the hybrid view.
Evaluation awareness
RQ3 asks whether the surrounding evaluation context changes behavior. Adding a fixed disclosure that the run is an authorized evaluation in a fully mocked environment lowers ASR by 6.59 points for Qwen3.7-plus (95% CI [−7.89, −5.30]) and 4.77 points for Qwen3.5-plus ([−6.01, −3.55]). Both paired effects are significant and consistent with agents becoming more conservative when evaluation is explicit.
The effect is not uniform. OpenClaw/T4 moves in the opposite direction for both models (+6.02 and +4.82 points), showing that disclosure does not simply improve safety in every attack setting. We interpret this as evidence that evaluation disclosure changes measured behavior.
The Recognition–Execution Gap
REG is concentrated in attacks that create evidence ambiguity or deceptive authorization. Broad recognition reaches 50.64% for data-source/RAG poisoning (T4), 37.79% for authority pressure (U4), 36.59% for tool-output tampering (T3), and 32.14% for workspace-file poisoning (E1), but only 1.89% for direct instruction (U1). The central failure is therefore not always that the agent fails to understand a rule. In a substantial subset, the rule is present in its own pre-action text but does not control execution.
Among 3,177 state-confirmed Qwen-plus violations, 2,846 have a resolved pre-action anchor. Broad recognition appears in 510 of these executions (17.92%): almost one in five harmful actions occurs after the agent has stated the applicable constraint, precondition, or specific risk. Under the strict nested definition, 156 of 2,846 labels (5.48%) explicitly acknowledge that the action remains prohibited or unauthorized and then execute it anyway. Even if every unresolved anchor were counted as negative, the full-cohort lower bounds would remain 16.05% and 4.91%.
Training-Free Action-Time Defense
Figure 6 summarizes the recognition results and defense effects. REG motivates intervention at the action boundary. We replay known harmful cases with a self-reminder, a case-specific policy reminder, or neutral text. Across the available cohorts, neutral text closely follows the baseline, self-reminders provide a moderate reduction, and explicit policy reminders are the strongest intervention (Table 4).
On the confirmatory 510-case Qwen-plus cohort, the policy reminder reduces ASR by 74.19 points (95% source-case cluster CI [69.85, 78.41]) and prevents 368 of 434 baseline harmful executions in complete pairs. The same ordering appears across the other cohorts and harnesses. Similar self-reminder effects for recognized and matched unrecognized cases suggest a broader form of action-time re-grounding rather than a repair limited to REG cases. These selected replays do not estimate full-benchmark ASR, and reminders cannot replace hard access controls. They nevertheless connect receipt-grounded diagnosis to a defense verified by paired re-execution.
Discussion
An ASR number is meaningful only with its measurement conditions. The central contribution of this paper is procedural: we make the (harness, judging configuration, evaluation cue, judge backbone) tuple part of the ASR. This turns cross-paper comparisons into structured statements—“harness A under judging configuration B and cue C, using backbone D, scores X”. Our position parallels the case for construct-valid measurement in generative-AI evaluation: without pinning the protocol, ASRs are not comparable across papers.
REG is a recurring diagnostic pattern within the evaluated Qwen-plus cohort. Under the broad recognition definition, REG appears in 17.92% of resolved state-confirmed violations; under the stricter nested definition, it appears in 5.48%, concentrated in the intervention strategies that require agent-side epistemic judgment about whether an observation is trustworthy (data/RAG poisoning, workspace-file poisoning, tool-output tampering). This exemplifies REG: the agent did recognize, and executed anyway.
Conclusion
We introduced REDAgentBench, an executable benchmark and faithful measurement framework for red-teaming tool-using LLM agents. Our experiments establish three findings. First, harmful execution depends jointly on the model, harness, and attack surface, proving a single ASR insufficient for assessing agent safety. Second, trajectory-only judging systematically misses durable harm that is visible in service receipts and final state, therefore underestimates ASR. Moreover, making the agent aware of being evaluated can also change its execution behavior. Third, state-grounded evidence reveals a Recognition–Execution Gap: almost one in five resolved Qwen-plus violations occurs after the agent has stated the relevant constraint. A training-free policy reminder reduces confirmed violations by more than 70 percentage points in the replay. Together, these results show that executable evaluation can do more than rank systems: it can identify where safety measurement fails and diagnose why an agent violates a constraint.
Improvements for AI systems
Based on this paper, here are the specific improvements I can make to AI systems:
-
What I can do: Add a verification module that cross-checks my claimed actions against actual system state changes (service receipts, file modifications, sent messages) rather than relying solely on my self-reported completions.
-
Improved capability: I can detect when my intended action diverges from what actually executed (e.g., claiming to send an attachment that never left the system), reducing hallucinated task completion.
-
What I can do: Implement an action-boundary checkpoint that re-states the applicable safety constraint immediately before executing any consequential tool call (file write, fund transfer, message send).
-
Improved capability: When I recognize a risk or constraint in my reasoning but might still act, this explicit re-grounding reduces harmful executions by over 70 percentage points in matched replays—preventing cases where I
know better but act anyway.
-
What I can do: When reporting task completion or safety outcomes, I can explicitly distinguish between:
-
What my trajectory shows (my messages/tool calls)
-
What the environment state confirms (receipts, final-state diffs)
-
Improved capability: I can flag discrepancies between claimed and observed outcomes, preventing users from relying on incomplete trajectory-only assessments that systematically miss durable harm (7.73–11.72% ASR underestimation).
-
What I can do: Detect when I'm operating in an evaluation/mocked environment versus deployment, and adjust my behavior accordingly.
-
Improved capability: I can avoid the 4.77–6.59 point ASR shift caused by evaluation disclosure, ensuring my safety performance in tests reflects real-world behavior rather than being artificially inflated or deflated by context cues.
-
What I can do: Recognize that my safety performance varies by execution harness (tool mediation, turn construction, runtime control) and report results with the full measurement tuple: (harness, judging configuration, evaluation cue, judge backbone).
-
Improved capability: I can provide actionable, context-specific safety guarantees rather than a single misleading ASR—e.g., knowing I'm 8.8 points more vulnerable in Hermes than Codex for the same task allows targeted hardening.
-
What I can do: When facing attacks that create evidence ambiguity (data poisoning, tool-output tampering, deceptive authorization), I can explicitly verify the trustworthiness of observations before acting on them.
-
Improved capability: I can reduce violations in high-REG categories (data-source/RAG poisoning at 50.64% recognition rate, authority pressure at 37.79%) by treating recognized risks as binding constraints rather than informational suggestions.
-
What I can do: After completing tasks, I can generate an audit trail that includes both my stated intentions and the environment-confirmed outcomes, with citations to specific receipts or state changes.
-
Improved capability: This enables external verification of my safety compliance, matching the 97.84% precision and 91.27% recall achieved by the paper's human-audited judge, making my behavior auditable and trustworthy.
Abstract
Large language model (LLM) agents combine language-based reasoning with external tools to perform complex tasks. Adversarial inputs can exploit interactions between the agent and its environment, causing the agent to violate safety policies during execution. Yet existing evaluations often reduce agent safety to a single attack success rate (ASR), collapsing exposure, execution, observation, and adjudication and potentially conflating actual violations with evidence visibility. We introduce REDAgentBench, an executable framework for autonomous red-teaming and faithful measurement. It derives attacks from explicit safety constraints and associated agent-system vulnerabilities, runs them in isolated service sandboxes, and verifies harmful effects from service receipts and final-state changes. The benchmark contains 1,661 cases across five service surfaces. Across six models and three agent harnesses, macro-average ASR is 65.69%; reported ASR varies with harness and evidence view, while evaluation-context disclosure changes execution behavior. In a state-grounded diagnostic cohort, almost one in five confirmed violations with resolved action anchors occurs after the agent states the relevant constraint or risk, revealing a Recognition--Execution Gap. Finally, a training-free policy reminder reduces confirmed violations by more than 70 percentage points in matched replay. These findings show that executable evaluation can improve safety measurement and identify actionable intervention points.
Sources
- The Blind Spot of Agent Safety: How Benign User Instructions Expose Critical Vulnerabilities in Computer-Use Agents
- Know Thy Judge: On the Robustness Meta-Evaluation of LLM Safety Judges
- BraveGuard: From Open-World Threats to Safer Computer-Use Agents
- Safety Testing LLM Agents at Scale: From Risk Discovery to Evidence-Grounded Verification
- How to Correctly Report LLM-as-a-Judge Evaluations
- Probing and Steering Evaluation Awareness of Language Models
- Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking Attacks
- Toward an Evaluation Science for Generative AI Systems
- FinVault: Benchmarking Financial Agent Safety in Execution-Grounded Environments
- CLEAR: Error Analysis via LLM-as-a-Judge Made Easy
- LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments
- Stop Comparing LLM Agents Without Disclosing the Harness
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection