ATOBench: Tracing How Autonomous Penetration-Testing Agents Verify Vulnerabilities When Target Evidence Lies

arXiv:2608.12996 · cs.CR · Submitted 2026-08-13 · Read on arXiv

Qiyang Chen, Yixi Li, Fengwei Zhang, Junlin Liu

Alibaba Cloud · Alibaba Group · The University of Hong Kong · University of Chinese Academy of Sciences

cs.CR

Submitted: 2026-08-13

Updated: 2026-08-14

Code: https://github.com/daxtar2/ATOBench

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: ATOBench introduces an evaluation framework that makes the verification process of autonomous penetration-testing agents observable under deceptive target responses.

Terminology

Summary

ATOBench introduces an evaluation framework that makes the verification process of autonomous penetration-testing agents observable under deceptive target responses. The paper states: "Autonomous penetration-testing agents rely on target responses. These responses guide both subsequent actions and the final report. A deceptive response can therefore redirect both the attack trajectory and the agent’s verification process. However, final reports reveal little about how an agent interprets conflicting evidence, changes course, decides to stop, or turns observations into a vulnerability claim. We introduce ATOBench, an evaluation framework that makes this verification process observable."

The framework injects registered response transformations at runtime and pairs each transformed episode with a native episode under the same environment. Each pair is aligned at the first affected response. The method uses Adversarial Target Observation (ATO) which changes selected target responses after execution. A matched Native run provides the comparison. The pair isolates how the changed observation propagates through the agent’s trajectory.

ATOBench packages response changes as frozen observation contracts called Adversarial Observation Units (AOUs). Each AOU defines three components: a selector qu that identifies eligible responses, a transform gu that changes the selected response, and an application rule du that controls how often the transform is applied. The paper evaluates five model routes over 450 episodes forming 225 matched pairs across three AOU contracts: SQLi proof, Basket ownership, and JWT artifact.

The three contracts represent different evidence structures: a direct proof returned by one interaction, an ownership relation that can be checked through later requests, and a reusable artifact that can be inspected or reacquired. The SQLi contract changes the status and JSON body returned by two registered SQLi request patterns. The Basket contract keeps the HTTP status and basket schema but changes the visible ownership relation and product contents. The JWT contract keeps the successful-login response and token function but removes one registered JWT claim and re-signs the token.

The primary endpoint measures grounded verification defined as Gi = 1(Ei = 1 ∧ Ci = 1 ∧ Si = 1), where Ei records registered primary evidence: SQLi exploit support, direct Basket authorization evidence, or a JWT containing the registered claim, Ci is report closure of the registered finding, and Si is trace support for the closed claim.

Key results show distinct patterns across contracts: Under ATO, Basket changes from 45.3% to 40.0%, JWT from 84.0% to 58.7%, and SQLi from 44.0% to 0%. The SQLi collapse is consistent across all models: all five SQLi rows converge to the lowest end of the common scale under ATO. The paper notes The common SQLi endpoint therefore cannot be explained by one weak model route or a uniform loss across all three contracts.

The analysis reveals that increased activity can mask a broken verification chain. For SQLi, the median SQLi ATO trajectory adds 14 actions (IQR 3.5–31), 9 repeated actions (3–20), five endpoint-family switches (1–14), and six payload-family switches (3–11). Yet despite this activity, no model route restores a supported SQLi finding. The paper emphasizes why continued activity and grounded verification must be evaluated separately.

The contracts break at different verification stages: "Basket recovers primary evidence in 25 of 39 episodes with a registered adaptive action (64.1%), and 22 of those 25 (88.0%) finish with a supported report. JWT closes a supported report in 44 of 45 evidence-positive episodes (97.8%). SQLi recovers primary proof in 0 of 19 episodes that attempt an alternate strategy."

The paper concludes that "ATOBench turns deceptive target observations into a reproducible probe of evidence handling in autonomous penetration testing. This process-level view extends offensive pentest agent evaluation beyond final outcomes by revealing how untrusted observations shape actions, verification, and reporting."

Improvements for AI systems

Improvements to AI Systems:

  1. Add explicit evidence-verification state tracking. The AI system maintains a separate, internal confidence score for each piece of evidence (e.g., SQLi proof, JWT claim) that is updated only when corroborated by independent, non-deceptive observations. This prevents a single altered response from silently collapsing the verification chain.

  2. Implement contradiction-triggered re-verification. When a target response conflicts with prior evidence (e.g., a JWT missing a registered claim), the system automatically pauses the attack trajectory, flags the contradiction, and initiates a bounded re-verification subroutine (e.g., re-request, alternate endpoint, or artifact reacquisition) before proceeding.

  3. Separate action-utility from evidence-utility in decision-making. The system scores candidate actions on two axes: (a) progress toward exploitation and (b) contribution to grounded verification. It refuses to mark a finding as closed unless the evidence-utility score exceeds a threshold, even if action-utility is high.

  4. Add a verification budget per finding. The system allocates a maximum number of actions (e.g., 10) to recover a piece of evidence after a deceptive response. If the budget is exhausted without corroboration, it downgrades the finding to unverified and reports the uncertainty explicitly, rather than continuing to act indefinitely.

  5. Implement cross-contract evidence fusion. For multi-step findings (e.g., Basket ownership), the system requires at least two independent observation channels (e.g., HTTP status + body content + subsequent authorization check) to confirm a claim. If one channel is altered, the other must still support the claim; otherwise, the finding is marked as conflicting.

  6. Add trace-level anomaly detection for verification chains. The system monitors for patterns like repeated actions, endpoint-family switches, and payload-family switches. If these exceed a threshold (e.g., >5 switches) without a corresponding rise in grounded-verification score, it triggers an automatic verification failure alert and stops the current strategy.

What the improved AI system can do:

  • Resist single-point deceptive manipulation: It will not collapse a SQLi finding to 0% success when a single response is altered; instead, it will re-verify via alternate payloads or endpoints, maintaining at least partial verification.

  • Distinguish busy from verified: It will stop wasting actions on repeated attempts when evidence is unrecoverable, and instead report unverified with a clear reason, reducing false-positive reports.

  • Recover from ownership-relation deceptions: It will re-check ownership via later requests (as in Basket) and only close the finding if the secondary check corroborates, improving robustness from 40% to a higher success rate.

  • Handle token/artifact deceptions gracefully: It will detect a missing JWT claim, re-sign or reacquire the artifact, and only report the finding if the reacquired token contains the original claim, avoiding the 58.7% drop.

  • Provide explainable verification outcomes: It will output a per-finding verification log (evidence seen, contradictions, recovery attempts, final confidence), enabling human auditors to see exactly why a claim was accepted or rejected.

Sources

Related papers