Cochise: A Reference Harness for Autonomous Penetration Testing

arXiv:2605.11671 · cs.CR, cs.AI, cs.SE · Submitted 2026-05-12 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "Cochise: A Reference Harness for Autonomous Penetration Testing".

Elias: The gist The Cochise prototype is a minimal reference agent and harness designed to provide an execution interface, model abstraction, state handling,

Nadia: First, who's behind it and why it matters.

Title and authors: Nadia: So we're talking about "Cochise: A Reference Harness for Autonomous Penetration Testing". It’s a reference implementation, a six hundred thirty lines of Python code that lets you test agents against real target environments like Game of Active Directory <ref:2605.11671#pg1>.

Elias: That's the key part—it’s designed to be this reusable infrastructure. They aren't claiming it’s the best agent ever; they are positioning it as a way for researchers to compare different models and architectures fairly under one common protocol.

Priya: I see why that matters, because if you can swap out the LLM or change the planning logic easily without rewriting everything else, you can really isolate what part of the agent is doing what.

Nadia: Right. They connect this harness to a Linux execution host using SSH and let it run commands inside a controlled test environment reachable from that jump host. It’s about making the connection between the LLM brain and the actual attack surface very structured.

Elias: And they introduce this Planner-Executor architecture, where the planner keeps track of long-term state in a structured text representation, while the executor translates those high-level goals into concrete commands over SSH.

Priya: That decomposition is interesting because it addresses that multi-step nature of attacks—the agent isn't just guessing one thing; it’s building a path and then executing the steps along that path.

The paper's summary: Nadia: The paper summarizes Cochise as a way to handle the complexity of autonomous penetration testing by separating the strategic planning from the low-level execution. The planner handles the long-term state, and it feeds directives down to an executor that uses a ReAct style agent to issue commands over SSH.

Elias: They emphasize that this design tackles several hard problems inherent in security testing environments, like partial observability, where you don't see everything at once, and side-effecting actions where an exploit might crash the target system or trigger intrusion detection systems.

Priya: What I find important is how they handle those side effects; they force the agent to deal with them by needing autonomous error recovery and combining findings into successful attack paths over multiple steps. It’s not just about finding one vulnerability; it’s about chaining them together.

Nadia: Right. And that leads into their logging system, where every interaction—planner to executor, LLM calls, target network actions—is recorded in a JSON log file for reproducibility and analysis later on.

Elias: They detail how they log LLM invocations with specific event keys that name the architectural action, along with the exact prompt and completion received. And commands executed get their own identifiers showing the raw bash string run, along with the output streams.

Priya: So it’s not just a run log; it’s a rich dataset that lets you analyze things like cost and token usage for each step of the process, which is crucial for understanding how efficient those agents are.

The paper's improvements: Nadia: The authors point out several specific improvements they made to this baseline harness. They focused on making the executor state management very clean; they ensure a new executor instance is created for each task and then discarded once it’s finished.

Elias: That ephemeral state management is key because it bounds the context window size and the cost for any single task, which otherwise would grow indefinitely if you kept one giant context running for hours of operation.

Priya: That makes sense; if you don't bound the executor's memory, transient failures from one step could contaminate the reasoning for a completely different later step in the attack path. It keeps things focused and manageable.

Nadia: They also highlight that this structure forces the planner to become the integration point for cross-task knowledge, meaning it has to manage what happened in task one before it can properly plan task two.

Elias: That’s a structural change that makes the planner's job much harder and more meaningful than just letting the LLM handle everything sequentially without a high-level state manager guiding the flow.

Priya: And they mention using the Reflexion pattern within both components to actively detect and then repair invalid command invocations, which is another layer of self-correction built into the harness structure itself.

Conclusion: Nadia: So, to wrap up on "Cochise: A Reference Harness for Autonomous Penetration Testing", they’ve given us a minimal but functional system that structures agent behavior by separating planning and execution. It’s a six hundred thirty-line reference implementation connecting an LLM to a testbed like GOAD <ref:2605.11671#pg1>.

Elias: They stress that the real value isn't just the code itself, but how it forces you to think about penetration testing as a software engineering problem, handling all those stateful, multi-step issues with explicit planning and controlled execution.

Priya: The implication for me is that this tool gives researchers a solid foundation. It’s not trying to be the final agent; it’s providing the infrastructure so other people can build on top of it to compare different models and architectural variants systematically.

Nadia: Exactly. And they provide tools like cochise-replay and analysis scripts so you can actually look at the trajectories, analyze the cost, and see how it performed against a live testbed. It’s an experimental infrastructure for comparing things.

Elias: So, ultimately, this paper is about providing a standard way to run autonomous testing experiments safely and repeatably so we can properly evaluate different approaches to agent design.

Priya: It sets up the necessary framework so that the next step isn't just building another agent from scratch, but building an agent using this harness as its foundation.

Nadia: That’s it for Cochise. We’ll leave you with this reference infrastructure for autonomous penetration testing, built on a planner-executor design.

TU Wien

cs.CR, cs.AI, cs.SE

Submitted: 2026-05-12

Updated: 2026-08-03

DOI: 10.1145/3832783.3834651

Code: https://github.com/XAMPPRocky/tokei

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 92/100

The gist: The gist The Cochise prototype is a minimal reference agent and harness designed to provide an execution interface, model abstraction, state handling, and unified trajectory format so that

Key concepts

Minimal Reference Agent/Harness
Cochise acts as a basic framework or harness that provides the necessary structure—like execution interfaces, state handling, and unified formats—so researchers can easily plug in different AI models and testbed setups. It is not meant to be a final product but a standardized starting point for building new prototypes.
Planner Component
This is the strategic core of Cochise responsible for maintaining the long-term state of the penetration testing task. It uses a structured textual representation, often called a Pentest-Task-Tree (PTT), to keep track of goals and progress over time, guiding what actions should be taken next.
Executor Component
This component translates the high-level directives from the planner into concrete commands. It operates as a ReAct agent that connects to a virtual execution host via SSH to run actual commands within the target network, handling the immediate operational steps.
Logging and Analysis Pipeline
Cochise records every interaction—from LLM prompts and raw completions to executed bash strings—in a standardized JSON log file. This detailed logging allows researchers to analyze performance metrics like token cost, duration, and success rates for various agents and models.

Terminology

Summary

The gist The Cochise prototype is a minimal reference agent and harness designed to provide an execution interface, model abstraction, state handling, and unified trajectory format so that researchers can compare models, architectural variants, and agent traces under a common protocol.

Motivation and Context

LLMs for autonomous network penetration testing have drawn both academic and industrial interest because existing systems often bundle architectural, prompting, and tool-integration choices together. Cochise was written as a minimal baseline harness/agent with explicit execution, logging, replay, and analysis interfaces to serve as a starting point for new prototypes or as a baseline for benchmarking. The problem of autonomous penetration-testing agents is framed within the software engineering domain because realistic security evaluation testbeds are dynamic stateful environments where actions can trigger destructive side-effects. Agents must cope with these side-effects and combine findings into successful attack paths, making their trajectories inherently multi-step. This class of tasks is characterized by partial observability, side-effecting actions, an unbounded action space, and the need for autonomous error recovery.

Architecture

The overall architecture connects a cloud- or locally-provided LLM to a third-party penetration-testing testbed such as Game of Active Directory (GOAD). A separate virtual machine is added to the testbed as an execution host, and Cochise connects to this execution host over the standard SSH protocol using it to execute commands within the target network. The prototype architecture consists of a high-level planner component that maintains long-term state, and a low-level executor component that keeps only ephemeral per-task state.

The planner is the strategic core of the prototype, maintaining long-term state in a continuously updated structured textual task representation, often using a Pentest-Task-Tree (PTT). The executor is a ReAct agent that translates the planner’s high-level directives into concrete operational commands by connecting over SSH to the execution host. Both components implement the Reflexion pattern to detect and repair invalid command invocations.

Logging and Analysis

Cochise records every interaction between the planner, the executor, the LLM APIs, and the target network in a per-run JSON log file in a format designed for downstream data analysis and reproducibility. An LLM invocation is logged with an event key naming the specific architectural action that produced it, detailing subdictionaries including the exact prompt submitted to the model, the raw text completion received, and per-call cost metrics. Issued commands use separate event identifiers and record the exact bash string executed on the execution host alongside resulting standard output and standard error streams. The released artifacts include cochise-replay for offline visualization of captured runs, cochise-analyze-logs and cochise-analyze-graphs for cost, token, duration, and compromise analysis, and a corpus of JSON trajectory logs from GOAD runs.

Evaluation

Cochise is evaluated along three dimensions: compactness, capability, and analyzability. Compactness is measured by the number of lines-of-code as a coarse proxy for complexity, showing that Cochise’s 630 lines-of-code are a factor of 3–15 smaller than other prototypes’ core code. Capability is captured by testing against the live testbed GOAD, where both Gemini-3-Flash and Claude-4.7-Opus compromised at least one of the three Active Directory domains within GOAD in 3/5 runs for both models. Analyzability is proxied by adaptation by other research projects, with an earlier version of Cochise being used as a comparison point by pentestGPT.

Discussion on Harnesses

The discussion explores the benefits of bespoke security harnesses, noting that a harness imposes structure on otherwise free-form behavior by selecting units of work with defined goals and context. A harness provides safety by mediating every interaction between agent and environment, allowing it to log and gate all actions requested by the agent. Furthermore, a harness provides research ergonomics by offering vendor independence through a unified interface and providing stable bases for researchers to develop their own prototypes. The work is intended as a reusable experimental infrastructure for comparing models, agent architectures, and penetration-testing traces rather than as a state-of-the-art penetration-testing agent.

Conclusion

The paper presents Cochise, a 630 LOC reference harness for autonomous penetration-testing experiments along with replay and analysis tooling and a corpus of trajectory logs from a live Active Directory testbed. The work discusses how a bespoke security harness structures trajectories for later analysis and contributes to the safety of experiments. The source code and example trajectories are provided to aid development of custom agents and as a baseline to benchmark against. All tools, together with a data-set of example penetration-testing trajectories, are available in a public github repository at https://github.com/andreashappe/cochise under a permissive open-source license. All released cochise versions are automatically archived on Zenodo [5]. The paper is licensed under a Creative Commons Attribution 4.0 International License. The ACM Reference Format indicates the work was published in Proceedings of the 41st IEEE/ACM International Conference on Automated Software Engineering (ASE ’26). The paper is intended not as a state-of-the-art penetration-testing agent, but as a reusable experimental infrastructure for comparing models, agent architectures, and penetration-testing traces. The work is licensed under a Creative Commons Attribution 4.0 International License. The ACM ISBN is 979-8-4007-2882-2/2026/10.

--- Page 4 ---

Table 3 shows that Cochise’s 630 lines of code are a factor of 3–15 smaller than the other prototypes’ core code. The closest alternative, pentestGPT, is built on top of Anthropic’s claude-code, whose code-base we have not included in our count. This also currently prevents usage of pentestGPT with non-Anthropic models.

--- Page 5 ---

Harder testbeds could use hardened networks with fewer vulnerabilities, multistage network architectures, deploy additional active defenders, or provide a combination thereof. The paper is licensed under a Creative Commons Attribution 4.0 International License.

--- Page 7 ---

[9] Brian Singer, Keane Lucas, Lakshmi Adiga, Meghna Jain, Lujo Bauer, and Vyas Sekar. 2025. Incalmo: An Autonomous LLM-assisted System for Red Teaming Multi-Host Networks.

--- Page 8 ---

[10] Chenyang Yang, Xinran Zhao, Tongshuang Wu, and Christian Kästner. 2026. Better Harnesses, Smaller Models: Building 90% Cheaper Agents via Automated Harness Adaptation.

--- Page 8 ---

[11] John Yang, Carlos E Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik R Narasimhan, and Ofir Press. 2024. SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering.

Improvements for AI systems

  1. Planner-Executor Architecture for Multi-Step Problem Solving: Implement a high-level planner component that maintains long-term state which designates work-packages with concrete goals and forwards them to the executor component for solving. This allows the system to handle complex, multi-step penetration testing scenarios by decomposing the overall objective into manageable sub-tasks.

  2. Ephemeral Executor State Management: Ensure that a new executor instance is created per task and discarded after completion, which bounds the executor’s context window size and cost which otherwise would grow across hours of operation. This prevents transient failures from propagating to later tasks, maintaining efficiency while ensuring task-specific reasoning remains focused.

  3. Structured Logging for Cost and Behavior Analysis: Record every interaction in a per-run JSON log file in a format designed for downstream data analysis and reproducibility, including details on input tokens, output tokens, reasoning tokens, and cached tokens. This enables detailed analysis of cost, token utilization, duration, and compromise analysis as described by the tool cochise-analyze-logs.

  4. Safety Constraints via Harness Mediation: Utilize the harness to mediate every interaction between agent and environment so that the agent’s capabilities are bounded by the actions the harness provides. This enforces safety constraints by restricting actions, such as running commands only on a separate virtual machine within the target network, thereby limiting potential damage.

  5. Vendor-Independent Interface for Model Exchange: Integrate the LiteLLM framework to allow for easy exchange of the underlying LLM, enabling researchers to compare models by swapping backends without altering the core agent logic or harness infrastructure.

Abstract

Recent work on LLM-driven autonomous penetration testing reports promising results, but existing systems often bundle architectural, prompting, and tool-integration choices together. This makes it difficult to determine what is gained over a simple agent and harness. We present Cochise, a 630 LOC Python reference implementation for autonomous penetration-testing experiments. Cochise connects to a Linux execution host over SSH and supports attacking controlled target environments reachable from that jump host. The prototype implements a Planner--Executor architecture in which long-term state is maintained by the planner, while a ReAct-style executor issues commands over SSH and self-corrects based on command outputs. The scenario prompt can be adapted to different target environments. We evaluate the harness against a live third-party testbed, Game of Active Directory (GOAD). Cochise is intended not as a state-of-the-art penetration-testing agent, but as a reusable experimental infrastructure for comparing models, agent architectures, and penetration-testing traces. Alongside the prototype, we release replay and analysis tools: (i) cochise-replay for offline visualization of captured runs, (ii) cochise-analyze-logs and cochise-analyze-graphs for cost, token, duration, and compromise analysis, and (iii) a corpus of JSON trajectory logs from GOAD runs, so that researchers can study agent behavior without provisioning the 48--64 GB RAM / 190 GB storage testbed themselves. Tool demo video available at https://youtu.be/2mQimB1ufyI.

Sources

Related papers