Cochise: A Reference Harness for Autonomous Penetration Testing
summary
The gist
The gist The Cochise prototype is a minimal reference agent and harness designed to provide an execution interface, model abstraction, state handling, and unified trajectory format so that
In short
Cochise is a minimal reference agent and harness designed for autonomous penetration testing experiments. It provides a structured interface for researchers to compare different AI models and agent architectures using a unified trajectory format. The system connects an LLM to a testbed, managing long-term planning and low-level execution while logging all interactions for detailed analysis.
Key concepts
- Minimal Reference Agent/Harness
- Cochise acts as a basic framework or harness that provides the necessary structure—like execution interfaces, state handling, and unified formats—so researchers can easily plug in different AI models and testbed setups. It is not meant to be a final product but a standardized starting point for building new prototypes.
- Planner Component
- This is the strategic core of Cochise responsible for maintaining the long-term state of the penetration testing task. It uses a structured textual representation, often called a Pentest-Task-Tree (PTT), to keep track of goals and progress over time, guiding what actions should be taken next.
- Executor Component
- This component translates the high-level directives from the planner into concrete commands. It operates as a ReAct agent that connects to a virtual execution host via SSH to run actual commands within the target network, handling the immediate operational steps.
- Logging and Analysis Pipeline
- Cochise records every interaction—from LLM prompts and raw completions to executed bash strings—in a standardized JSON log file. This detailed logging allows researchers to analyze performance metrics like token cost, duration, and success rates for various agents and models.
Terminology used across episodes
This episode discusses
- Cochise: A Reference Harness for Autonomous Penetration Testing · Paper Radio
- What Makes a Good LLM Agent for Real-world Penetration Testing?
- CAI: An Open, Bug Bounty-Ready Cybersecurity AI
- Incalmo: An Autonomous LLM-assisted System for Red Teaming Multi-Host Networks
- Better Harnesses, Smaller Models: Building 90% Cheaper Agents via Automated Harness Adaptation
- SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
- ReAct: Synergizing Reasoning and Acting in Language Models
The paper
Cochise: A Reference Harness for Autonomous Penetration Testing · Read on arXiv
TU Wien
Recent work on LLM-driven autonomous penetration testing reports promising results, but existing systems often bundle architectural, prompting, and tool-integration choices together. This makes it difficult to determine what is gained over a simple agent and harness. We present Cochise, a 630 LOC Python reference implementation for autonomous penetration-testing experiments. Cochise connects to a Linux execution host over SSH and supports attacking controlled target environments reachable from that jump host. The prototype implements a Planner--Executor architecture in which long-term state is maintained by the planner, while a ReAct-style executor issues commands over SSH and self-corrects based on command outputs. The scenario prompt can be adapted to different target environments. We evaluate the harness against a live third-party testbed, Game of Active Directory (GOAD). Cochise is intended not as a state-of-the-art penetration-testing agent, but as a reusable experimental infrastructure for comparing models, agent architectures, and penetration-testing traces. Alongside the prototype, we release replay and analysis tools: (i) cochise-replay for offline visualization of captured runs, (ii) cochise-analyze-logs and cochise-analyze-graphs for cost, token, duration, and compromise analysis, and (iii) a corpus of JSON trajectory logs from GOAD runs, so that researchers can study agent behavior without provisioning the 48--64 GB RAM / 190 GB storage testbed themselves. Tool demo video available at https://youtu.be/2mQimB1ufyI.
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: Today's paper: "Cochise: A Reference Harness for Autonomous Penetration Testing".
Elias: The gist The Cochise prototype is a minimal reference agent and harness designed to provide an execution interface, model abstraction, state handling,
Nadia: First, who's behind it and why it matters.
Title and authors: Nadia: So we're talking about "Cochise: A Reference Harness for Autonomous Penetration Testing". It’s a reference implementation, a six hundred thirty lines of Python code that lets you test agents against real target environments like Game of Active Directory <ref:2605.11671#pg1>.
Elias: That's the key part—it’s designed to be this reusable infrastructure. They aren't claiming it’s the best agent ever; they are positioning it as a way for researchers to compare different models and architectures fairly under one common protocol.
Priya: I see why that matters, because if you can swap out the LLM or change the planning logic easily without rewriting everything else, you can really isolate what part of the agent is doing what.
Nadia: Right. They connect this harness to a Linux execution host using SSH and let it run commands inside a controlled test environment reachable from that jump host. It’s about making the connection between the LLM brain and the actual attack surface very structured.
Elias: And they introduce this Planner-Executor architecture, where the planner keeps track of long-term state in a structured text representation, while the executor translates those high-level goals into concrete commands over SSH.
Priya: That decomposition is interesting because it addresses that multi-step nature of attacks—the agent isn't just guessing one thing; it’s building a path and then executing the steps along that path.
The paper's summary: Nadia: The paper summarizes Cochise as a way to handle the complexity of autonomous penetration testing by separating the strategic planning from the low-level execution. The planner handles the long-term state, and it feeds directives down to an executor that uses a ReAct style agent to issue commands over SSH.
Elias: They emphasize that this design tackles several hard problems inherent in security testing environments, like partial observability, where you don't see everything at once, and side-effecting actions where an exploit might crash the target system or trigger intrusion detection systems.
Priya: What I find important is how they handle those side effects; they force the agent to deal with them by needing autonomous error recovery and combining findings into successful attack paths over multiple steps. It’s not just about finding one vulnerability; it’s about chaining them together.
Nadia: Right. And that leads into their logging system, where every interaction—planner to executor, LLM calls, target network actions—is recorded in a JSON log file for reproducibility and analysis later on.
Elias: They detail how they log LLM invocations with specific event keys that name the architectural action, along with the exact prompt and completion received. And commands executed get their own identifiers showing the raw bash string run, along with the output streams.
Priya: So it’s not just a run log; it’s a rich dataset that lets you analyze things like cost and token usage for each step of the process, which is crucial for understanding how efficient those agents are.
The paper's improvements: Nadia: The authors point out several specific improvements they made to this baseline harness. They focused on making the executor state management very clean; they ensure a new executor instance is created for each task and then discarded once it’s finished.
Elias: That ephemeral state management is key because it bounds the context window size and the cost for any single task, which otherwise would grow indefinitely if you kept one giant context running for hours of operation.
Priya: That makes sense; if you don't bound the executor's memory, transient failures from one step could contaminate the reasoning for a completely different later step in the attack path. It keeps things focused and manageable.
Nadia: They also highlight that this structure forces the planner to become the integration point for cross-task knowledge, meaning it has to manage what happened in task one before it can properly plan task two.
Elias: That’s a structural change that makes the planner's job much harder and more meaningful than just letting the LLM handle everything sequentially without a high-level state manager guiding the flow.
Priya: And they mention using the Reflexion pattern within both components to actively detect and then repair invalid command invocations, which is another layer of self-correction built into the harness structure itself.
Conclusion: Nadia: So, to wrap up on "Cochise: A Reference Harness for Autonomous Penetration Testing", they’ve given us a minimal but functional system that structures agent behavior by separating planning and execution. It’s a six hundred thirty-line reference implementation connecting an LLM to a testbed like GOAD <ref:2605.11671#pg1>.
Elias: They stress that the real value isn't just the code itself, but how it forces you to think about penetration testing as a software engineering problem, handling all those stateful, multi-step issues with explicit planning and controlled execution.
Priya: The implication for me is that this tool gives researchers a solid foundation. It’s not trying to be the final agent; it’s providing the infrastructure so other people can build on top of it to compare different models and architectural variants systematically.
Nadia: Exactly. And they provide tools like cochise-replay and analysis scripts so you can actually look at the trajectories, analyze the cost, and see how it performed against a live testbed. It’s an experimental infrastructure for comparing things.
Elias: So, ultimately, this paper is about providing a standard way to run autonomous testing experiments safely and repeatably so we can properly evaluate different approaches to agent design.
Priya: It sets up the necessary framework so that the next step isn't just building another agent from scratch, but building an agent using this harness as its foundation.
Nadia: That’s it for Cochise. We’ll leave you with this reference infrastructure for autonomous penetration testing, built on a planner-executor design.
More episodes
- 2610.10644-SoK: Failure Modes in Common Criteria Product Evaluation - A Taxonomy and Design-for-Evaluability Guidance
- 2610.10617-MRCert: Towards Post-deployment Patch Robustness Certification for Adversarially Patched Samples via Type-specific Masking
- 2610.10620-When AI Finds Hidden Messages, Does It Report?
- 2610.10625-Safe at One Loop, Risky at Another: Aligning Safety Across Recurrent Depths in Looped Language Models
- 2610.10992-The Hint Weight of ML-DSA Signatures Is Key-Dependent: An Empirical Study across the Three FIPS 204 Parameter Sets
- 2610.10659-Applying Security by Design at the Point of Execution: How Governed Security Requirements Affect the Security of AI-Generated Code
- 2610.10735-DITTO: A Context-aware Pickle-based Pre-Trained Model Scanner for Effective Security Audits
- 2610.10742-BRANCH: Bypassing Multi-Scanner AI Guardrails
- 2610.10752-Detection-Guided Adaptive Purification with Diffusion Models for Robust Audio Deepfake Detection
- 2610.10766-CPU-Auth: Device Fingerprinting for Authentication via DVFS Side-Channel