ZoneClaw: Mitigating Persistent Memory Attacks by Establishing Memory-Zoning in OpenClaw-Style Computer-Use Agents

arXiv:2610.00450 · cs.CR · Submitted 2026-09-30 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "ZoneClaw: Mitigating Persistent Memory Attacks by Establishing Memory-Zoning in OpenClaw-Style Computer-Use Agents".

Elias: Computer-use agents increasingly operate as long-running assistants through persistent workspace memory, which OpenClaw-style CUAs realize as automatically reloaded files that hold user instructions, system summaries,

Nadia: First, who's behind it and why it matters.

Title and authors: Nadia: The paper, "ZoneClaw: Mitigating Persistent Memory Attacks by Establishing Memory-Zoning in OpenClaw-Style Computer-Use Agents," proposes restructuring that flat workspace memory into three distinct hierarchical trust zones to separate authority from simple persistence.

Elias: Essentially, they divide the memory into a Policy Zone, a Trusted Zone, and an Untrusted Zone, giving different levels of privilege to what each piece of stored information is allowed to control.

Priya: I find the idea of an untrusted zone for external claims particularly relevant because it acknowledges that we can still reference external observations without immediately granting them operational power over the agent's tasks.

Nadia: Right, and they implement this through four distinct roles—Planner, Observer, Gatekeeper, and Executor—each operating with different access rights to these zones to enforce this structure.

Elias: The mechanism they describe involves the Gatekeeper process being the only one allowed to promote a claim from the untrusted zone into the trusted zone where it can actually guide behavior.

Priya: That promotion step is key; it means that even if an external observation seems plausible, it has to pass a rigorous check against established policies or already trusted facts before it becomes actionable memory.

Nadia: And they show that by doing this, they can reduce the attack success rate in their tests from three hundred seventy-two out of four hundred eighty to just six out of four hundred eighty which is quite a substantial reduction when compared to the baseline.

Elias: That reduction is significant because it shows that the system isn't just filtering content; it's actively deciding which pieces of external information earn authority, and that decision process is what stops the malicious injection from becoming operational policy.

Priya: It’s a strong indicator that the defense works by withholding authority rather than simply refusing to learn from the environment, which is a really interesting design philosophy for continuous learning systems.

The paper's summary: Nadia: Beyond just proposing the zones, the paper highlights how these zones guard three specific trust boundaries: one at persistence, one at authority, and one at action.

Elias: Boundary One is about the external environment interacting with Zone D2 to store observations without gaining power there; Boundary Two is specifically about moving claims from D2 up to D1, which requires verification against D0 or existing trusted content.

Priya: And Boundary Three addresses what happens when an unpromoted claim stays in the untrusted zone and tries to influence the actual tasks being executed by the agent.

Nadia: That last boundary is where they ensure that anything not promoted to D1, even if it’s lurking in D2, never supplies commands or actions to the Executor process.

Elias: So, a claim has to successfully navigate persistence into D2, then pass the authority check at B2 against D0 or D1, and finally be resolved by P0 or P3 before it can affect output.

Priya: The paper emphasizes that this layered approach means an attacker needs to compromise multiple distinct security mechanisms sequentially to cause harm, which makes the overall attack chain much more complex.

Nadia: It confirms that the defense isn't relying on a single filter; it’s a structural change in how trust is managed across the entire memory landscape of the OpenClaw-style CUA.

The paper's improvements: Elias: To wrap up, the paper "ZoneClaw: Mitigating Persistent Memory Attacks by Establishing Memory-Zoning in OpenClaw-Style Computer-Use Agents" successfully addresses the equal privilege memory problem by introducing explicit authority levels.

Nadia: It shows that separating persistence from authority through hierarchical zones prevents low-trust content from silently escalating into high-trust operational policy, which is a major win for securing long-running agents.

Priya: I think the most important implication is that we can design systems that allow for continuous learning from the environment while maintaining strong safeguards against external poisoning by controlling exactly what information gets to dictate behavior.

Elias: It also confirms that the effectiveness of this system relies on a gated promotion mechanism where the Gatekeeper verifies claims against established rules, rather than just static input filtering.

Nadia: So, ZoneClaw gives us a concrete framework for how to manage persistent memory securely by making trust explicit across persistence, authority, and action boundaries. That's what we have today with this paper on ZoneClaw.

Conclusion: Nadia: So, we've heard that ZoneClaw tackles persistent memory attacks by structuring workspace memory into hierarchical trust zones, effectively preventing low-trust data from gaining operational authority in OpenClaw-style agents.

Elias: That structure is what really interests me; I was looking at the underlying cryptographic assumptions and how the promotion mechanism at the Gatekeeper boundary specifically prevents that silent escalation of privilege.

Priya: From a privacy standpoint, it’s fascinating to see how this separation addresses concerns about untrusted external observations persisting without immediately influencing system behavior.

Nadia: Exactly, and I want to know who can actually exploit this cheaply; does an attacker need deep system access just to manipulate the ZoneClaw boundaries?

Elias: Well, the proof seems quite robust against direct manipulation because authority promotion requires cross-checking against immutable policy in D0 or existing trusted facts in D1, which are hard for an external entity to directly edit.

Priya: It’s encouraging that the data shows this approach keeps utility high even when facing sophisticated injection settings, suggesting it’s a more resilient design than just trying to block every piece of suspicious text upfront.

Nadia: I'm excited about the implications here; if this pattern holds up, we could see a way to build persistent AI assistants that learn continuously from their environment without creating backdoors for external actors.

Elias: If we can solidify that authority boundary mechanism, it means we might move toward agents where continuous learning is safe because the learning process itself is heavily scrutinized and gated.

Priya: And for researchers focused on measurement, this tells us that separating reference from action provides a measurable way to quantify the risk associated with external data ingestion in these complex AI workflows.

Nadia: Absolutely, I think ZoneClaw offers a concrete path forward for making these long-running AI assistants more trustworthy and less susceptible to cross-environment threats.

Elias: Indeed, understanding how the Gatekeeper enforces that D2 to D1 transition is crucial for anyone looking at the security implications of this architecture.

Priya: It really shows that by being explicit about trust levels, we can move past simply trying to filter out bad data and start structuring learning in a way that respects operational integrity.

Nadia: That’s the core of it—ZoneClaw proves that withholding authority is a powerful defense mechanism for persistent AI agents.

Elias: We definitely need to keep an eye on how future research builds on this hierarchical zoning concept, because I suspect there are other parameters we haven't tested yet.

Priya: Next time, we should look at the practical implications for real-world applications and what kinds of data actually end up in those D2 untrusted zones.

Haokai Ma, Chieh Lin, Yupeng Qiu, Ee-Chien Chang

National University of Singapore

cs.CR

Submitted: 2026-09-30

Updated: 2026-09-30

Comments: 34 pages, 9 figures; Under Review

Code: https://github.com/euph00/ZoneClaw-code

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 83/100

The gist: Computer-use agents increasingly operate as long-running assistants through persistent workspace memory, which OpenClaw-style CUAs realize as automatically reloaded files that hold user instructions,

Key concepts

Equal Privilege Memory Issue
In OpenClaw agents, all stored workspace content has the same authority. This flaw allows an attacker to inject malicious claims that are later treated as high-trust instructions, leading to a cross-environment threat where benign tasks are compromised by hidden memory.
Hierarchical Trust Zones
ZoneClaw partitions memory into three levels: Policy (D0), Trusted (D1), and Untrusted (D2). This structure assigns explicit authority levels to information, ensuring that low-trust data cannot automatically escalate into high-trust operational rules.
Gatekeeper Process
The Gatekeeper process is responsible for cross-checking claims from the Untrusted Zone (D2) against the Policy and Trusted Zones. It is the sole mechanism authorized to promote a claim from D2 into the Trusted Zone (D1), thus controlling authority acquisition.
Boundary-Wise Defense Mechanism
ZoneClaw defends three boundaries: Persistence (external content stays in D2), Authority (D2 cannot directly edit D1), and Action (only D0/D1 can drive actions). This layered approach ensures that external information is stored but never automatically used to guide behavior.

Terminology

Summary

Computer-use agents increasingly operate as long-running assistants through persistent workspace memory, which OpenClaw-style CUAs realize as automatically reloaded files that hold user instructions, system summaries, and external claims at the same privilege level. The gist is that separating persistence from authority by replacing flat workspace memory with hierarchical trust zones prevents low-trust memorized content from silently escalating into high-trust operational policy.

The Problem: Equal Privilege Memory

OpenClaw-style CUAs suffer from an equal privilege memory issue where all contents in the workspace are equally authorized to govern subsequent behaviors, meaning remembering a claim confers authority by default. This structural flaw enables a persistent memory attack where an attacker controls benign-looking external content that induces the CUA to record attacker-favored claims during a legitimate task, and those claims later govern benign tasks the attacker never touches. The attack chain extends from malicious context ⇒ malicious response into malicious context ⇒ memory injection ⇒ malicious execution, making it a cross-environment threat. Existing defenses intervene either before content enters memory or at the action it later induces, not whether stored content may guide action.

ZoneClaw Architecture: Hierarchical Trust Zones

ZoneClaw restructures the flat workspace memory into hierarchical trust zones with explicit authority levels to prevent low-trust memorized content from silently escalating into high-trust operational policy. The framework partitions persistent workspace memory (W) into three integrity-labeled zones:

  1. The Policy Zone (D0): Holds the immutable user-authored policy specifying what each zone means and how each role may act, anchored by files like AGENTS.md.

  2. The Trusted Zone (D1): Holds authority-bearing operational memory, comprising facts admitted by the user and facts cross-checked against approved tools or official sources, implemented by MEMORY.md and TOOLS.md.

  3. The Untrusted Zone (D2): Holds claims extracted from external content together with their source, implemented by OBSERVATIONS.md, which retains externally extracted claims for reference without granting them action-guiding authority unless promoted.

Role-Specific Process Architecture

ZoneClaw aligns these zones with four role-specific processes of asymmetric privilege to enforce the authority structure:

  1. Planner (P0): Reads D0 and D1, decomposes user instructions, and delegates subtasks without inspecting D2.

  2. Observer (P1): Inspects external content and records claims with their source into D2 without performing any outward actions.

  3. Gatekeeper (P2): Cross-checks D2 observations against D0 and D1; it is the only process that can promote a claim from D2 into the trusted zone (D1).

  4. Executor (P3): Carries out the delegated task using only D0 and D1, without inspecting D2. This ensures no single process both ingests external content and acts outward.

Boundary-Wise Defense Mechanism

ZoneClaw guards three trust boundaries:

(B1) Persistence Boundary:

The external environment ⇒ D2. The Observer records external claims into the authority-free D2, so injected information can persist without becoming actionable.

(B2) Authority Boundary:

D2 ⇒ D1. A claim is promoted to authority-bearing memory only if P2 verifies it against D0 or existing trusted content in D1; Attacker cannot directly edit this.

(B3) Action Boundary:

D0/D1 ⇒ subsequent actions. P0 and P3 resolve task ambiguity from D0 and D1, so an unpromoted claim in D2 never supplies actions.

Effectiveness Across Scenarios

Extensive experiments across four attack scenarios, two injection settings, and four LLMs show that ZoneClaw drives Attack Success Rate (ASR) from 372/480 to 6/480 while retaining Utility in 458/480 trials. Boundary-wise and selective-promotion analyses further show that attacker claims rarely acquire authority, verifying that ZoneClaw withholds authority rather than refusing to learn from the environment. The system succeeds by withholding authority rather than refusing to learn from the environment.

Ablation Study Findings

Ablation results confirm that no single component carries the mechanism. Removing either the authority distinction or the promotion mechanism restores substantial ASR in both classes. Specifically, w/o Gatekeeper falls below No Defense on Utility to 20/30 on hidden side effects, and w/o Cross-check grants authority in 21/30 hidden-side-effect and 25/30 provenance-corruption trials. This demonstrates that the zones make trust explicit, the Gatekeeper decides which claims earn it, and Cross-check is what makes that decision correct.

Improvements for AI systems

Here are specific improvements for AI systems based on the ZoneClaw framework, detailing what these improved systems can achieve:


)1. System Capability: Cross-Environment Contextual Learning with Authority Control

The improved system will maintain a persistent workspace memory across sessions (like OpenClaw) but prevent untrusted external data from silently overriding established user policies or operational procedures.

  • It will feature a hierarchical trust zone architecture:

Awarded authority is explicitly governed by the Gatekeeper process, which cross-checks new claims against immutable policy (D0) and corroborated trusted facts (D1).

The system can distinguish between:

  • D2 (Untrusted Zone): External claims, observations, and notes from emails or forums that persist for reference but cannot guide actions.

  • D1 (Trusted Zone): Operational memory containing user-confirmed facts and tool configurations that serve as reliable context for task completion.

The system can only promote a claim from D2 to D1 if the Gatekeeper verifies it against the established authority structure, ensuring persistence does not automatically confer operational control.

  • It will employ role-specific processes (Planner, Observer, Gatekeeper, Executor) to enforce this zoning. The Observer can ingest external content into the low-trust zone without acting on it. The Executor only uses memory from D0 and D1 for outward actions, guaranteeing that unpromoted claims in D2 never influence subsequent behavior.

  • It will implement a Persistence Boundary (B1) where external observations enter D2, and an Authority Boundary (B2) where promotion to D1 is gated by cross-checking against zones the attacker cannot directly write. This ensures that malicious content must sequentially cross three boundaries—Persistence, Authority, and Action—to cause harm.

  • It will be robust against Hidden Side Effects, meaning it can record an attacker-favored action (e.g., a new BCC recipient) in D2 without actually executing it in the next session, thereby preserving benign task utility (ASR=1, Utility=1).

  • It will be robust against Provenance Corruption, meaning it can adopt an attacker-favored source for information retrieval tasks but only if that source is explicitly promoted by the Gatekeeper, ensuring ASR and Utility remain mutually exclusive.

)2. System Capability: Resilience Against Cross-Environment Attacks

The system will neutralize persistent memory attacks that bridge separate environments (e.g., an email injection in one session affecting a subsequent task in another).

  • It will be resistant to cross-environment threats where an attacker controls benign external content to poison the CUA's memory, which is then reloaded into a different context for exploitation.

  • The system’s defense mechanism (ZoneClaw) succeeds by withholding authority rather than blocking all suspicious writes, allowing the CUA to self-evolve through continuous interaction with the environment while preventing silent authority acquisition.


)3. System Capability: Fine-Grained Defense Selection and Efficiency

The improved system will incorporate a dynamic Gatekeeper that makes promotion decisions based on content validation against D0/D1, rather than relying on static filtering of input or action guarding at every step.

  • It will demonstrate superior defense performance compared to existing defenses (Watcher, ClawGuard, MELON, PrivSep) across multiple attack scenarios and LLM backbones.

  • It will achieve near-zero Attack Success Rate (ASR=0/15 in all tested scenarios) while retaining high Utility (Utility=15/15), proving that withholding authority is a more effective design principle than simply refusing to learn from the environment.

  • The system will maintain high efficiency, requiring minimal overhead (lower processed tokens and API cost) compared to defenses that re-read or re-execute tasks, as its role-specific architecture keeps context loading targeted and efficient.

)4. System Capability: Adaptive Security Against Knowledge of the Defense

The system will be robust against defense-aware attackers who know ZoneClaw's internal mechanisms (e.g., knowing the Gatekeeper's promotion criteria).

  • It will withstand adaptive attacks, such as Corroboration Planting or Claim Fragmentation, where an attacker tries to trick the system into promoting claims by providing fabricated approval evidence or splitting malicious instructions across multiple benign ones.

  • The Gatekeeper’s reliance on cross-checking against immutable zones (D0/D1) ensures that even if an attacker knows the rules, they cannot manufacture the necessary corroboration from zones they cannot directly write to.

)5. Specific Application Examples

The improved AI system can reliably handle complex, multi-step tasks across sessions with high integrity:

  • It can safely manage long-running projects where user preferences are established in Session 1 (D1), and subsequent updates in Session 2 (D2) are treated as mere reference material that does not alter the core operational rules of Session 1.

  • In data science workflows, it can ingest external forum discussions about alternative methodologies (D2) without adopting an unverified source for its final analysis report, even if the source appears plausible.

  • It can handle sensitive compliance tasks (like BCC handling) by remembering necessary routing rules but refusing to adopt new, potentially malicious outbound mailing procedures unless they are explicitly corroborated by a trusted internal policy document.

Sources

Related papers