DualView: Preventing Indirect Prompt Injection in Personal AI Agents

arXiv:2607.03821 · cs.CR, cs.AI · Submitted 2026-07-04 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "DualView: Preventing Indirect Prompt Injection in Personal AI Agents".

Jane: The paper was written by Juhee Kim, Woohyuk Choi, Taehyun Kang, Youngmin Kim and Byoungyoung Lee from Seoul National University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: We are starting our show today with a look at 'DualView: Preventing Indirect Prompt Injection in Personal AI Agents'.

Jane: The researchers behind this work come from Seoul National University.

Tom: They are tackling the scary problem of how an assistant can be hijacked by things it reads on the internet.

Jane: Imagine a website having hidden instructions that tell your assistant to steal your files.

Meng: That sounds like a nightmare for anyone letting an agent run commands in their terminal.

Lu: It is even worse when you realize the attacker could hide those instructions in an email or a simple webpage!

Jane: You're right, Lu, because the agent can't easily tell the difference between your data and an attacker's command.

Tom: This specific type of attack is what they call indirect prompt injection.

Meng: If I had an agent managing my files, I would be terrified to let it browse the web without something like this.

Lu: We could end up with agents that act like silent spies inside our own computers!

Jane: That's a very vivid way to put it, Lu.

Tom: It really makes you wonder how we can ever trust these agents with our most sensitive tasks.

Lalam: Building this kind of security is the only way we can eventually let AI handle our personal lives safely.

Jane: I agree, Lalam, because without safety, there is no real utility for the user.

Tom: But how do they actually separate these two worlds without making a mess of your files?

Summary: Tom: We've established the danger of hijacked agents, so let's look at how DualView actually works.

Jane: The core idea is that the system provides two different ways to see your computer.

Tom: They call them AgentView and HumanView.

Jane: In AgentView, the agent only sees symbols like "s1" instead of the actual dangerous text.

Meng: So if a webpage says "delete everything," the agent just sees a weird placeholder?

Tom: Exactly, so the model never actually reads the malicious instruction.

Jane: Meanwhile, HumanView keeps everything looking normal for you and your other programs to read.

Meng: How do they keep those two versions from getting completely out of sync?

Jane: They use Git to manage different versions of your files, almost like a developer does.

Tom: It's a clever way to make sure when you edit a file, the agent sees that change too.

Meng: Using Git worktrees seems like a very practical engineering choice for this.

Lu: It's like giving the AI specialized glasses so it can see the structure of your data without being blinded by any traps!

Jane: I love that analogy, Lu!

Tom: It also solves the problem of "stored injection" where an agent saves a bad command and reads it back later.

Meng: So if it writes a symbol to a file, that symbol stays there when it reads it again?

Tom: Precisely, because the tracking survives the whole trip through your file system.

Lu: This means an agent can finally have a long-term memory that isn't poisoned by enemies!

Jane: That would be such a huge leap forward for making these tools reliable.

Lalam: A secure memory is what will allow an AI to become a true, dependable partner in our lives.

Tom: So it's secure, but does it actually perform well in the real world?

Improvements: Tom: We've seen the architecture, so now let's talk about the actual performance and trade-offs.

Jane: The researchers tested this against various attacks and found a zero percent success rate.

Tom: That is a massive win, especially since they hit both immediate and stored injections.

Jane: They even compared it to sandboxing, which actually makes the agent much less useful.

Meng: I noticed in the paper that there is a significant token overhead for some models.

Tom: You're right, Meng, they saw costs as high as ninety-three percent on Claude Sonnet.

Jane: That's a lot of extra processing power to pay for!

Meng: If I need absolute security for my business, I might actually be willing to pay that extra cost.

Lu: The way they route tools between the two views is just incredibly smart!

Jane: It really is about finding that balance between being safe and being helpful.

Tom: They even showed that agent utility stays very high, only losing a few percentage points in task success.

Meng: That's much better than the alternatives where the agent just stops working entirely.

Lu: It's like teaching an agent to walk a tightrope instead of just putting it in a cage!

Jane: That is such a great way to describe it, Lu.

Tom: They also added a data trust policy so you can choose which sources are safe.

Lalam: Giving humans that kind of granular control is how we build lasting trust in technology.

Meng: It makes the whole system feel like a professional tool rather than just a toy.

Jane: Let's bring this all together and look at what this means for the future.

Conclusion: Tom: We are wrapping up our discussion on 'DualView: Preventing Indirect Prompt Injection in Personal AI Agents'.

Jane: This paper really shows that security and utility don't have to be enemies.

Tom: By using these dual views, they've found a way for agents to live in our real world without being controlled by it.

Lu: I can see a future where every digital interaction is protected by this kind of invisible layer!

Meng: From an engineering standpoint, seeing that zero percent success rate makes this look like a real standard.

Lalam: This advances our culture by ensuring we don't delegate our safety along with our tasks.

Jane: It really changes the conversation about how much we can rely on these assistants.

Tom: If we can secure the environment, the possibilities for what these agents can do are endless.

Lu: We could have agents that manage our entire digital lives with total peace of mind!

Meng: I'll be watching closely to see how this gets implemented in real-world runtimes.

Lalam: It's about moving from a state of constant suspicion to one of true collaboration.

Jane: That is such a powerful thought to end on, Lalam.

Tom: Thank you all for joining us for this deep dive into the research!

Jane: We'll be back soon with another fascinating paper, but for now, goodbye everyone!

Seoul National University

cs.CR, cs.AI

Submitted: 2026-07-04

Updated: 2026-09-14

Code: https://github.com/compsec-snu/dualview

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 86/100

The gist: This paper presents DualView, a defense mechanism designed to protect personal AI agents from indirect prompt injection (IPI) attacks.

Key concepts

Indirect Prompt Injection
This occurs when an AI agent reads malicious instructions hidden in external data, such as a website or email. These instructions can hijack the agent, causing it to perform unauthorized actions like stealing files or running commands without the user's knowledge.
AgentView and HumanView
DualView provides two ways to see a computer. In AgentView, the agent sees symbols like 's1' instead of dangerous text, preventing it from reading malicious instructions. HumanView keeps files looking normal for users. The system uses Git to manage different versions and keep both views synchronized.
Stored Injection
This occurs when an agent saves a malicious command into a file and reads it back later. DualView prevents this by using symbols that survive the trip through the file system, ensuring that even if a bad command is written to a file, it remains unreadable as an instruction.

Terminology

Summary

This paper presents DualView, a defense mechanism designed to protect personal AI agents from indirect prompt injection (IPI) attacks. As agents gain the ability to automate tasks using local files, shell commands, and network requests, they become susceptible not only to immediate IPI but also to stored IPI, where malicious instructions are saved in the user's environment and later re-read as trusted data. DualView addresses this by extending untrusted data tracking from the agent’s context into the user’s computer environment.

The challenge of stored IPI

Existing defenses, such as the Dual LLM pattern, attempt to block IPI by replacing untrusted data with opaque symbols. However, these defenses track untrusted data only inside the agent's context. When an agent saves and later rereads untrusted data from its environment, that data can return as trusted data rather than as a symbol, allowing an attacker to regain control. This creates a fundamental conflict between security, agent utility (the ability to complete tasks), and human utility (the ability for humans to read and use the agent's output). A defense must either keep data symbolized—which fills the environment with unreadable symbols—or resolve symbols into original data, which remains vulnerable to stored IPI.

How DualView works

DualView resolves this conflict by providing two views of the user's environment to satisfy all three requirements simultaneously:

  • AgentView: In this view, untrusted data appears to the agent only as symbols, even after it has been written to a file and read back. This ensures that tracking survives the write-then-read path and blocks stored IPI. AgentView utilizes an Agent File System and a local-only AgentShell.

  • HumanView: This view is designed for humans, non-agent programs, and remote network endpoints. In HumanView, users see original data, not symbols, ensuring the environment remains usable.

To maintain consistency, DualView employs tool view routing and synchronization. It routes tool calls to the appropriate view—for example, running network requests in HumanView because remote endpoints need original data—and synchronizes changes so that human edits are visible to the agent and agent writes are visible to humans in real time.

Policy-driven protection

The system's security and usability are governed by two distinct policies:

  1. Data Trust Policy: This policy classifies returned data using tool schemas and origin rules. It determines whether data should be symbolized or kept as original data based on whether it derives from remote content or a trusted source.

  2. Data Usage Policy: This acts as a best-effort layer that monitors how untrusted data is utilized. It specifically targets the risk of an agent desymbolizing data into HumanShell; if a symbol would be used as a command or code, DualView withholds the execution and asks the user for approval.

Evaluation results

The researchers evaluated DualView using an IPI benchmark and PinchBench, testing it against Claude Haiku 4.5 and Claude Sonnet 4.6 models. The results demonstrated that:

  • DualView blocked every IPI attack, including stored IPI, achieving a 0% attack success rate.

  • Agent utility remained close to the unprotected baseline, with success rates staying within a small margin of the unprotected OpenClaw agent.

  • Human utility was preserved, as the system leaves human-facing files, messages, and tool outputs free of symbols.

Improvements for AI systems

1. Improvement: Environment-Level Symbolic Tracking via Dual-View Architecture

  • Capability: The AI system can maintain a strict distinction between AgentView (where untrusted data is replaced by opaque symbols like s1) and HumanView (where original data is preserved). This allows the agent to write malicious instructions into a local file and later read that same file back without the instructions being interpreted as commands, effectively neutralizing Stored Indirect Prompt Injection.

2. Improvement: Dual-Worktree Filesystem Synchronization using Git

  • Capability: The system can manage two simultaneous versions of the user's filesystem. The agent operates on a symbolic Git worktree (AgentView) where untrusted file content is represented by symbols, while the human user interacts with a real worktree (HumanView). This ensures the agent remains secure during long-term memory retrieval while ensuring humans and non-agent programs see legible, original text.

3. Improvement: Schema-Based Data Trust Policy

  • Capability: The system can perform granular, field-level symbolization of tool outputs. For example, when an agent fetches a web page or receives a webhook, the system can automatically treat the status code and URL as trusted (original text) while strictly symbolizing the body content or message text. This prevents the agent from being steered by untrusted content while still allowing it to act on reliable metadata.

4. Improvement: Adaptive Tool Routing (AgentShell vs. HumanShell)

  • Capability: The system can dynamically route tool calls based on their requirements. It can use a restricted AgentShell (with network access disabled and symbolic filesystem access) for safe tasks like grep or ls, and switch to a HumanShell (with full network/original data access) only when the agent must perform external actions like curl. This minimizes the attack surface by limiting when untrusted data is ever desymbolized.

5. Improvement: Symbolic Execution Guardrails and Command Rewriting

  • Capability: The system can prevent arbitrary code execution by intercepting commands in the HumanShell that attempt to use symbols as command arguments or interpreters (e.g., python-c " symbol"). Additionally, it can implement command rewriting (e.g., transforming python script.py into an inline command) to detect and block attempts to execute untrusted code hidden within a file, requiring explicit human approval before execution.

6. Improvement: Provenance-Aware Symbol Table

  • Capability: The system can maintain a SQLite-based symbol table that maps every symbol to its original data and its specific provenance (e.g., which tool or network endpoint produced it). This allows the agent to implement advanced security policies, such as blocking any command that attempts to use a symbol originating from a specifically untrusted domain or unverified webhook port.

Abstract

Personal AI agents that run on the user's local machine automate daily tasks including web search, email, and file management. Their access to computer resources, including the network, file system, and shell, exposes them to indirect prompt injection (IPI) attacks. Prior Dual LLM defenses block IPI by replacing untrusted data with symbols that the agent can reference but not read. However, they track untrusted data only inside the agent's context, so when the agent saves and later rereads untrusted data, that data, possibly an attacker's prompt, can return as trusted data rather than as a symbol, which we call stored IPI. Operating on the user's real environment is what makes agents like OpenClaw practical, and is exactly why a defense that ignores it is incomplete. Preserving symbols in such an environment is hard, because humans and programs need original data. We present DualView, which extends untrusted data tracking from the agent's context to the user's environment, including the file system, shell, network, and other agents, by giving each channel two views. In AgentView, the agent sees untrusted data as symbols even after writing it out and reading it back, blocking stored IPI, while HumanView preserves original data for humans and tools. DualView routes each tool call to the right view and synchronizes data across the two views. DualView deploys as an OpenClaw plugin using only tool hooks, without changing the agent's tool-call logic or tool implementations. DualView deterministically prevents instructions in untrusted data from directly steering the agent's tool calls; this guarantee does not depend on recognizing the evaluated attack templates. In our evaluation on an IPI benchmark and PinchBench, DualView blocked every tested IPI attack, including stored IPI. On PinchBench, its utility drop was within 1.8 to 6.4 points.

Sources

Related papers