DualView: Preventing Indirect Prompt Injection in Personal AI Agents
summary
The gist
This paper presents DualView, a defense mechanism designed to protect personal AI agents from indirect prompt injection (IPI) attacks.
In short
This episode discusses the paper 'DualView,' which addresses indirect prompt injection in personal AI agents. The researchers propose a system using two views: AgentView, where the agent sees symbols instead of potentially malicious text, and HumanView, which remains normal for users. This method prevents attacks while maintaining high agent utility.
Key concepts
- Indirect Prompt Injection
- This occurs when an AI agent reads malicious instructions hidden in external data, such as a website or email. These instructions can hijack the agent, causing it to perform unauthorized actions like stealing files or running commands without the user's knowledge.
- AgentView and HumanView
- DualView provides two ways to see a computer. In AgentView, the agent sees symbols like 's1' instead of dangerous text, preventing it from reading malicious instructions. HumanView keeps files looking normal for users. The system uses Git to manage different versions and keep both views synchronized.
- Stored Injection
- This occurs when an agent saves a malicious command into a file and reads it back later. DualView prevents this by using symbols that survive the trip through the file system, ensuring that even if a bad command is written to a file, it remains unreadable as an instruction.
Terminology used across episodes
This episode discusses
- DualView: Preventing Indirect Prompt Injection in Personal AI Agents · Paper Radio
- Prompt Flow Integrity to Prevent Privilege Escalation in LLM Agents
- Defeating Prompt Injections by Design
- Securing AI Agents with Information-Flow Control
- Benchmarking and Defending Against Indirect Prompt Injection Attacks on Large Language Models
- LlamaFirewall: An open source guardrail system for building secure AI agents
- DataSentinel: A Game-Theoretic Detection of Prompt Injection Attacks
- Defending Against Indirect Prompt Injection Attacks With Spotlighting
- MELON: Provable Defense Against Indirect Prompt Injection Attacks in AI Agents
- DRIFT: Dynamic Rule-Based Defense with Injection Isolation for Securing LLM Agents
- ShieldAgent: Shielding Agents via Verifiable Safety Policy Reasoning
- GuardAgent: Safeguard LLM Agents by a Guard Agent via Knowledge-Enabled Reasoning
- AgentSpec: Customizable Runtime Enforcement for Safe and Reliable LLM Agents · Paper Radio
- RTBAS: Defending LLM Agents Against Prompt Injection and Privacy Leakage
- Permissive Information-Flow Analysis for Large Language Models
- Prompt Infection: LLM-to-LLM Prompt Injection within Multi-Agent Systems
- System-Level Defense against Indirect Prompt Injection Attacks: An Information Flow Control Perspective
- Ignore Previous Prompt: Attack Techniques For Language Models
- Jailbroken: How Does LLM Safety Training Fail?
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
The paper
DualView: Preventing Indirect Prompt Injection in Personal AI Agents · Read on arXiv
Seoul National University
Personal AI agents that run on the user's local machine automate daily tasks including web search, email, and file management. Their access to computer resources, including the network, file system, and shell, exposes them to indirect prompt injection (IPI) attacks. Prior Dual LLM defenses block IPI by replacing untrusted data with symbols that the agent can reference but not read. However, they track untrusted data only inside the agent's context, so when the agent saves and later rereads untrusted data, that data, possibly an attacker's prompt, can return as trusted data rather than as a symbol, which we call stored IPI. Operating on the user's real environment is what makes agents like OpenClaw practical, and is exactly why a defense that ignores it is incomplete. Preserving symbols in such an environment is hard, because humans and programs need original data. We present DualView, which extends untrusted data tracking from the agent's context to the user's environment, including the file system, shell, network, and other agents, by giving each channel two views. In AgentView, the agent sees untrusted data as symbols even after writing it out and reading it back, blocking stored IPI, while HumanView preserves original data for humans and tools. DualView routes each tool call to the right view and synchronizes data across the two views. DualView deploys as an OpenClaw plugin using only tool hooks, without changing the agent's tool-call logic or tool implementations. DualView deterministically prevents instructions in untrusted data from directly steering the agent's tool calls; this guarantee does not depend on recognizing the evaluated attack templates. In our evaluation on an IPI benchmark and PinchBench, DualView blocked every tested IPI attack, including stored IPI. On PinchBench, its utility drop was within 1.8 to 6.4 points.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "DualView: Preventing Indirect Prompt Injection in Personal AI Agents".
Jane: The paper was written by Juhee Kim, Woohyuk Choi, Taehyun Kang, Youngmin Kim and Byoungyoung Lee from Seoul National University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: We are starting our show today with a look at 'DualView: Preventing Indirect Prompt Injection in Personal AI Agents'.
Jane: The researchers behind this work come from Seoul National University.
Tom: They are tackling the scary problem of how an assistant can be hijacked by things it reads on the internet.
Jane: Imagine a website having hidden instructions that tell your assistant to steal your files.
Meng: That sounds like a nightmare for anyone letting an agent run commands in their terminal.
Lu: It is even worse when you realize the attacker could hide those instructions in an email or a simple webpage!
Jane: You're right, Lu, because the agent can't easily tell the difference between your data and an attacker's command.
Tom: This specific type of attack is what they call indirect prompt injection.
Meng: If I had an agent managing my files, I would be terrified to let it browse the web without something like this.
Lu: We could end up with agents that act like silent spies inside our own computers!
Jane: That's a very vivid way to put it, Lu.
Tom: It really makes you wonder how we can ever trust these agents with our most sensitive tasks.
Lalam: Building this kind of security is the only way we can eventually let AI handle our personal lives safely.
Jane: I agree, Lalam, because without safety, there is no real utility for the user.
Tom: But how do they actually separate these two worlds without making a mess of your files?
Summary: Tom: We've established the danger of hijacked agents, so let's look at how DualView actually works.
Jane: The core idea is that the system provides two different ways to see your computer.
Tom: They call them AgentView and HumanView.
Jane: In AgentView, the agent only sees symbols like "s1" instead of the actual dangerous text.
Meng: So if a webpage says "delete everything," the agent just sees a weird placeholder?
Tom: Exactly, so the model never actually reads the malicious instruction.
Jane: Meanwhile, HumanView keeps everything looking normal for you and your other programs to read.
Meng: How do they keep those two versions from getting completely out of sync?
Jane: They use Git to manage different versions of your files, almost like a developer does.
Tom: It's a clever way to make sure when you edit a file, the agent sees that change too.
Meng: Using Git worktrees seems like a very practical engineering choice for this.
Lu: It's like giving the AI specialized glasses so it can see the structure of your data without being blinded by any traps!
Jane: I love that analogy, Lu!
Tom: It also solves the problem of "stored injection" where an agent saves a bad command and reads it back later.
Meng: So if it writes a symbol to a file, that symbol stays there when it reads it again?
Tom: Precisely, because the tracking survives the whole trip through your file system.
Lu: This means an agent can finally have a long-term memory that isn't poisoned by enemies!
Jane: That would be such a huge leap forward for making these tools reliable.
Lalam: A secure memory is what will allow an AI to become a true, dependable partner in our lives.
Tom: So it's secure, but does it actually perform well in the real world?
Improvements: Tom: We've seen the architecture, so now let's talk about the actual performance and trade-offs.
Jane: The researchers tested this against various attacks and found a zero percent success rate.
Tom: That is a massive win, especially since they hit both immediate and stored injections.
Jane: They even compared it to sandboxing, which actually makes the agent much less useful.
Meng: I noticed in the paper that there is a significant token overhead for some models.
Tom: You're right, Meng, they saw costs as high as ninety-three percent on Claude Sonnet.
Jane: That's a lot of extra processing power to pay for!
Meng: If I need absolute security for my business, I might actually be willing to pay that extra cost.
Lu: The way they route tools between the two views is just incredibly smart!
Jane: It really is about finding that balance between being safe and being helpful.
Tom: They even showed that agent utility stays very high, only losing a few percentage points in task success.
Meng: That's much better than the alternatives where the agent just stops working entirely.
Lu: It's like teaching an agent to walk a tightrope instead of just putting it in a cage!
Jane: That is such a great way to describe it, Lu.
Tom: They also added a data trust policy so you can choose which sources are safe.
Lalam: Giving humans that kind of granular control is how we build lasting trust in technology.
Meng: It makes the whole system feel like a professional tool rather than just a toy.
Jane: Let's bring this all together and look at what this means for the future.
Conclusion: Tom: We are wrapping up our discussion on 'DualView: Preventing Indirect Prompt Injection in Personal AI Agents'.
Jane: This paper really shows that security and utility don't have to be enemies.
Tom: By using these dual views, they've found a way for agents to live in our real world without being controlled by it.
Lu: I can see a future where every digital interaction is protected by this kind of invisible layer!
Meng: From an engineering standpoint, seeing that zero percent success rate makes this look like a real standard.
Lalam: This advances our culture by ensuring we don't delegate our safety along with our tasks.
Jane: It really changes the conversation about how much we can rely on these assistants.
Tom: If we can secure the environment, the possibilities for what these agents can do are endless.
Lu: We could have agents that manage our entire digital lives with total peace of mind!
Meng: I'll be watching closely to see how this gets implemented in real-world runtimes.
Lalam: It's about moving from a state of constant suspicion to one of true collaboration.
Jane: That is such a powerful thought to end on, Lalam.
Tom: Thank you all for joining us for this deep dive into the research!
Jane: We'll be back soon with another fascinating paper, but for now, goodbye everyone!
More episodes
- 2610.10857-Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization