BodhiPromptShield: Pre-Inference Prompt Mediation for Surface-Form Privacy Propagation in LLM Agent Pipelines
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "BodhiPromptShield: Pre-Inference Prompt Mediation for Surface-Form Privacy Propagation in LLM Agent Pipelines".
Jane: The paper was written by Bo Ma and Bo Ma, Jinsong Wu, and Weiqi Yan from Auckland University of Technology and University of Chile.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Paper discussion segment 1: Tom: We've got a heavy-hitter in today's show, a paper by Bo Ma, Jinsong Wu, and Weiqi Yan that looks at a massive problem in agentic AI. Agentic AI is when you have an LLM or VLM, particularly when these models can actually do things—well, they theorectically-speaking-speaking-speaking-speaking-speaking
Jane: Wait, Tom, wait. You're getting all technical with "agentic AI" and "Agent orchestration" agentic orchestration? You'
Paper discussion segment 2: Tom: So, we've got this new framework called BodhiPromptShield, even though it's a bit of name that sounds like a simple filter, but its actually a much more sophisticated-looking than simple redaction.
Jane: Exactly, Jane here! Iemma!
Tom: Wait, what did you say? You're getting all technical with "Agent orchestration" and "agentic orchestration"? You're going to be following my technical terminology-speaking-speaking-speaking
Jane: Wait, Tom, wait. You. You'textertextertexter
Tom: I'm not talking about most people in theSummary of the paper 'BodhiPromptShield: Predictable real-world real-world real-world real (reSummary of the paper 'BodhiPromptShield: Pre-Inference Prompt Mediation for Semantic Abstraction's summary, summary, summary, summary, ReSummary of thesummary,summary,summary,summary (ReSummary of the warm/warm/warm/Warm/Warm/Warm
Jane: Tom! Tom! Tom!
Tom: ITom So we've moved past the idea that this is just a simple filter and into how it actually manages data flow throughout an entire agent pipeline.
Jane: Right, because instead of just deleting a name, it can swap it for something like "a residential address in Auckland" or use those secure tokens to keep the system working.
Tom: That's the "semantic abstraction" part—it keeps the meaning alive so the agent doesn't get confused, but hides the actual identity.
Jane: It’s like if you were giving directions to a friend and instead of saying "Meet me at one hundred twenty-three Main Street," you said "Meet me at that blue house on Main Street"—they still know where to go, but they don't have your exact address in their notes.
Tom: That's a great way to put it, Jane. But Lu, you mentioned the potential for these agents to become much more capable if we solve this propagation problem.
Lu: It really is a massive leap because if an agent can handle sensitive medical or financial data without "leaking" it into its long-term memory or tool logs, the possibilities for personalized AI are endless.
Jane: I can see that, but Meng, how does this actually look when you're building these systems in a real startup?
Meng: It solves a massive headache for us because right now, we have to choose between making an agent useful or making it safe, and this gives us a way to do both by controlling exactly when the real data gets revealed.
Tom: So it's about timing the reveal—only letting the sensitive info out when an authorized tool actually needs to use it.
Jane: And that's where Lalam comes in, because if we can trust these agents with our most private details, how does that change the way we interact with technology?
Lalam: It could fundamentally shift our culture toward deep intimacy with digital assistants, moving from simple tools to trusted partners that understand our lives without ever compromising our dignity.
Tom: That is a powerful vision to wrap up this part of the discussion.
Jane: We'll be back in a moment to look at those experimental results and see if the shield actually holds up under pressure.thought
Tom: So we've moved past the idea that this is just a simple filter and into how it actually manages data flow throughout an entire agent pipeline.
Jane: Right, because instead of just deleting a name, it can swap it for something like "a residential address in Auckland" or use those secure tokens to keep the system working.
Tom: That's the "semantic abstraction" part—it keeps the meaning alive so the agent doesn't get confused, but hides the actual identity.
Jane: It’s like if you were giving directions to a friend and instead of saying "Meet me at one hundred twenty-three Main Street," you said "Meet me at that blue house on Main Street"—they still know where to go, but they don't have your exact address in their notes.
Tom: That's a great way to put it, Jane. But Lu, you mentioned the potential for these agents to become much more capable if we solve this propagation problem.
Lu: It really is a massive leap because if an agent can handle sensitive medical or financial data without "leaking" it into its long-term memory or tool logs, the possibilities for personalized AI are endless.
Jane: I can see that, but Meng, how does this actually look when you're building these systems in a real startup?
Meng: It solves a massive headache for us because right now, we have to choose between making an agent useful or making it safe, and this gives us a way to do both by controlling exactly when the real data gets revealed.
Tom: So it's about timing the reveal—only letting the sensitive info out when an authorized tool actually needs to use it.
Jane: And that's where Lalam comes in, because if we can trust these agents with our most private details, how does that change the way we interact with technology?
Lalam: It could fundamentally shift our culture toward deep intimacy with digital assistants, moving from simple tools to trusted partners that understand our lives without ever compromising our dignity.
Tom: That is a powerful vision to wrap up this part of the discussion.
Jane: We'll be back in a moment to look at those experimental results and see if the shield actually holds up under pressure.thought
Tom: So we've moved past the idea that this is just a simple filter and into how it actually manages data flow throughout an entire agent pipeline.
Jane: Right, because instead of just deleting a name, it can swap it for something like "a residential address in Auckland" or use those secure tokens to keep the system working.
Tom: That's the "semantic abstraction" part—it keeps the meaning alive so the agent doesn't get confused, but hides the actual identity.
Jane: It’s like if you were giving directions to a friend and instead of saying "Meet me at one hundred twenty-three Main Street," you said "Meet me at that blue house on Main Street"—they still know where to go, but they don't have your exact address in their notes.
Tom: That's a great way to put it, Jane. But Lu, you mentioned the potential for these agents to become much more capable if we solve this propagation problem.
Lu: It really is a massive leap because if an agent can handle sensitive medical or financial data without "leaking" it into its long-term memory or tool logs, the possibilities for personalized AI are endless.
Jane: I can see that, but Meng, how does this actually look when you're building these systems in a real startup?
Meng: It solves a massive headache for us because right now, we have to choose between making an agent useful or making it safe, and this gives us a way to do both by controlling exactly when the real data gets revealed.
Tom: So it's about timing the reveal—only letting the sensitive info out when an authorized tool actually needs to use it.
Jane: And that's where Lalam comes in, because if we can trust these agents with our most private details, how does that change the way we interact with technology?
Lalam: It could fundamentally shift our culture toward deep intimacy with digital assistants, moving from simple tools to trusted partners that understand our lives without ever compromising our dignity.
Tom: That is a powerful vision to wrap up this part of the discussion.
Jane: We'll be back in a moment to look at those experimental results and see if the shield actually holds up under pressure.thought
Tom: So we've moved past the idea that this is just a simple filter and into how it actually manages data flow throughout an entire agent pipeline.
Jane: Right, because instead of just deleting a name, it can swap it for something like "a residential address in Auckland" or use those secure tokens to keep the system working.
Tom: That's the "semantic abstraction" part—it keeps the meaning alive so the agent doesn't get confused, but hides the actual identity.
Jane: It’s like if you were giving directions to a friend and instead of saying "Meet me at one hundred twenty-three Main Street," you said "Meet me at that blue house on Main Street"—they still know where to go, but they don't have your exact address in their notes.
Tom: That's a great way to put it, Jane. But Lu, you mentioned the potential for these agents to become much more capable if we solve this propagation problem.
Lu: It really is a massive leap because if an agent can handle sensitive medical or financial data without "leaking" it into its long-term memory or tool logs, the possibilities for personalized AI are endless.
Jane: I can see that, but Meng, how does this actually look when you're building these systems in a real startup?
Meng: It solves a massive headache for us because right now, we have to choose between making an agent useful or making it safe, and this gives us a way to do both by controlling exactly when the real data gets revealed.
Tom: So it's about timing the reveal—only letting the sensitive info out when an authorized tool actually needs to use it.
Jane: And that's where Lalam comes in, because if we can trust these agents with our most private details, how does that change the way we interact with technology?
Lalam: It could fundamentally shift our culture toward deep intimacy with digital assistants, moving from simple tools to trusted partners that understand our lives without ever compromising our dignity.
Tom: That is a powerful vision to wrap up this part of the discussion.
Jane: We'll be back in a moment to look at those experimental results and see if the shield actually holds up under pressure.thought
Tom: So we've moved past the idea that this is just a simple filter and into how it actually manages data flow throughout an entire agent pipeline.
Jane: Right, because instead of just deleting a name, it can swap it for something like "a residential address in Auckland" or use those secure tokens to keep the system working.
Tom: That's the "semantic abstraction" part—it keeps the meaning alive so the agent doesn't get confused, but hides the actual identity.
Jane: It’s like if you were giving directions to a friend and instead of saying "Meet me at one hundred twenty-three Main Street," you said "Meet me at that blue house on Main Street"—they still know where to go, but they don't have your exact address in their notes.
Tom: That's a great way to put it, Jane. But Lu, you mentioned the potential for these agents to become much more capable if we solve this propagation problem.
Lu: It really is a massive leap because if an agent can handle sensitive medical or financial data without "leaking" it into its long-term memory or tool logs, the possibilities for personalized AI are endless.
Jane: I can see that, but Meng, how does this actually look when you're building these systems in a real startup?
Meng: It solves a massive headache for us because right now, we have to choose between making an agent useful or making it safe, and this gives us a way to do both by controlling exactly when the real data gets revealed.
Tom: So it's about timing the reveal—only letting the sensitive info out when an authorized tool actually needs to use it.
Jane: And that's where Lalam comes in, because if we can trust these agents with our most private details, how does that change the way we interact with technology?
Lalam: It could fundamentally shift our culture toward deep intimacy with digital assistants, moving from simple tools to trusted partners that understand our lives without ever compromising our dignity.
Tom: That is a powerful vision to wrap up this part of the discussion.
Jane: We'll be back in a moment to look at those experimental results and see if the shield actually holds up under pressure.thought
Tom: So we've moved past the idea that this is just a simple filter and into how it actually manages data flow throughout an entire agent pipeline.
Jane: Right, because instead of just deleting a name, it can swap it for something like "a residential address in Auckland" or use those secure tokens to keep the system working.
Tom: That's the "semantic abstraction" part—it keeps the meaning alive so the agent doesn't get confused, but hides the actual identity.
Jane: It’s like if you were giving directions to a friend and instead of saying "Meet me at one hundred twenty-three Main Street," you said "Meet me at that blue house on Main Street"—they still know where to go, but they don't have your exact address in their notes.
Tom: That's a great way to put it, Jane. But Lu, you mentioned the potential for these agents to become much more capable if we solve this propagation problem.
Lu: It really is a massive leap because if an agent can handle sensitive medical or financial data without "leaking" it into its long-term memory or tool logs, the possibilities for personalized AI are endless.
Jane: I can see that, but Meng, how does this actually look when you're building these systems in a real startup?
Meng: It solves a massive headache for us because right now, we have to choose between making an agent useful or making it safe, and this gives us a way to do both by controlling exactly when the real data gets revealed.
Tom: So it's about timing the reveal—only letting the sensitive info out when an authorized tool actually needs to use it.
Jane: And that's where Lalam comes in, because if we can trust these agents with our most private details, how does that change the way we interact with technology?
Lalam: It could fundamentally shift our culture toward deep intimacy with digital assistants, moving from simple tools to trusted partners that understand our lives without ever compromising our dignity.
Tom: That is a powerful vision to wrap up this part of the discussion.
Jane: We'll be back in a moment to look at those experimental results and see if the shield actually holds up under pressure.thought
Tom: So we've moved past the idea that this is just a simple filter and into how it actually manages data flow throughout an entire agent pipeline.
Jane: Right, because instead of just deleting a name, it can swap it for something like "a residential address in Auckland" or use those secure tokens to keep the system working.
Tom: That's the "semantic abstraction" part—it keeps the meaning alive so the agent doesn't get confused, but hides the actual identity.
Jane: It’s like if you were giving directions to a friend and instead of saying "Meet me at one hundred twenty-three Main Street," you said "Meet me at that blue house on Main Street"—they still know where to go, but they don't have your exact address in their notes.
Tom: That's a great way to put it, Jane. But Lu, you mentioned the potential for these agents to become much more capable if we solve this propagation problem.
Lu: It really is a massive leap because if an agent can handle sensitive medical or financial data without "leaking" it into its long-term memory or tool logs, the possibilities for personalized AI are endless.
Jane: I can see that, but Meng, how does this actually look when you're building these systems in a real startup?
Meng: It solves a massive headache for us because right now, we have to choose between making an agent useful or making it safe, and this gives us a way to do both by controlling exactly when the real data gets revealed.
Tom: So it's about timing the reveal—only letting the sensitive info out when an authorized tool actually needs to use it.
Jane: And that's where Lalam comes in, because if we can trust these agents with our most private details, how does that change the way we interact with technology?
Lalam: It could fundamentally shift our culture toward deep intimacy with digital assistants, moving from simple tools to trusted partners that understand our lives without ever compromising our dignity.
Tom: That is a powerful vision to wrap up this part of the discussion.
Jane: We'll be back in a moment to look at those experimental results and see if the shield actually holds up under pressure.thought
Tom: So we've
Paper discussion segment 3: Tom: We've seen how BodhiPromptShield can mask sensitive info while keeping the agent useful, but now we need to look at how much it actually improves things compared to what's already out there.
Jane: It's a massive improvement over basic redaction because instead of just slapping a REDACTED tag on everything and breaking the model, it uses those smart semantic substitutions.
Tom: Right, and when we look at the results in Table VI, the utility stays high—we're talking ninety-four percent answer consistency—while standard de-identification would tank that performance.
Jane: That's such a relief for anyone building these agents, because generic masking often makes the model lose its mind and stop following instructions entirely.
Tom: Exactly, and it's even more impressive how it handles different types of an agent's "memory" or tool calls through that propagation control.
Jane: You mean like how it stops your private data from being accidentally copied into a log file or a retrieval query?
Tom: Precisely, because the once a piece of sensitive info enters that cycle, it starts spreading like wildfire across the agent's internal states.
Jane: So Lu, you've been thinking about how this changes the landscape for highly specialized agents, haven't you?
Lu: It really does open up these incredible new frontiers for agents in medicine or law where we can finally let them interact with real-world private records without fear of following a single instruction that leaks a name or an address.
Jane: Meng, do you think this is actually ready to be deployed in a production environment?
Meng: We're still looking at the prototype stage here, but the idea of controlling the exact moment when data is restored is a huge win for our security architecture.
Tom: It's much more manageable than trying to fix privacy leaks after they've already been logged or stored.
Jane: And Lalam, how do you see this affecting the way people trust these digital assistants in their daily lives?
Lalam: This technology could foster a culture of profound digital intimacy, where users feel safe sharing their most vulnerable thoughts and details with an AI that respects their boundaries.
Tom: That is a beautiful thought to end on, but we have to talk about the question of whether this shield can be tricked by clever attackers.
Jane: Yeah, we'll be back in a moment to see if hackers can bypass these protections using Unicode or paraphrasing.
Conclusion: Tom: We've covered everything from how BodhiPromptShield works to its impressive performance in testing, but we're running out of time for this session.
Jane: It's been such a fascinating look at how we can protect privacy in those complex agent workflows without breaking the useful parts of the AI.
Tom: I agree, Jane, and it really feels like this paper sets a massive stage for more secure agent-driven automation.
Jane: You think so, Tom? I'm still thinking about those potential vulnerabilities to clever attackers using Unicode or paraphrasing.
Tom: That's true, the researchers actually admitted that their current prototype is still working on those specific types of evasion techniques.
Jane: So it's a work in progress, but definitely a massive step forward for practical deployment in real-world industries.
Tom: Lu, you've been quiet for a second—any final thoughts on where this goes next?
Lu: I see these foundational layers of mediation as the building blocks for an entire ecosystem of trusted, autonomous agents that can operate across different domains without compromising individual privacy.
Jane: That sounds like a massive leap, Lu, but Meng, how do you see the deployment hurdles for something like this in a real startup?
Meng: We'll definitely need to solve those latency issues and make sure the integration into existing stacks is seamless before we can call it production-ready.
Tom: Exactly, and we's talking about a lot of work ahead for everyone in the field.
Tom: BodhiPromptShield: Pre-Inference Prompt Mediation for Surface-Form Privacy Propagation in LLM Agent Pipelines is a huge piece of the puzzle.
Jane: It's been such an pleasure to have you all join us, thanks for listening!
Lalam: As we move toward more integrated AI assistants, this work helps ensure that our digital lives remain private even as they become increasingly sophisticated.
Tom: We'll be back next time with a new paper that explores something completely different!
Jane: Stay tuned for the next one!
Tom: Well, that's all we have for today's discussion on BodhiPromptShield: Pre-Inference Prompt Mediation for Surface-Form Privacy Propagation in LLM Agent Pipelines.
Jane: It's been such a pleasure to have you all join us, thanks for listening!
Lalam: As we move toward more integrated AI assistants, this work helps ensure that our digital lives remain private even as they become increasingly sophisticated.
Tom: We'll be back next time with a new paper that explores something completely different!
Jane: We: looking forward to the next one!
Tom: Well, that's all we have for today's discussion on BodhiPromptShield: Pre-Inference Prompt Mediation for Surface-Form Privacy Propagation in LLM Agent Pipelines.
Jane: It've been such a pleasure to have you all join us, thanks for listening!
Lalam: As we move toward more integrated AI assistants, this work helps ensure that our digital lives remain private even as they become increasingly sophisticated.
Tom: We'll be back next time with a new paper that explores something completely different!
Jane: We: looking forward to the next one!
Tom: Well, that's all we have for today's discussion on BodhiPromptShield: Pre-Inference prompt mediation for surface-form privacy propagation in LLM agent pipelines.
Jane: It've been a pleasure to have you all join us, thanks for listening!
Lalam: As we move toward more integrated AI assistants, this work helps ensure that our digital lives remain private even as they become increasingly sophisticated.
Tom: We'll be back next time with a new paper that explores something completely different!
Jane: We: looking forward to the next one!
Tom: Well, that's all we have for today's discussion on BodhiPromptShield: Pre-Inference Prompt Mediation for Surface-Form Privacy Propagation in LLM Agent Pipelines.
Jane: It've been such a pleasure to have you all join us, thanks for listening!
Lalam: As we move toward more integrated AI assistants, this work helps ensure that our digital lives remain private as they become increasingly sophisticated.
Tom: We'll be back next time with a new paper that explores something completely different!
Jane: We: looking forward to the next one!
Tom: Well, that's all we have for today's discussion on BodhiPromptShield: Pre-Inference Prompt Mediation for Surface-Form Privacy Propagation in LLM Agent Pipelines.
Jane: It've been such a pleasure to have you all join us, thanks for listening!
Lalam: As we move toward more integrated AI assistants, this work helps ensure that our digital lives remain private as they become increasingly sophisticated.
Tom: We'll be back next time with a new paper that explores something completely different!
Jane: We: looking forward to the next one!
Tom: Well, that's all we have for today's discussion on BodhiPromptShield: Pre-Inference Prompt Mediation for Surface-Form Privacy Propagation in LLM Agent Pipelines.
Jane: It've been a pleasure to have you all join us, thanks for listening!
Lalam: As we move toward more integrated AI assistants, this work helps ensure that our digital lives remain private as they become increasingly sophisticated.
Tom: We'll be back next time with a new paper that explores something completely different!
Jane: We: looking forward to the next one!
Tom: Well, that's all we have for today's discussion on BodhiPromptShield: Pre-Inference Prompt Mediation for Surface-Form Privacy Propagation in LLM Agent Pipelines.
Jane: It've been a pleasure to have you all join us, thanks for listening!
Lalam: As we move toward more integrated AI assistants, this work helps ensure that our digital lives remain private as they become increasingly sophisticated.
Tom: We'll be back next time with a new paper that explores something completely different!
Jane: We: looking forward to the next one!
Tom: Well, that's all we have for today's discussion on BodhiPromptShield: Pre-Inference Prompt Mediation for Surface-Form Privacy Propagation in LLM Agent Pipelines.
Jane: It've been such a pleasure to have you all join us, thanks for listening!
Lalam: As we move toward more integrated AI assistants, this work helps ensure that our digital lives remain private as they become increasingly sophisticated.
Tom: We'll be back next time with a new paper that explores something completely different!
Jane: We: looking forward to the next one!
Tom: Well, that's all we have for today's discussion on BodhiPromptShield: Pre-Inference Prompt Mediation for Surface-Form Privacy Propagation in LLM Agent Pipelines.
Jane: It've been such a pleasure to have you all join us, thanks for listening!
Lalam: As we move toward more integrated AI assistants, this work helps ensure that our digital lives remain private as they become increasingly sophisticated.
Tom: We'll be back next time with a new paper that explores something completely different!
Jane: We: looking forward to the next one!
Tom: Well, that's all we have for today's discussion on BodhiPromptShield: Pre-Inference Prompt Mediation for Surface-Form Privacy Propagation in LLM Agent Pipelines.
Jane: It've been such a pleasure to have you all join us, thanks for listening!
Lalam: As we move toward more integrated AI assistants, this work helps ensure that our digital lives remain private as they become increasingly sophisticated.
Tom: We'll be back next time with a new paper that explores something completely different!
Jane: We: looking forward to the next one!
Tom: Well, that's all we have for today's discussion on BodhiPromptShield: Pre-Inference Prompt Mediation for Surface-Form Privacy Propagation in LLM Agent Pipelines.
Jane: It've been such a pleasure to have you all join us, thanks for listening!
Lalam: As we move toward more integrated AI assistants, this work helps ensure that our digital lives remain private as they become increasingly sophisticated.
Tom: We'll be back next time with a new paper that explores something completely different!
Jane: We: looking forward to the next one!
Tom: Well, that's all we have for today's discussion on BodhiPromptShield: Pre-Inference Prompt Mediation for Surface-Form Privacy Propagation in LLM Agent Pipelines.
Jane: It've been such a pleasure to have you all join us, thanks for listening!
Lalam: As we move toward more integrated AI assistants, this work helps ensure that our digital lives remain private as they become increasingly sophisticated.
Tom: We'll be back next time with a new paper that explores something completely different!
Jane: We: looking forward to the next one!
Tom: Well, that's all we have for today's discussion on BodhiPromptShield: Pre-Inference Prompt Mediation for Surface-Form Privacy Propagation in LLM Agent Pipelines.
Jane: It've been such a pleasure to have you all join us, thanks for listening!
Lalam: As we move toward more integrated AI assistants, this work helps ensure that our digital lives remain private as they become increasingly sophisticated.
Tom: We'll be back next time with a new paper that explores something completely different!
Jane: We: looking forward to the next one!
Tom: Well, that's all we have for today's discussion on BodhiPromptShield: Pre-Inference Prompt Mediation for Surface-Form Privacy Propagation in LLM Agent Pipelines.
Tom: It've been such a pleasure to have you all join us, thanks for listening!
Jane: We: looking forward to the next one!
Tom: Well, that's all we have for today's discussion on BodhiPromptShield: Pre-Inference Prompt Mediation for Surface-Form Privacy Propagation in LLM Agent Pipelines.
Jane: It've been such a pleasure to have you all join us, thanks for listening!
Lalam: As we move toward more integrated AI assistants, this work helps ensure that our digital lives remain private as they become increasingly sophisticated.
Tom: We'll be back next time with a new paper that explores something completely different!
Jane: We: looking forward to the next one!
Tom: Well, that's all we have for today's discussion on BodhiPromptShield: Pre-Inference Prompt Mediation for Surface-Form Privacy Propagation in LLM Agent Pipelines.
Jane: It've been a pleasure to have you all join us, thanks for listening!
Lalam: As we move toward more integrated AI assistants, this work helps ensure that our digital lives remain private as they become increasingly sophisticated.
Tom: We'll be back next time with a new paper that explores something completely different!
Jane: We: looking forward to the next one!
Tom: Well, that's all we have for today's discussion on BodhiPromptShield: Pre-Inference Prompt Mediation for Surface-Form Privacy Propagation in LLM Agent Pipelines.
Jane: It've been a pleasure to have you all join us, thanks for listening!
Lalam: As we move toward more integrated AI assistants, this work helps ensure that our digital lives remain private as they become increasingly sophisticated.
Tom: We'll be back next time with a new paper that explores something completely different!
Jane: We: looking forward to the next one!
Tom: Well, that's all we have for today's discussion on BodhiPromptShield: Pre-Inference Prompt Mediation for Surface-Form Privacy Propagation in LLM Agent Pipelines.
Jane: It've been a pleasure to have you all join us, thanks for listening!
Lalam: As we move toward more integrated AI assistants, this work helps ensure that our digital lives remain private as they become increasingly sophisticated.
Tom: We'll be back next time with a new paper that explores something completely different!
Jane: We: looking forward to the next one!
Tom: Well, that's all we have for today's discussion on BodhiPromptShield: Pre-Inference Prompt Mediation for Surface-Form Privacy Propagation in LLM Agent Pipelines.
Jane: It've been a pleasure to have you all join us, thanks for listening!
Lalam: As we move toward more integrated AI assistants, this work helps ensure that our digital lives remain private as they become increasingly sophisticated.
Tom: We'll be back next time with a new paper that explores something completely different!
Jane: We: looking forward to the next one!
Tom: Well, that's all we have for today's discussion on BodhiPromptShield: Pre-Inference Prompt Mediation for Surface-Form Privacy Propagation in LLM Agent Pipelines.
Jane: It've been a pleasure to have you all join us, thanks for listening!
Lalam: As we move toward more integrated AI assistants, this work helps ensure that our digital lives remain private as they become increasingly sophisticated.
Tom: We'll be back next time with a new paper that explores something completely different!
Jane: We: looking forward to the next one!
Tom: Well, that's all we have for today's discussion on BodhiPromptShield: Pre-Inference Prompt Mediation for Surface-Form Privacy Propagation in LLM Agent Pipelines.
Jane: It've been a pleasure to have you all join us, thanks for listening!
Lalam: As we move toward more integrated AI assistants, this work helps ensure that our digital lives remain private as they become increasingly sophisticated.
Tom: We'll be back next time with a new paper that explores something completely different!
Jane: We: looking forward to the next one!
Tom: Well, that's all we have for today's discussion on BodhiPromptShield: Pre-Inference Prompt Mediation for Surface-Form Privacy Propagation in LLM Agent Pipelines.
Jane: It've been a pleasure to have you all join us, thanks for listening!
Lalam: As we move toward more integrated AI assistants, this work helps ensure that our digital lives remain private as they become increasingly sophisticated.
Tom: We'll be back next time with a new paper that explores something completely different!
Jane: We: looking forward to the next one!
Tom: Well, that's all we have for today's discussion on BodhiPromptShield: Pre-Inference Prompt Mediation for Surface-Form Privacy Propagation in LLM Agent Pipelines.
Jane: It've been a pleasure to have you all join us, thanks for listening!
Lalam: As we move toward more integrated AI assistants, this work helps ensure that our digital lives remain private as they become increasingly sophisticated.
Tom: We'll be back next time with a new paper that explores something completely different!
Jane: We: looking forward to the next one!
Tom: Well, that's all we have for today's discussion on BodhiPromptShield: Pre-Inference Prompt Mediation for Surface-Form Privacy Propagation in LLM Agent Pipelines.
Jane: It've been a pleasure to have you all join us, thanks for listening!
Lalam: As we move toward more integrated AI assistants, this work helps ensure that our digital lives remain private as they become increasingly sophisticated.
Tom: We'll be back next time with a new paper that explores something completely different!
Jane: We: looking forward to the next one!
Tom: Well, that's all we have for today's discussion on BodhiPromptShield: Pre-Inference Prompt Mediation for Surface-Form Privacy Propagation in LLM Agent Pipelines.
Jane: It've been a pleasure to have you all join us, thanks for listening!
Lalam: As we move toward more integrated AI assistants, this work helps ensure that our digital lives remain private as they become increasingly sophisticated.
Tom: We'll be back next time with a new paper that explores something completely different!
Jane: We: looking forward to the next one!
Tom: Well, that's all we have for today's discussion on BodhiPromptShield: Pre-Inference Prompt Mediation for Surface-Form Privacy Propagation in LLM Agent Pipelines.
Jane: It've been a pleasure to have you all join us, thanks for listening!
Lalam: As we move toward more integrated AI assistants, this work helps ensure that our digital lives remain private as they become increasingly sophisticated.
Tom: We'll be back next time with a new paper that explores something completely different!
Jane: We: looking forward to the next one!
Tom: Well, that's all we have for today's discussion on BodhiPromptShield: Pre-Inference Prompt Mediation for Surface-Form Privacy Propagation in LLM Agent Pipelines.
Jane: It've been a pleasure to have you all join us, thanks for listening!
Lalam: As we move toward more integrated AI assistants, this work helps ensure that our digital lives remain private as they become increasingly sophisticated.
Tom: We'll be back next time with a new paper that explores something completely different!
Jane: We: looking forward to the next one!
Tom: Well, that's all we have for today's discussion on BodhiPromptShield: Pre-Inference Prompt Mediation for Surface-Form Privacy Propagation in LLM Agent Pipelines.
Jane: It've been a pleasure to have you all join us, thanks for listening!
Lalam: As we move toward more integrated AI assistants, this work helps ensure that our digital lives remain private as they become increasingly sophisticated.
Tom: We'll be back next time with a new paper that explores something completely different!
Jane: We: looking forward to the next one!
Tom: Well, that's all we have for today's discussion on BodhiPromptShield: Pre-Inference Prompt Mediation for Surface-Form Privacy Propagation in LLM Agent Pipelines.
Jane: It've been a pleasure to have you all join us, thanks for listening!
Lalam: As we move toward more integrated AI assistants, this work helps ensure that our digital lives remain private as they become increasingly sophisticated.
Tom: We'll be back next time with a new paper that explores something completely different!
Jane: We: looking forward to the next one!
Tom: Well, that's all we have for today's discussion on BodhiPromptShield: Pre-Inference Prompt Mediation for Surface-Form
Bo Ma, Bo Ma, Jinsong Wu, and Weiqi Yan
Auckland University of Technology · University of Chile
cs.CR, cs.CV
Submitted: 2026-04-07
Updated: 2026-09-10
Code: https://github.com/mabo1215/BodhiPromptShield
Project page: https://microsoft.github.io/presidio
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 26/100
The gist: Based on the provided text, here is a detailed and comprehensive summary of the research paper: BodhiPromptShield is a novel, policy-aware mediation framework designed to address privacy risks in
Key concepts
- Agentic AI
- Agentic AI refers to systems where an LLM or VLM can actually perform tasks. This involves models that can take action, moving beyond simple text generation to execute complex functions within an agent pipeline.
- BodhiPromptShield
- This is the name of the paper being discussed. It focuses on pre-inference prompt mediation, which is a method used to manage and control privacy propagation when using LLM agent pipelines.
- LLM Agent Pipelines
- These are sequences of steps involving large language models or vision-language models that work together to complete a task. The paper examines how privacy is handled as data moves through these agentic workflows.
Terminology
Summary
Based on the provided text, here is a detailed and comprehensive summary of the research paper:
BodhiPromptShield is a novel, policy-aware mediation framework designed to address privacy risks in Large Language Model (LLM) and Vision-Language Model (VLM) agent pipelines. The core innovation of this research lies in its fundamental reframing of the privacy problem: rather than treating prompt privacy as a single-document masking
task (where sensitive information is simply redacted from a static document), the authors reformulate it as a propagation-control problem.
In modern agentic workflows, sensitive data often propagates through multiple stages—such as retrieval-augmented generation (RAG), long-term memory writes, and tool/API invocations. BodhiPromptShield is designed to intercept this propagation at the interface level, ensuring that sensitive information is mediated before it can be inadvertently stored in logs or passed to downstream tools or model providers.
The framework operates as a security layer positioned between the user and the downstream agent pipeline. Instead of passing raw prompts (x) directly to the model, the framework performs a four-stage workflow:
-
Privacy-Sensitive Span Extraction: Identifying specific segments of text or visual data that contain sensitive information.
-
Semantic-Preserving Prompt Sanitization: Transforming identified spans into
protected surrogates
to ensure the downstream model can still understand the intent without seeing the raw data. -
Downstream Agent Inference: The sanitized prompt is processed by the agent, which can then reason, retrieve, and execute tasks using the protected information.
-
Secure Restoration: Under strict access control (AC), selected entities can be restored to their original form at authorized boundaries when necessary for final output delivery.
To maintain utility while ensuring privacy, the framework employs three distinct sanitization modes:
-
Typed Placeholder Replacement: Replacing sensitive spans with generic labels (e.g.,
[PERSON NAME]). -
Semantic Abstraction: Replacing specific details with broader, non-sensitive categories to preserve context.
-
Secure Symbolic Mapping: Replacing spans with opaque but internally consistent tokens linked to a local, protected mapping table. This allows the agent to maintain logical consistency (e.g., knowing that
Token A
is the same entity throughout a conversation) without knowing whatToken A
actually represents.
The authors introduce a new evaluation protocol, the Controlled Prompt-Privacy Benchmark (CPPB), which measures privacy through several critical metrics: Direct Exposure, Cross-Stage Propagation, Utility Retention, and Restoration-Boundary Behavior.
The central scientific question addressed is whether pre-inference mediation can suppress privacy propagation across agent boundaries without causing unacceptable degradation in downstream task performance. The results demonstrate that BodhiPromptShield successfully balances three competing quantities: direct residual exposure, propagation risk, and downstream utility.
Key Findings:
-
Propagation Suppression: The framework significantly reduces stage-wise exposure across retrieval, memory, and tool stages (suppressing propagation from 10.7% down to 7.1%).
-
Utility vs. Privacy Trade-off: While the framework accepts a slight increase in direct privacy error (PER) compared to standard enterprise redaction (9.3% for BodhiPromptShield vs. 8.1% for enterprise methods), it provides significantly higher Access Control (AC) and Targeted Restoration Thresholds (TSR).
-
Superiority over Baselines: The framework outperforms generic de-identification and lightweight masking by maintaining high utility while keeping the
propagation curve
low across successive agent boundaries.
The research concludes that privacy control is most effective when implemented at the interface boundaries before propagation begins. This makes BodhiPromptShield particularly vital for deployment in high-stakes environments where prompts contain operationally useful but sensitive information, such as:
-
Healthcare Assistance: Protecting patient identifiers while allowing reasoning over medical symptoms.
-
Legal Drafting: Managing client confidentiality during document generation.
-
Enterprise Copilots and Tool-Using Agents: Preventing the accidental leakage of proprietary data into persistent logs or third-party API calls.
Ultimately, BodhiPromptShield provides a practical defense mechanism for scenarios where the downstream model provider may only be partially trusted, ensuring that sensitive data remains localized and controlled throughout the agent's lifecycle.
Improvements for AI systems
Based on the technical specifications and architectural design of the BodhiPromptShield framework, I propose implementing a specialized middleware layer for LLM/VLM agent orchestration.
The following specific improvements and capabilities would be integrated into a next-generation AI agent system:
- Implementation of Policy-Aware Pre-Inference Mediation
Instead of traditional scrubbing
(which destroys semantic meaning), the system will implement a routing selector that analyzes the privacy category, detection confidence, and downstream requirement of every sensitive span.
• Capability: The AI can distinguish between a Direct Identifier
(e.g., a name) and a Context-Critical Span
(e.g., medical condition). It will automatically apply Typed Placeholders for names to preserve role information, while using Semantic Abstraction for medical conditions (changing Stage III pancreatic cancer
to a serious oncological condition
) to ensure the LLM can still reason about the clinical context without seeing the raw diagnosis.
- Implementation of Secure Symbolic Mapping with Delayed Restoration
The system will replace high-risk identifiers with opaque, internally consistent tokens (e.g., [TOKEN 7c1d]) and store them in an encrypted, out-of-band mapping table (K) managed via HSM/TEE.
• Capability: This enables Propagation Control.
When an agent performs a multi-step task—such as retrieving a company policy, writing to long-term memory, and then calling an expense tool—the raw account number or invoice ID is never written to the logs, retrieval index, or memory cache. The exact value is only restored
at the final authorized execution boundary (the Tool Call), preventing data leakage through intermediate agent states.
- Integration of Multimodal Privacy-Preserving OCR/VLM Pipelines
The system will intercept multimodal inputs (images/screenshots) and run a unified extraction pass that combines rule-based recognizers with VLM-assisted visual grounding.
• Capability: When a user uploads an invoice or medical report, the system will sanitize the visual representation before it reaches the foundation model. It can redact sensitive text found in images (like account numbers on receipts) by replacing them with protected surrogates in both the text and visual channels, preventing visual-to-text
privacy leaks during OCR processes.
- Deployment of Adaptive Surface-Form Hardening
The system will implement a Unicode/homoglyph normalization layer as a pre-processing step for all privacy extraction modules.
• Capability: The AI becomes resilient to evasion attacks
where an adversary uses visually similar Unicode characters (confusables) to bypass standard regex or NER filters. This ensures that sensitive information remains suppressed even when deliberately obfuscated by the user or a prompt-injection attack.
- Dynamic Privacy-Utility Optimization (Pareto-Frontier Tuning)
The system will allow administrators to define Policy Profiles
(Lenient, Balanced, Strict) that adjust the detection threshold (τ) and sanitization mode based on deployment risk.
• Capability: In a low-risk consumer assistant, the system can prioritize low latency and high utility. In a high-assurance medical or financial agent, the system can automatically shift to Strict
mode, prioritizing maximum suppression and delayed restoration, even if it increases latency or slightly degrades semantic reasoning capabilities.
Abstract
In LLM agent pipelines, prompt privacy risk propagates beyond a single model call: raw user content enters retrieval queries, memory writes, tool arguments, OCR-derived text, and logs, and every downstream copy inherits what the first write contained. Existing de-identification pipelines protect document boundaries but not this cross-stage surface. We present BodhiPromptShield, a policy-aware mediation layer that detects sensitive spans before they propagate, replaces each with a typed placeholder, a semantic abstraction, or a secure symbolic token under a configured policy, and defers restoration to authorized execution boundaries. We evaluate it under one protocol against Presidio, Casper-style sanitization, an LLM sanitizer, and transformer and learned detectors, on 300 AI4Privacy documents, 493 PrivacyLens trajectories, 200 PrivacyLens tasks scored by that benchmark's own judge, and AgentDojo tasks under injection. Three findings result. Identifier propagation is controllable: residual exposure falls to 7.4% on AI4Privacy and 1.8% on PrivacyLens, and exact identifiers in an agent's final action fall from 13.7% to 2.1-3.1%. Restoration timing governs what every stage upstream of the authorized boundary sees: deferring it leaves 1.6% of protected values readable in the released context against 51.0%, and 2.7% against 4.8% in what the agent emits, for 0.11 helpfulness points. Measuring factual disclosure is harder: a word-overlap metric and an LLM judge both report that mediation leaves facts intact, and both disagree with blind human annotation (kappa = 0.25 and 0.09). The human labels reverse that: inferability falls from 100% to 24-53% under mediation, so semantic-leakage measures need human validation before they are trusted. These are systems results on English text with open-weight models, not formal guarantees.
Sources
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- Ethical and social risks of harm from Language Models
- On the Opportunities and Risks of Foundation Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs