Policy-Backed Selective Regeneration under Tainted Inter-Agent Communication
cs.CR, cs.AI
Submitted: 2026-08-02
Updated: 2026-08-02
License: http://creativecommons.org/licenses/by/4.0/
The gist: Inter-agent communication is essential to multi-agent language-model systems, yet a single message may combine task-critical information with instructions not authorized by the original request.
Terminology
Abstract
Inter-agent communication is essential to multi-agent language-model systems, yet a single message may combine task-critical information with instructions not authorized by the original request. Prompt-based defenses leave enforcement to models exposed to adversarial messages, while indiscriminate message removal discards useful information. We introduce Executable Semantic Commitments with Clean-Room Recovery (ESC-CR), a policy-backed framework for secure inter-agent code generation and recovery. It separates message claims from authorization, constructs executable commitments from trusted tasks, evidence, and policy, and enforces them at an external release boundary. Upon a violation, ESC-CR taints the responsible message and rejected artifact, reconstructs a clean context from evidence-backed task information, and regenerates under the same policy. We evaluate ESC-CR across communication-essential and standard code-generation benchmarks, multiple model families and communication topologies, and adaptive attacks spanning direct, obfuscated, and verifier-aware payloads. Results show that polluted-context retry frequently fails to remove unauthorized influence, while complete message removal can discard information required by communication-essential tasks. ESC-CR preserves evidence-backed claims while suppressing unauthorized releases under matched computational budgets, and the same design transfers to end-to-end agent trajectories.
Sources
- Program Synthesis with Large Language Models
- Evaluating Large Language Models Trained on Code
- StruQ: Defending Against Prompt Injection with Structured Queries
- LlamaFirewall: An open source guardrail system for building secure AI agents
- Securing AI Agents with Information-Flow Control
- Defeating Prompt Injections by Design
- Taming OpenClaw: Security Analysis and Mitigation of Autonomous LLM Agent Threats
- AgentLeak: A Benchmark for Internal-Channel Privacy Leakage in Multi-Agent LLM Systems
- Optimizing Agent Planning for Security and Autonomy
- Breaking the Protocol: Security Analysis of the Model Context Protocol Specification and Prompt Injection Vulnerabilities in Tool-Integrated LLM Agents
- PAC-BENCH: Evaluating Multi-Agent Collaboration under Privacy Constraints
- Benefits and Limitations of Communication in Multi-Agent Reasoning
- Progent: Securing AI Agents with Privilege Control
- PromptArmor: Simple yet Effective Prompt Injection Defenses
- The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
- AgentSpec: Customizable Runtime Enforcement for Safe and Reliable LLM Agents
- Contextualized Privacy Defense for LLM Agents
- Multi-User Large Language Model Agents
- AgenTracer: Who Is Inducing Failure in the LLM Agentic Systems?
- Poison Once, Exploit Forever: Environment-Injected Memory Poisoning Attacks on Web Agents
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs