When Local Monitors Miss Compositional Harm: Diagnosing Distributed Backdoors in Multi-Agent Systems
Yibo Hu, Ren Wang
cs.CR, cs.LG, cs.MA
Submitted: 2026-07-13
Code: https://github.com/yibo-hu-lab/observabilityboundary
License: http://creativecommons.org/licenses/by/4.0/
The gist: As multi-agent, tool-using LLM systems are deployed, a common safety net is a runtime monitor that checks each message, tool call, or step on its own.
Terminology
Abstract
As multi-agent, tool-using LLM systems are deployed, a common safety net is a runtime monitor that checks each message, tool call, or step on its own. We show this net has a fundamental hole. A distributed backdoor splits a harmful payload across agents, so every local check passes while the assembled object is the attack. The monitor can be right on every step and still miss the attack. The problem is not splitting itself: split fragments can still leak suspicious tokens or provenance edges. The hard case is local benignness. No fragment carries the harm, and what is left looks like ordinary benign traffic. We formalize this as an observability boundary: a monitor catches only what its view can tell apart from benign traffic. We prove that once the fragments look benign in the monitored view, no detector on that view can catch them, however strong it is. Across a controlled testbed, an external benchmark, and end-to-end agent runs, local monitors lose the signal exactly as local evidence disappears, and it returns only when the monitor sees the assembled object. A monitor trained only on benign traffic recovers the attack's code structure across held-out encodings (0.874 mean AUROC). A decoded-view gate, given the encoding family, blocks every tested attack. But seeing more is not enough: full-trace monitors and decoders still fail unless they reach the representation where the payload is exposed. Local safety is not global safety when harm is compositional, and the open problem is finding that representation.
Sources
- Ghost in the Agent: Redefining Information Flow Tracking for LLM Agents
- TraceSafe: A Systematic Assessment of LLM Guardrails on Multi-Step Tool-Calling Trajectories
- SafetyDrift: Predicting When AI Agents Cross the Line Before They Actually Do
- PeerGuard: Defending Multi-Agent Systems Against Backdoor Attacks Through Mutual Reasoning
- SkillTrojan: Backdoor Attacks on Skill-Based Agent Systems
- BackdoorAgent: A Unified Framework for Backdoor Attacks on LLM-based Agents
- From Prompt Injections to Protocol Exploits: Threats in LLM-Powered AI Agents Workflows
- When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift
- LLM Censorship: A Machine Learning Challenge or a Computer Security Problem?
- When Better Codebooks Are Not Enough: Predictive Performance and Behavioral Reliability in LLM Political Event Coding
- Lying with Truths: Open-Channel Multi-Agent Collusion for Belief Manipulation via Generative Montage
- Most LLM Conformity Needs No Speaker: Measuring the Speaker-Free Floor in Peer-Pressure Benchmarks
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Breaking and Fixing Defenses Against Control-Flow Hijacking in Multi-Agent Systems
- Reliable Weak-to-Strong Monitoring of LLM Agents
- Prompt Infection: LLM-to-LLM Prompt Injection within Multi-Agent Systems
- Your Agent Can Defend Itself against Backdoor Attacks
- DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers
- AgentArmor: Enforcing Program Analysis on Agent Runtime Trace to Defend Against Prompt Injection
- Collaborative Shadows: Distributed Backdoor Attacks in LLM-Based Multi-Agent Systems
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs