Early Detection of Distributed Backdoors in Multi-Agent LLM Systems: A Characterization Study
cs.CR, cs.AI
Submitted: 2026-07-27
Updated: 2026-10-02
Code: https://github.com/yibo-hu-lab/distributed-backdoor-early-warning
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security
- TrojanRAG: Retrieval-Augmented Generation Can Be Backdoor Driver in Large Language Models
- BackdoorAgent: A Unified Framework for Backdoor Attacks on LLM-based Agents
- SkillTrojan: Backdoor Attacks on Skill-Based Agent Systems
- PrefixGuard: Online Failure Warning and Trace-Grounded Diagnosis for LLM Agents
- HINTBench: Horizon-agent Intrinsic Non-attack Trajectory Benchmark
- AgentForesight: Online Auditing for Early Failure Prediction in Multi-Agent Systems
- Collaborative Shadows: Distributed Backdoor Attacks in LLM-Based Multi-Agent Systems
- DemonAgent: Dynamically Encrypted Multi-Backdoor Implantation Attack on LLM-based Agent
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs