Cross-Agent Campaign Attribution: Linking Asynchronous Attacks Across LLM Agents
SangJin Park, Myungsub Choi, Jineok Kim, Minseung Kang
cs.CR, cs.AI
Submitted: 2026-07-21
Comments: 22 pages, 5 figures. Accepted at the Second Workshop on Agents in the Wild: Safety, Security, and Beyond (AIWILD) at ICML 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: LLM-agent defenses are typically evaluated one session at a time.
Terminology
Abstract
LLM-agent defenses are typically evaluated one session at a time. In deployment, however, attacks can be distributed across independent agents, teams, and runtimes, leaving each local guardrail with only a sparse fragment. We formalize cross-agent asynchronous campaign attribution: linking sessions from the same latent adversarial campaign without shared runtime state, test-time campaign labels, or attacker identity oracles. We introduce Asynchronous Attribution Fingerprint Vectors (A 2FV), a lightweight proxy-side reference protocol for scoring pairwise campaign similarity from proxy-observable tool-use, timing, and prompt residue. We also construct SCD-v1, a controlled persona-matched benchmark with benign traffic, isolated attacks, multi-session campaigns, matched non-oracle evasion, and leakage audits. On SCD-v1, A 2FV achieves 0.82 pairwise AUC for campaign linking, while score-only adaptations of per-session detectors and chunked LLM judges remain near chance under the same task. The strongest fixed signal is carried by structural and stylometric residue, while timing is retained as a diagnostic channel for richer proxy traces. Crossed-style controls show that the signal is partly style-sensitive but not reducible to style alone. Static and dimension-aware non-oracle stress tests further show that pairwise separability persists under controlled evasion. These results establish cross-agent campaign attribution as a distinct evaluation layer for securing LLM agents in the wild.
Sources
- Cross-Session Threats in AI Agents: Benchmark, Evaluation, and Algorithms
- Kairos: Practical Intrusion Detection and Investigation using Whole-system Provenance
- Defeating Prompt Injections by Design
- AgentOps: Enabling Observability of LLM Agents
- Memory Injection Attacks on LLM Agents via Query-Only Interaction
- Attention Tracker: Detecting Prompt Injection Attacks in LLMs
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Securing the Model Context Protocol (MCP): Risks, Controls, and Governance
- WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
- PromptShield: Deployable Detection for Prompt Injection Attacks
- BinaryShield: Cross-Service Threat Intelligence in LLM Services using Privacy-Preserving Fingerprints
- A Critical Evaluation of Defenses against Prompt Injection Attacks
- Seven Security Challenges in Cross-domain Multi-agent LLM Systems
- Prompt Infection: LLM-to-LLM Prompt Injection within Multi-Agent Systems
- NODLINK: An Online System for Fine-Grained APT Attack Detection and Investigation
- Prompt Injection Detection and Mitigation via AI Multi-Agent NLP Frameworks
- HADES: Detecting Active Directory Attacks via Whole Network Provenance Analytics
- AttriGuard: Defeating Indirect Prompt Injection in LLM Agents via Causal Attribution of Tool Invocations
- A Multi-Agent LLM Defense Pipeline Against Prompt Injection Attacks
- DataSentinel: A Game-Theoretic Detection of Prompt Injection Attacks
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs