Out of Sight, Not Out of Mind: Unveiling Latent Attack in Latent-based Multi-Agent Systems
cs.CR, cs.LG, cs.MA
Submitted: 2026-05-27
Updated: 2026-08-27
Code: https://github.com/mnmn-f/Out-of-Sight-LatentAttack
Terminology
Sources
- Compressed Chain of Thought: Efficient Reasoning Through Dense Representations
- Training Verifiers to Solve Math Word Problems
- Enabling Agents to Communicate Entirely in Latent Space
- The Llama 3 Herd of Models
- Prompt Injection attack against LLM-integrated Applications
- Ignore Previous Prompt: Attack Techniques For Language Models
- Qwen3 Technical Report
- Steering Language Models With Activation Engineering
- SafeAgentBench: A Benchmark for Safe Task Planning of Embodied LLM Agents
- Representation Engineering: A Top-Down Approach to AI Transparency
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- Latent Collaboration in Multi-Agent Systems
- The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook
- Agent-SafetyBench: Evaluating the Safety of LLM Agents
- A Survey on Latent Reasoning
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs