Isolation as a First-Class Principle for LLM-Agent System Safety: Concepts, Taxonomy, Challenges and Future Directions
cs.AI
Submitted: 2026-07-14
Updated: 2026-09-02
Code: https://github.com/openclaw/openclaw
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Do Multimodal RAG Systems Leak Data? A Comprehensive Evaluation of Membership Inference and Image Caption Retrieval Attacks
- The Obvious Invisible Threat: LLM-Powered GUI Agents' Vulnerability to Fine-Print Injections
- HWE-Bench: Benchmarking LLM Agents on Real-World Hardware Bug Repair Tasks
- Trojan Hippo Bench: A Dynamic Benchmark for Persistent Memory Attacks and Defenses in LLM Agents
- MEMSAD: Gradient-Coupled Anomaly Detection for Memory Poisoning in Retrieval-Augmented Agents
- AttriGuard: Defeating Indirect Prompt Injection in LLM Agents via Causal Attribution of Tool Invocations
- Parallax: Why AI Agents That Think Must Never Act
- The Attack and Defense Landscape of Agentic AI: A Comprehensive Survey
- Easier Said Than Done: Unpacking Intent-Behavior Gap in Jailbreaking LLM-Based Robots
- ToolSafe: Enhancing Tool Invocation Safety of LLM-based agents via Proactive Step-level Guardrail and Feedback
- MOSAIC-Bench: Measuring Compositional Vulnerability Induction in Coding Agents
- AgentSentry: Mitigating Indirect Prompt Injection in LLM Agents via Temporal Causal Diagnostics and Context Purification
- AgentVisor: Defending LLM Agents Against Prompt Injection via Semantic Virtualization
- AutoDAN: Interpretable Gradient-Based Adversarial Attacks on Large Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection