An AI Agent Execution Environment to Safeguard User Data
cs.CR, cs.AI, cs.OS
Submitted: 2026-04-21
Updated: 2026-09-24
Code: https://github.com/DonutShinobu/claude-code-fork
Project page: https://openai.github.io/openai-agents-python/guardrails
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
- Memory Injection Attacks on LLM Agents via Query-Only Interaction
- Text-Based Personas for Simulating User Privacy Decisions
- Personalizing Agent Privacy Decisions via Logical Entailment
- Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
- Can LLMs Make (Personalized) Access Control Decisions?
- Defending Against Indirect Prompt Injection Attacks With Spotlighting
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Preventing Prompt Injection with Type-Directed Privilege Separation
- Baseline Defenses for Adversarial Attacks Against Aligned Language Models
- Detecting Language Model Attacks with Perplexity
- AirGapAgent: Protecting Privacy-Conscious Conversational Agents
- AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases
- LlamaFirewall: An open source guardrail system for building secure AI agents
- Securing AI Agents with Information-Flow Control
- Defeating Prompt Injections by Design
- WebGPT: Browser-assisted question-answering with human feedback
- The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
- Formal Policy Enforcement for Real-World Agentic Systems
- Red Teaming the Mind of the Machine: A Systematic Evaluation of Prompt Injection and Jailbreak Vulnerabilities in LLMs
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs