AgentXploit: Autonomous Repository-to-Runtime Red-Teaming for AI Agents
cs.CR, cs.AI
Submitted: 2026-09-25
Updated: 2026-09-25
Code: https://github.com/lwd17/AgentXploit
Terminology
Sources
- StruQ: Defending Against Prompt Injection with Structured Queries
- SecAlign: Defending Against Prompt Injection with Preference Optimization
- AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
- A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis
- The Emerged Security and Privacy of LLM Agent: A Survey with Case Studies
- Co-RedTeam: Orchestrated Security Discovery and Exploitation with LLM Agents
- Defending Against Indirect Prompt Injection Attacks With Spotlighting
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security
- EIA: Environmental Injection Attack on Generalist Web Agents for Privacy Leakage
- DataSentinel: A Game-Theoretic Detection of Prompt Injection Attacks
- WebGPT: Browser-assisted question-answering with human feedback
- Gorilla: Large Language Model Connected with Massive APIs
- Ignore Previous Prompt: Attack Techniques For Language Models
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
- Progent: Securing AI Agents with Privilege Control
- PromptArmor: Simple yet Effective Prompt Injection Defenses
- The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs