Self-Evolving Defense: Continual Security Policy Learning for LLM Agents
cs.CR
Submitted: 2026-09-29
Updated: 2026-09-29
Code: https://github.com/Infini-AI-Lab/SED
Project page: https://infini-ai-lab.github.io/SED
Terminology
Sources
- DecodingTrust-Agent Platform (DTap): A Controllable and Interactive Red-Teaming Platform for AI Agents
- Defeating Prompt Injections by Design
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- GLM-5: from Vibe Coding to Agentic Engineering
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
- Kimi K3: Open Frontier Intelligence
- AgentDyn: Are Your Agent Security Defenses Deployable in Real-World Dynamic Environments?
- Adversarial Attack-Defense Co-Evolution for LLM Safety Alignment via Tree-Group Dual-Aware Search and Optimization
- Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models
- AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs
- The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
- AdvEvo-MARL: Shaping Internalized Safety through Adversarial Co-Evolution in Multi-Agent Reinforcement Learning
- X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-Agents
- MemoryGraft: Persistent Compromise of LLM Agents via Poisoned Experience Retrieval
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- Be Your Own Red Teamer: Safety Alignment via Self-Play and Reflective Experience Replay
- A-MemGuard: A Proactive Defense Framework for LLM-Based Agent Memory
- SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs