GPT-Red: Automated Red Teaming via Self-Play at Scale
Eric Wallace, Christopher A. Choquette-Choo, Nikhil Kandpal, Sam Toyer, Dylan Hunn, Stephanie Lin, Yuxin Wen, Xiangyu Qi, Christopher Wolff, Zizhao Wang, Milad Nasr, Sicheng Zhu, Chuan Guo, Juan Felipe Cerón Uribe, Kaiwen Wang, Aiden Low, Kai Xiao, Kai Chen
cs.CR, cs.AI, cs.CL, cs.LG
Submitted: 2026-07-28
Comments: 28 pages.13 main pages and 13 main figures
Code: https://github.com/UKGovernmentBEIS/inspect_ai
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Constitutional AI: Harmlessness from AI Feedback
- Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning
- Jailbreaking Black Box Large Language Models in Twenty Queries
- AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
- Enhancing LLM Safety Through a Theoretical Minimax Game Lens
- How Vulnerable Are AI Agents to Indirect Prompt Injections? Insights from a Large-Scale Public Competition
- MART: Improving LLM Safety with Multi-round Automatic Red-Teaming
- Explaining and Harnessing Adversarial Examples
- Safety Alignment of LMs via Non-cooperative Games
- Deliberative Alignment: Reasoning Enables Safer Language Models
- Red Teaming Language Models with Language Models
- IH-Challenge: A Training Dataset to Improve Instruction Hierarchy on Frontier LLMs
- Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models
- Lessons from Defending Gemini Against Indirect Prompt Injections
- MUZZLE: Adaptive Agentic Red-Teaming of Web Agents Against Indirect Prompt Injection Attacks
- Evolving Diverse Red-team Language Models in Multi-round Multi-agent Games
- Towards Deep Learning Models Resistant to Adversarial Attacks
- The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
- Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
- AgentVigil: Generic Black-Box Red-teaming for Indirect Prompt Injection against LLM Agents
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs