Learning to Attack and Defend: Adaptive Red Teaming of Language Models via GRPO
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Learning to Attack and Defend".
Jane: The gist The AdvGRPO framework introduces a co-training method that makes Group Relative Policy Optimization (GRPO) viable for joint attacker-defender optimization using dense multi-channel rewards and decoupled advantage normalization,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: We're starting with the title of Learning to Attack and Defend: Adaptive Red Teaming of Language Models via GRPO. It immediately tells us this paper is focused on training both an attacker and a defender together, adaptively, using this GRPO method.
Jane: The authors are Blake Bullwinkel, Eugenia Kim, Amanda Minnich, and Mark Russinovich from Microsoft AI Red Team. They bring a lot of experience in the red teaming side of things to this work.
Lu: What they’re doing is tackling the fact that static methods for attacker-defender co-training using GRPO tend to be unstable, and they propose AdvGRPO as a way around that by adding those dense reward channels and the decoupled advantage normalization.
Meng: So they're taking an existing RL technique, GRPO, and making it stable enough to actually train two models in tandem for red teaming. How much stability are we talking about?
Lalam: It suggests that we can achieve better results than previous work because they’ve addressed the instability issues that have been reported with co-training methods like PPO and DPO, which is a key problem in this area.
The paper's summary: Tom: They summarize AdvGRPO as a curriculum that moves from single-turn training to closed-loop multi-turn attacks, after which they bootstrap the co-training by alternating updates between the attacker and defender models every N steps.
Jane: The core mechanism involves using dense, multi-channel rewards—they score responses on intent alignment, content harms, detail level for the attack reward A—and then decoupling advantage normalization for each channel.
Lu: They use four specific reward channels: the attack reward A, which measures how well the response satisfies the objective; a prompt reward P to check strategy adherence; a thinking-trace reward T to penalize failure modes like lack of commitment; and a helpfulness reward H for benign objectives when training the defender.
Meng: It seems they’re trying to score every aspect of the interaction, not just one overall success metric, which is smart because it gives you more feedback on *why* an attack worked or failed.
Lalam: And they leverage GDPO to normalize these channels independently before combining them with a weighted sum, which helps stabilize the training process by standardizing each reward channel before any advantage computation happens.
The paper's improvements: Tom: They highlight several improvements that AdvGRPO brings, specifically noting that it produces strong single-turn, multi-turn, and reasoning attackers that generalize well to unseen defenders and out-ofdistribution objectives.
Jane: A big improvement they point to is the transferability of these attacks; the paper shows they can produce highly effective attacks that transfer to unseen model families.
Lu: For reasoning models, they introduce the thinking-trace reward T, which allows the attacker to overcome self-censoring by evaluating conciseness and commitment during a rollout. They also show that co-trained defenders outperform baselines on safety benchmarks.
Meng: The numbers are pretty compelling here; in their best configuration, the attackers achieve ninety–ninety-one percent ASR on both benchmarks, and the defenders achieve the lowest ASR, reducing HarmBench ASR to less than two percent. That's a huge drop compared to baseline methods.
Lalam: They also address stability directly by using GDPO to standardize each reward channel before computing advantages, which mitigates that issue people have seen with vanilla GRPO setups where the advantage scale shifts constantly.
Conclusion: Tom: So, to wrap up, the paper on Learning to Attack and Defend: Adaptive Red Teaming of Language Models via GRPO shows a method called AdvGRPO that uses multi-channel rewards and decoupled normalization to make GRPO work for joint optimization.
Jane: The main implication is that this framework creates attackers capable of generalizing well and defenders that are significantly safer, which is a major step forward in adaptive red teaming.
Lu: The authors show they can handle the complexity of multi-turn attacks while maintaining stability through their structured curriculum and the specific reward components they use.
Meng: Practically speaking, this means we have a more reliable way to train both sides of this adversarial game than we had before, which is crucial for building safety guardrails.
Lalam: It’s a practical alternative to PPO and DPO-based approaches because it seems stable enough for real-world application, though they do flag that only vanilla benign prompts were used during co-training, which limits the benign compliance of the defenders a bit.
Tom: That’s right, so the paper on Learning to Attack and Defend: Adaptive Red Teaming of Language Models via GRPO offers a stable and effective way to train adaptive attackers and robust defenders using this structured approach.
Microsoft AI Red Team 2 · Microsoft Azure
cs.CL, cs.AI, cs.LG
Submitted: 2026-06-08
Updated: 2026-10-07
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 92/100
The gist: The gist The AdvGRPO framework introduces a co-training method that makes Group Relative Policy Optimization (GRPO) viable for joint attacker-defender optimization using dense multi-channel rewards
Key concepts
- AdvGRPO
- A co-training framework for joint attacker-defender optimization using GRPO. It trains the attacker and defender models by alternating updates, progressing from single-turn to closed-loop multi-turn attacks.
- Dense Multi-channel Rewards
- Multiple LLM judges score different aspects of the interaction, including attack success, prompt faithfulness, thinking quality, and helpfulness. These scores are combined using GDPO for normalization before calculating policy updates.
- GDPO Advantage Normalization
- Group Discriminator Policy Optimization (GDPO) is used to independently normalize each reward channel before combining them. This standardizes the advantage scale, which prevents the advantage computation from being skewed by shifting reward distributions during co-training.
Terminology
Summary
The gist The AdvGRPO framework introduces a co-training method that makes Group Relative Policy Optimization (GRPO) viable for joint attacker-defender optimization using dense multi-channel rewards and decoupled advantage normalization, showing that it produces highly effective and transferable attacks and that co-trained defenders outperform baselines on safety benchmarks.
How it works
AdvGRPO is a co-training framework designed for joint attacker-defender optimization using GRPO. It progresses through a curriculum from single-turn to closed-loop multi-turn attacks before bootstrapping co-training, where attacker and defender models are updated in alternation. The method utilizes dense, multi-channel rewards and decoupled advantage normalization to make GRPO viable for this setting.
Policy Optimization
The training involves alternating phases between the attacker and defender updates every N steps during co-training. For single-turn training, the attacker produces a single prompt p ∼ πA(· s, o) given a strategy system prompt s and objective o, and the defender responds once: y ∼ πD(· p). In a K-turn attack, the attacker observes and adapts to the defender’s response at each turn. Each turn is included as a separate training example with its own attack reward A(rk, o), measuring the extent to which the response rk satisfies the attack objective.
Reward Functions and GDPO Advantage
Multiple dense reward channels are scored on a [0, 1] scale by an LLM judge. These channels include:
-
Attack reward A, which evaluates the extent to which the defender’s response y satisfies the adversarial objective o along intent alignment, content harms, and detail level.
-
Attack prompt reward P, which closes this credit-assignment gap by evaluating objective faithfulness, strategy compliance (conditioned on the attack strategy system prompt s), and coherence.
-
Thinking-trace reward T, which penalizes both failure modes by evaluating conciseness, attacker commitment, and objective faithfulness.
-
Helpfulness reward H, which scores intent alignment and detail for benign objectives during defender co-training.
GDPO is leveraged to normalize these reward channels independently before combining them via a weighted sum and re-normalizing across the batch. For the attacker, the channels are A, P, and optionally T, while for the defender in co-training, the channels are 1−A for adversarial objectives and H for benign objectives.
Training Modes
The framework supports two primary training modes. Attacker-only training freezes the defender and trains only πA, serving as a curriculum learning stage. In single-turn training (K = 1), the base model learns to overcome alignment-induced refusal and generate effective attack prompts. The attacker’s objective combines all active reward channels via GDPO with adversarial objectives sampled from Dadv.
For co-training, models are updated in alternation every N steps. During the attacker phase, πD is frozen and πA is updated with the attacker objective. During the defender phase, πA is frozen, and each batch mixes a fraction α of adversarial objectives with 1 − α benign objectives from a separate dataset Dbenign. The defender’s objective involves maximizing both adversarial rewards (1−A) and benign rewards (H).
Key Findings
AdvGRPO produces strong singleturn, multi-turn, and reasoning attackers that generalize well to unseen defenders and out-ofdistribution (OOD) objectives. The method can produce highly effective attacks that transfer to unseen model families. Furthermore, the framework demonstrates that GRPO can be effective for co-training despite prior reports of instability by using GDPO to standardize each reward channel before advantage computation. The curriculum pre-training of the attacker prevents the defender from dominating in co-training by seeding it with a capable attacker.
The results show that AdvGRPO attackers achieve substantial gains across all model sizes, with the best configuration achieving 90–91% ASR on both benchmarks. Defender evaluations indicate that AdvGRPO defenders achieve the lowest ASR (highest safety) on all benchmarks, reducing HarmBench ASR to <2% compared to baseline methods. The attacker and defender in this setup are coupled via A with opposing signs, analogous to the generator-discriminator dynamic in GANs.
The limitations noted include reduced benign compliance for co-trained defenders because only vanilla benign prompts were used during cotraining, and some entropy collapse in attacker prompts over training which reduced attack diversity. The authors suggest developing entropy-aware exploration compatible with their approach as a useful direction for future work.
The ethical considerations emphasize that the goal of this research is to support AI red teaming by systematically identifying model weaknesses so that they can be mitigated before real-world harm occurs. All experiments were conducted in controlled research settings, and no harmful content was published in this work. The datasets used are publicly available and commonly used in AI safety research. The successful attack examples were selectively redacted to avoid disseminating unnecessarily harmful content while still illustrating model behaviors. All core research ideas, design decisions, experiments, analyses, and conclusions were conceived and verified by the authors. The paper concludes that GRPO-based co-training can be both stable and effective, offering a practical alternative to PPO and DPO-based approaches. All experiments were conducted on a single node with 4× NVIDIA A100 80GB GPUs using bfloat16 precision and gradient checkpointing. The training parameters include setting the clipping parameter ε=0.2, KL coefficient β=0.1 for co-training, and performing E=2 inner gradient steps per batch of rollouts. The attacker reward is a weighted combination of the attack reward A (weight 1.0), prompt reward P (weight 0.5), and, for thinking models, the thinking-trace reward T (weight 0.5). The attacker uses temperature 1.0, top-p 1.0, and a maximum of 512 new tokens for rollout generation. The defender generates up to 500 tokens with temperature 1.0. The paper provides detailed reward function scoring rubrics where all rewards are computed by an LLM judge (GPT-4.1) using a multiplicative structure. All experiments were conducted on a single node with 4× NVIDIA A100 80GB GPUs using bfloat16 precision and gradient checkpointing. The training runs for 300 steps during co-training. The paper shows that the defender learns to generate safer responses after 7–8 alternations in the curriculum-based co-training approach. All experiments were conducted on a single node with 4× NVIDIA A100 80GB GPUs using bfloat16 precision and gradient checkpointing. The paper reports that general utility scores are unaffected and even improve on IFBench, indicating that co-training does not degrade factual knowledge, reasoning, or instruction following abilities. The paper notes that initializing both models from scratch caused the defender to dominate because deflecting weak attacks is easier than discovering novel attack strategies. All experiments were conducted in controlled research settings. The authors acknowledge the dual-use nature of such capabilities but state the goal is to support AI red teaming by systematically identifying model weaknesses so that they can be mitigated before real-world harm occurs. The paper concludes that GRPO-based co-training can be both stable and effective, offering a practical alternative to PPO and DPO-based approaches. All experiments were conducted on a single node with 4× NVIDIA A100 80GB GPUs using bfloat16 precision and gradient checkpointing. The paper shows that AdvGRPO can simultaneously optimize multiple reward channels. The paper notes that the group normalization in vanilla GRPO couples the advantage scale to a continually shifting reward distribution as the attacker and defender both change, which is mitigated by using GDPO to standardize each reward channel before advantage computation. The paper shows that AdvGRPO can produce strong attackers in single-turn, reasoning, and closed-loop multi-turn settings. All experiments were conducted on a single node with 4× NVIDIA A100 80GB GPUs using bfloat16 precision and gradient checkpointing. The paper demonstrates that GRPO can be effective for cotraining robust defenders despite prior reports of instability. The paper presents AdvGRPO, a framework for training adaptive language model attackers and robust defenders via GRPO. All experiments were conducted in controlled research settings. The authors used GPT-4.1 as the target model in their training runs. The paper reports that transfer ASR is especially well for multi-turn attacks, with the Qwen2.
Improvements for AI systems
-
Improved Attacker Capability: The resulting attacker can produce
strong singleturn, multi-turn, and reasoning attackers that generalize well to unseen defenders and out-ofdistribution (OOD) objectives,
as shown by achieving90–91% ASR on both benchmarks
in the best configuration. -
Improved Co-Trained Defender Robustness: The co-trained defender exhibits superior safety performance, as the paper shows that "AdvGRPO defenders achieve the lowest ASR (highest safety) on all benchmarks, reducing HarmBench ASR to <2% compared to 18.8% for the base model."
-
Improved Strategy Transferability: The framework enables
transfer ASR against held-out target models not seen during training,
indicating that attackers cangeneralize well to unseen defenders and out-ofdistribution (OOD) objectives.
-
Improved Reasoning Attack Efficacy: For reasoning-capable models, the system generates attackers that can overcome self-censoring by using a
thinking-trace reward T,
allowing them to produce prompts that successfully jailbreak models like GPT-4.1, as demonstrated by the Qwen3.5-9B AdvGRPO model achieving79.1 (+79.1) 71.0 (+70.5)
ASR on HarmBench in a single turn with reasoning traces (ST-Think). -
Improved Co-Training Stability: The framework addresses prior instability reports by using
Group reward-Decoupled Policy Optimization (GDPO) to normalize these reward channels independently before combining them,
which helpsmitigate reward signal collapse.
Abstract
Language model safety must continually adapt to evolving attacks. Recent works have demonstrated that reinforcement learning can be used to train stronger attacker and defender models in tandem by applying PPO-style self-play and DPO-style online preference optimization. In this work, we explore the efficacy of GRPO in this setting. Co-training can be challenging because it requires jointly optimizing multiple properties of both the attacker and defender. We therefore shape model outputs using multiple LLM judge-based reward channels and compute advantages with GDPO, which prevents any single channel from dominating. Our method uses a curriculum that progresses from attacker-only single-turn and multi-turn training to co-training, where attacker and defender models are updated in alternation. We show that this method produces highly effective and transferable attacks, and that co-trained defenders reach competitive safety while preserving general utility. Through a controlled ablation, we further identify which components of our training pipeline most affect the resulting balance between safety and utility. Finally, we find that GRPO tends to collapse attacker diversity over training and discuss possible ways to address this limitation.
Sources
- Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning
- Learning to Inject: Automated Prompt Injection via Reinforcement Learning
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
- Jailbreak-R1: Exploring the Jailbreak Capabilities of LLMs via Reinforcement Learning
- WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
- Measuring Massive Multitask Language Understanding
- LoRA: Low-Rank Adaptation of Large Language Models
- WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
- TruthfulQA: Measuring How Models Mimic Human Falsehoods
- Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models
- GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- PyRIT: A Framework for Security Risk Identification and Red Teaming in Generative AI System
- The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
- Safety Alignment of LMs via Non-cooperative Games
- Automated Red Teaming with GOAT: the Generative Offensive Agent Tester
- Red Teaming Language Models with Language Models
- Generalizing Verifiable Instruction Following
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering