Learning to Attack and Defend: Adaptive Red Teaming of Language Models via GRPO

summary

Video file (mp4)

The gist

The gist The AdvGRPO framework introduces a co-training method that makes Group Relative Policy Optimization (GRPO) viable for joint attacker-defender optimization using dense multi-channel rewards

In short

AdvGRPO introduces a co-training method for joint attacker-defender optimization using Group Relative Policy Optimization (GRPO). It uses dense, multi-channel rewards and decoupled advantage normalization to make GRPO work effectively. This results in highly transferable attacks and defenders that significantly outperform baselines on safety benchmarks.

Key concepts

AdvGRPO
A co-training framework for joint attacker-defender optimization using GRPO. It trains the attacker and defender models by alternating updates, progressing from single-turn to closed-loop multi-turn attacks.
Dense Multi-channel Rewards
Multiple LLM judges score different aspects of the interaction, including attack success, prompt faithfulness, thinking quality, and helpfulness. These scores are combined using GDPO for normalization before calculating policy updates.
GDPO Advantage Normalization
Group Discriminator Policy Optimization (GDPO) is used to independently normalize each reward channel before combining them. This standardizes the advantage scale, which prevents the advantage computation from being skewed by shifting reward distributions during co-training.

Terminology used across episodes

This episode discusses

The paper

Learning to Attack and Defend: Adaptive Red Teaming of Language Models via GRPO · Read on arXiv

Microsoft AI Red Team 2 · Microsoft Azure

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Learning to Attack and Defend".

Jane: The gist The AdvGRPO framework introduces a co-training method that makes Group Relative Policy Optimization (GRPO) viable for joint attacker-defender optimization using dense multi-channel rewards and decoupled advantage normalization,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: We're starting with the title of Learning to Attack and Defend: Adaptive Red Teaming of Language Models via GRPO. It immediately tells us this paper is focused on training both an attacker and a defender together, adaptively, using this GRPO method.

Jane: The authors are Blake Bullwinkel, Eugenia Kim, Amanda Minnich, and Mark Russinovich from Microsoft AI Red Team. They bring a lot of experience in the red teaming side of things to this work.

Lu: What they’re doing is tackling the fact that static methods for attacker-defender co-training using GRPO tend to be unstable, and they propose AdvGRPO as a way around that by adding those dense reward channels and the decoupled advantage normalization.

Meng: So they're taking an existing RL technique, GRPO, and making it stable enough to actually train two models in tandem for red teaming. How much stability are we talking about?

Lalam: It suggests that we can achieve better results than previous work because they’ve addressed the instability issues that have been reported with co-training methods like PPO and DPO, which is a key problem in this area.

The paper's summary: Tom: They summarize AdvGRPO as a curriculum that moves from single-turn training to closed-loop multi-turn attacks, after which they bootstrap the co-training by alternating updates between the attacker and defender models every N steps.

Jane: The core mechanism involves using dense, multi-channel rewards—they score responses on intent alignment, content harms, detail level for the attack reward A—and then decoupling advantage normalization for each channel.

Lu: They use four specific reward channels: the attack reward A, which measures how well the response satisfies the objective; a prompt reward P to check strategy adherence; a thinking-trace reward T to penalize failure modes like lack of commitment; and a helpfulness reward H for benign objectives when training the defender.

Meng: It seems they’re trying to score every aspect of the interaction, not just one overall success metric, which is smart because it gives you more feedback on *why* an attack worked or failed.

Lalam: And they leverage GDPO to normalize these channels independently before combining them with a weighted sum, which helps stabilize the training process by standardizing each reward channel before any advantage computation happens.

The paper's improvements: Tom: They highlight several improvements that AdvGRPO brings, specifically noting that it produces strong single-turn, multi-turn, and reasoning attackers that generalize well to unseen defenders and out-ofdistribution objectives.

Jane: A big improvement they point to is the transferability of these attacks; the paper shows they can produce highly effective attacks that transfer to unseen model families.

Lu: For reasoning models, they introduce the thinking-trace reward T, which allows the attacker to overcome self-censoring by evaluating conciseness and commitment during a rollout. They also show that co-trained defenders outperform baselines on safety benchmarks.

Meng: The numbers are pretty compelling here; in their best configuration, the attackers achieve ninety–ninety-one percent ASR on both benchmarks, and the defenders achieve the lowest ASR, reducing HarmBench ASR to less than two percent. That's a huge drop compared to baseline methods.

Lalam: They also address stability directly by using GDPO to standardize each reward channel before computing advantages, which mitigates that issue people have seen with vanilla GRPO setups where the advantage scale shifts constantly.

Conclusion: Tom: So, to wrap up, the paper on Learning to Attack and Defend: Adaptive Red Teaming of Language Models via GRPO shows a method called AdvGRPO that uses multi-channel rewards and decoupled normalization to make GRPO work for joint optimization.

Jane: The main implication is that this framework creates attackers capable of generalizing well and defenders that are significantly safer, which is a major step forward in adaptive red teaming.

Lu: The authors show they can handle the complexity of multi-turn attacks while maintaining stability through their structured curriculum and the specific reward components they use.

Meng: Practically speaking, this means we have a more reliable way to train both sides of this adversarial game than we had before, which is crucial for building safety guardrails.

Lalam: It’s a practical alternative to PPO and DPO-based approaches because it seems stable enough for real-world application, though they do flag that only vanilla benign prompts were used during co-training, which limits the benign compliance of the defenders a bit.

Tom: That’s right, so the paper on Learning to Attack and Defend: Adaptive Red Teaming of Language Models via GRPO offers a stable and effective way to train adaptive attackers and robust defenders using this structured approach.

More episodes

← Home