Learning to Attack and Defend: Adaptive Red Teaming of Language Models via GRPO
summary
The gist
The gist The AdvGRPO framework introduces a co-training method that makes Group Relative Policy Optimization (GRPO) viable for joint attacker-defender optimization using dense multi-channel rewards
In short
AdvGRPO introduces a co-training method for joint attacker-defender optimization using Group Relative Policy Optimization (GRPO). It uses dense, multi-channel rewards and decoupled advantage normalization to make GRPO work effectively. This results in highly transferable attacks and defenders that significantly outperform baselines on safety benchmarks.
Key concepts
- AdvGRPO
- A co-training framework for joint attacker-defender optimization using GRPO. It trains the attacker and defender models by alternating updates, progressing from single-turn to closed-loop multi-turn attacks.
- Dense Multi-channel Rewards
- Multiple LLM judges score different aspects of the interaction, including attack success, prompt faithfulness, thinking quality, and helpfulness. These scores are combined using GDPO for normalization before calculating policy updates.
- GDPO Advantage Normalization
- Group Discriminator Policy Optimization (GDPO) is used to independently normalize each reward channel before combining them. This standardizes the advantage scale, which prevents the advantage computation from being skewed by shifting reward distributions during co-training.
Terminology used across episodes
This episode discusses
- Learning to Attack and Defend: Adaptive Red Teaming of Language Models via GRPO · Paper Radio
- Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning
- Learning to Inject: Automated Prompt Injection via Reinforcement Learning
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
- Jailbreak-R1: Exploring the Jailbreak Capabilities of LLMs via Reinforcement Learning
- WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
- Measuring Massive Multitask Language Understanding
- LoRA: Low-Rank Adaptation of Large Language Models
- WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
- TruthfulQA: Measuring How Models Mimic Human Falsehoods
- Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models
- GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- PyRIT: A Framework for Security Risk Identification and Red Teaming in Generative AI System
- The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
- Safety Alignment of LMs via Non-cooperative Games
- Automated Red Teaming with GOAT: the Generative Offensive Agent Tester
- Red Teaming Language Models with Language Models
- Generalizing Verifiable Instruction Following
The paper
Learning to Attack and Defend: Adaptive Red Teaming of Language Models via GRPO · Read on arXiv
Microsoft AI Red Team 2 · Microsoft Azure
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Learning to Attack and Defend".
Jane: The gist The AdvGRPO framework introduces a co-training method that makes Group Relative Policy Optimization (GRPO) viable for joint attacker-defender optimization using dense multi-channel rewards and decoupled advantage normalization,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: We're starting with the title of Learning to Attack and Defend: Adaptive Red Teaming of Language Models via GRPO. It immediately tells us this paper is focused on training both an attacker and a defender together, adaptively, using this GRPO method.
Jane: The authors are Blake Bullwinkel, Eugenia Kim, Amanda Minnich, and Mark Russinovich from Microsoft AI Red Team. They bring a lot of experience in the red teaming side of things to this work.
Lu: What they’re doing is tackling the fact that static methods for attacker-defender co-training using GRPO tend to be unstable, and they propose AdvGRPO as a way around that by adding those dense reward channels and the decoupled advantage normalization.
Meng: So they're taking an existing RL technique, GRPO, and making it stable enough to actually train two models in tandem for red teaming. How much stability are we talking about?
Lalam: It suggests that we can achieve better results than previous work because they’ve addressed the instability issues that have been reported with co-training methods like PPO and DPO, which is a key problem in this area.
The paper's summary: Tom: They summarize AdvGRPO as a curriculum that moves from single-turn training to closed-loop multi-turn attacks, after which they bootstrap the co-training by alternating updates between the attacker and defender models every N steps.
Jane: The core mechanism involves using dense, multi-channel rewards—they score responses on intent alignment, content harms, detail level for the attack reward A—and then decoupling advantage normalization for each channel.
Lu: They use four specific reward channels: the attack reward A, which measures how well the response satisfies the objective; a prompt reward P to check strategy adherence; a thinking-trace reward T to penalize failure modes like lack of commitment; and a helpfulness reward H for benign objectives when training the defender.
Meng: It seems they’re trying to score every aspect of the interaction, not just one overall success metric, which is smart because it gives you more feedback on *why* an attack worked or failed.
Lalam: And they leverage GDPO to normalize these channels independently before combining them with a weighted sum, which helps stabilize the training process by standardizing each reward channel before any advantage computation happens.
The paper's improvements: Tom: They highlight several improvements that AdvGRPO brings, specifically noting that it produces strong single-turn, multi-turn, and reasoning attackers that generalize well to unseen defenders and out-ofdistribution objectives.
Jane: A big improvement they point to is the transferability of these attacks; the paper shows they can produce highly effective attacks that transfer to unseen model families.
Lu: For reasoning models, they introduce the thinking-trace reward T, which allows the attacker to overcome self-censoring by evaluating conciseness and commitment during a rollout. They also show that co-trained defenders outperform baselines on safety benchmarks.
Meng: The numbers are pretty compelling here; in their best configuration, the attackers achieve ninety–ninety-one percent ASR on both benchmarks, and the defenders achieve the lowest ASR, reducing HarmBench ASR to less than two percent. That's a huge drop compared to baseline methods.
Lalam: They also address stability directly by using GDPO to standardize each reward channel before computing advantages, which mitigates that issue people have seen with vanilla GRPO setups where the advantage scale shifts constantly.
Conclusion: Tom: So, to wrap up, the paper on Learning to Attack and Defend: Adaptive Red Teaming of Language Models via GRPO shows a method called AdvGRPO that uses multi-channel rewards and decoupled normalization to make GRPO work for joint optimization.
Jane: The main implication is that this framework creates attackers capable of generalizing well and defenders that are significantly safer, which is a major step forward in adaptive red teaming.
Lu: The authors show they can handle the complexity of multi-turn attacks while maintaining stability through their structured curriculum and the specific reward components they use.
Meng: Practically speaking, this means we have a more reliable way to train both sides of this adversarial game than we had before, which is crucial for building safety guardrails.
Lalam: It’s a practical alternative to PPO and DPO-based approaches because it seems stable enough for real-world application, though they do flag that only vanilla benign prompts were used during co-training, which limits the benign compliance of the defenders a bit.
Tom: That’s right, so the paper on Learning to Attack and Defend: Adaptive Red Teaming of Language Models via GRPO offers a stable and effective way to train adaptive attackers and robust defenders using this structured approach.
More episodes
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck