Constrained Group Relative Policy Optimization
cs.LG, cs.CL, cs.RO
Submitted: 2026-02-05
Updated: 2026-09-02
License: http://creativecommons.org/licenses/by/4.0/
The gist: Group Relative Policy Optimization (GRPO) remains the dominant critic-free approach for fine-tuning LLMs and VLMs, but its compatibility with constrained policy optimization (e.g.
Terminology
Abstract
Group Relative Policy Optimization (GRPO) remains the dominant critic-free approach for fine-tuning LLMs and VLMs, but its compatibility with constrained policy optimization (e.g. for safety-critical domains) has not been carefully examined. In this work, we introduce Constrained GRPO, a Lagrangian-based extension of GRPO for constrained policy optimization. We show that the standard practice of scalarizing rewards before normalization introduces a critical Lagrangian-specific failure mode: GRPO's within-group normalization makes constrained optimization highly sensitive to how multi-component learning signals are aggregated. We show that scalarizing rewards before normalization introduces shared-denominator coupling, so that changing one multiplier alters not only the emphasis on its corresponding constraint, but also the relative weighting of the reward and other constraints. We address this with a simple but crucial modification: scalarizing standardized advantages rather than rewards. This yields a better-conditioned update by addressing the coupling induced by reward scalarization, resulting in better-behaved multiplier dynamics and more stable constraint enforcement in practice. Empirically, across a controlled gridworld, a real-world autonomous driving benchmark, and a mathematical reasoning task, Constrained GRPO consistently achieves better adherence to specified constraints while maintaining or improving task performance.
Sources
- RIFT: Group-Relative RL Fine-Tuning for Realistic and Controllable Traffic Simulation
- Deep reinforcement learning from human preferences
- SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
- Training Verifiers to Solve Math Word Problems
- Qwen2.5-VL Technical Report
- Safe RLHF: Safe Reinforcement Learning from Human Feedback
- Constitutional AI: Harmlessness from AI Feedback
- Using Natural Language for Reward Shaping in Reinforcement Learning
- Soft Actor-Critic Algorithms and Applications
- Inverse Reward Design
- GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
- MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems
- CaRL: Learning Scalable Planning Policies with Simple Rewards
- Active Inverse Reward Design
- Safe RLHF-V: Safe Reinforcement Learning from Multi-modal Human Feedback
- Solving Quantitative Reasoning Problems with Language Models
- MTGS: Multi-Traversal Gaussian Splatting
- Optimizing Safe and Aligned Language Generation: A Multi-Objective GRPO Approach
- Drive-R1: Bridging Reasoning and Planning in VLMs for Autonomous Driving with Reinforcement Learning
- Confronting Reward Model Overoptimization with Constrained RLHF
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks