RSPO: Regularized Self-Play Alignment of Large Language Models
cs.LG, cs.AI
Submitted: 2025-02-24
Updated: 2026-08-29
Comments: Accepted at ICML 2026
Code: https://github.com/xiaohangt/RSPO
License: http://creativecommons.org/licenses/by/4.0/
The gist: Self-play-based policy optimization has emerged as an effective approach for fine-tuning large language models (LLMs), formulating preference optimization as a two-player game.
Terminology
Abstract
Self-play-based policy optimization has emerged as an effective approach for fine-tuning large language models (LLMs), formulating preference optimization as a two-player game. However, the regularization with respect to the reference policy, which is crucial for mitigating over-optimization, has been insufficiently investigated in self-play alignment. To study the impact of different regularization strategies, we propose Regularized Self-Play Policy Optimization (RSPO), a novel framework that unifies prior methods and enables simple plug-and-play regularizers, meanwhile preserving convergence to Nash equilibrium of the corresponding regularized game. We empirically show that RSPO with appropriate regularizers can substantially improve the length-controlled win rate (LCWR) on AlpacaEval-2 across a range of base models, while also achieving consistently superior performance on Arena-Hard, MT-Bench, ArmoRM, and response diversity. In particular, RSPO improves unregularized self-play baseline (SPPO) on AlpacaEval-2 LCWR from 28.5% to 35.4% with base model Mistral-7B, from 38.77% to 43.66% with LLaMA-8B, and from 50.54% to 51.83% with Gemma-2B. Combining simplicity, convergence guarantees, and significant empirical gains, RSPO offers a strong foundation for exploring regularized self-play in alignment. Code is available at https://github.com/xiaohangt/RSPO
Sources
- Diffusion-GAN: Training GANs with Diffusion
- Deep Reinforcement Learning from Self-Play in Imperfect-Information Games
- Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
- Textual Aesthetics in Large Language Models
- Self-Play Preference Optimization for Language Model Alignment
- Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
- Nash Learning from Human Feedback
- Human Alignment of Large Language Models through Online Preference Optimisation
- A Minimaximalist Approach to Reinforcement Learning from Human Feedback
- Magnetic Preference Optimization: Achieving Last-iterate Convergence for Language Model Alignment
- REBEL: Reinforcement Learning via Regressing Relative Rewards
- Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning
- Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
- A Unified Approach to Reinforcement Learning, Quantal Response Equilibria, and Two-Player Zero-Sum Games
- Mirror Descent Policy Optimization
- From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
- Mistral 7B
- SimPO: Simple Preference Optimization with a Reference-Free Reward
- Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization
- Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks