ESSA: Evolutionary Strategies for Scalable Alignment
cs.LG
Submitted: 2025-07-06
Updated: 2026-09-09
License: http://creativecommons.org/licenses/by/4.0/
The gist: Online alignment of large language models (LLMs) is dominated by reinforcement learning from human feedback (RLHF) with gradient-based optimizers such as PPO or GRPO.
Terminology
Abstract
Online alignment of large language models (LLMs) is dominated by reinforcement learning from human feedback (RLHF) with gradient-based optimizers such as PPO or GRPO. While effective, these pipelines require backpropagation through long rollouts, gradient synchronization across devices, and careful hyperparameter tuning, all of which become increasingly costly at scale. We present ESSA (Evolutionary Strategies for Scalable Alignment), a gradient-free online alignment stage that follows supervised fine-tuning (SFT) and replaces the gradient loop with inference-only black-box search. ESSA optimizes only the singular values of low-rank adaptation (LoRA) factors after a short SFT warm-start, restricting the search to a compact, task-aligned subspace where evolutionary search is practical even for 72B-parameter models. Because the loop is inference-only, ESSA is compatible with INT4/INT8 weight quantization and reduces inter-GPU communication to a few bytes per iteration. Across instruction following (IFEval), preference-based assistant tuning (HelpSteer2, HH-RLHF), and mathematical reasoning (GSM8K, MATH500), ESSA matches or exceeds LoRA-GRPO in the reported LoRA comparisons; on GSM8K it also outperforms Online DPO and PPO, while remaining competitive with both methods on IFEval. At scale, ESSA reaches a fixed MATH500 accuracy on Qwen2.5-32B/PRM800K up to 7.8x faster than LoRA-GRPO on 128 GPUs.
Sources
- Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
- LoTR: Low Tensor Rank Weight Adaptation
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale
- ORPO: Monolithic Preference Optimization without Reference Model
- Parameter-Efficient Transfer Learning for NLP
- LoRA: Low-Rank Adaptation of Large Language Models
- REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
- LoraHub: Efficient Cross-Task Generalization via Dynamic LoRA Composition
- Derivative-Free Optimization for Low-Rank Adaptation in Large Language Models
- VeRA: Vector-based Random Matrix Adaptation
- Prefix-Tuning: Optimizing Continuous Prompts for Generation
- DoRA: Weight-Decomposed Low-Rank Adaptation
- Guided evolutionary strategies: Augmenting random search with surrogate gradients
- Simple random search provides a competitive approach to reinforcement learning
- SimPO: Simple Preference Optimization with a Reference-Free Reward
- Training language models to follow instructions with human feedback
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Evolution Strategies as a Scalable Alternative to Reinforcement Learning
- Proximal Policy Optimization Algorithms
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks