Beyond the Stability-Exploration Dilemma: Environmental Regularization for LLM Policy Optimization
cs.CL
Submitted: 2026-08-24
Updated: 2026-08-25
Code: https://github.com/AlibabaResearch/ERPO
Terminology
Sources
- Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- SALSA: Soup-based Alignment Learning for Stronger Adaptation in RLHF
- Mitigating Reward Over-Optimization in RLHF via Behavior-Supported Regularization
- On-Policy RL with Optimal Reward Baseline
- OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
- Understanding R1-Zero-Like Training: A Critical Perspective
- Qwen2.5 Technical Report
- Reinforcement and Imitation Learning via Interactive No-Regret Learning
- Proximal Policy Optimization Algorithms
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Solving math word problems with process- and outcome-based feedback
- Language Models Learn to Mislead Humans via RLHF
- Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
- Scalable Reinforcement Post-Training Beyond Static Human Prompts: Evolving Alignment via Asymmetric Self-Play
- DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering