Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying
cs.LG, cs.AI
Submitted: 2026-05-29
Updated: 2026-09-12
Comments: ICML 2026 camera-ready version; Github: https://github.com/nissymori/remax-rl
Code: https://github.com/nissymori/remax-rl
License: http://creativecommons.org/licenses/by/4.0/
The gist: In reinforcement learning (RL), agents benefit from exploration only because they repeatedly encounter similar states: trying different actions can improve performance or reduce uncertainty; without
Terminology
Abstract
In reinforcement learning (RL), agents benefit from exploration only because they repeatedly encounter similar states: trying different actions can improve performance or reduce uncertainty; without such retries, a greedy policy is optimal. We formalize this intuition with ReMax, an objective that evaluates a policy by the expected maximum return over M samples, where M is a positive integer, while accounting for return uncertainty. Optimizing this objective induces stochastic exploration as an emergent property, without explicit bonus terms. For efficient policy optimization, we derive a new policy-gradient formulation for ReMax and introduce ReMax PPO (RePPO), a PPO variant that optimizes ReMax while generalizing the discrete retry count M to a continuous parameter m > 0, enabling fine-grained control of exploration. Empirically, RePPO promotes exploration, without any explicit exploration bonuses, on the MinAtar and Craftax benchmarks. The code is available at https://github.com/nissymori/remax-rl.
Sources
- Evaluating Large Language Models Trained on Code
- Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
- Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review
- Playing Atari with Deep Reinforcement Learning
- Proximal Policy Optimization Algorithms
- Incentivizing Exploration In Reinforcement Learning With Deep Predictive Models
- MinAtar: An Atari-Inspired Testbed for Thorough and Reproducible Reinforcement Learning Experiments
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks