Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning
cs.LG, cs.AI
Submitted: 2026-06-02
Updated: 2026-09-08
Comments: Core contributors: Anthony GX-Chen, Ankit Anand, Gheorghe Comanici, André Barreto, Mark Rowland
License: http://creativecommons.org/licenses/by/4.0/
The gist: Classical reinforcement learning (RL) typically seeks a deterministic policy that maximizes the expected sum of a scalar reward.
Terminology
Abstract
Classical reinforcement learning (RL) typically seeks a deterministic policy that maximizes the expected sum of a scalar reward. Yet, modern applications such as language model fine-tuning or scientific discovery demand diversity. Existing remedies such as entropy regularization or diversity bonuses often require fragile trade-offs that sacrifice performance for stochasticity or rely on heuristic metrics that can misalign policy rankings. We argue that diversity is more naturally understood as the rational response to uncertainty in the reward. When the reward function is not perfectly known--as is the case with ambiguous preferences or imperfect reward models--committing to a single action can be sub-optimal. Building on this, we propose a fundamental reformulation of the RL objective by replacing the scalar reward with a distribution over reward functions, and applying a non-linear objective over sets of actions. The result is a framework in which calibrated behavioural diversity emerges naturally, remains controllable through the reward function distribution, and is obtained without sacrificing expected reward. Focusing on the contextual bandit setting as commonly used in large language model (LLM) post-training, we derive a principled gradient estimator for this objective and prove that our formulation naturally generalizes both vanilla policy gradient and more recently developed action-set approaches. We provide didactic experiments which complement our theoretical results, and our large-scale empirical results in LLM reasoning further demonstrate that this framework offers a robust and theoretically grounded alternative for complex RL tasks where the traditional formulation of the problem fails to induce the desired breadth of agent behaviour.
Sources
- An AI system to help scientists write expert-level empirical software
- Vector Policy Optimization: Training for Diversity Improves Test-Time Search
- The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
- Breaking the Bias Barrier in Concave Multi-Objective Reinforcement Learning
- Scaling Laws for Reward Model Overoptimization
- Understanding Performance Gap Between Parallel and Sequential Sampling in Large Reasoning Models
- Jointly Reinforcing Diversity and Quality in Language Model Generations
- Poly-EPO: Training Exploratory Reasoning Models
- Deep Reinforcement Learning with Attention for Slate Markov Decision Processes with High-Dimensional States and Actions
- LLM Probability Concentration: How Alignment Shrinks the Generative Horizon
- The Price of Format: Diversity Collapse in LLMs
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks