Shared Actors Need Not Share Critics: Effects of Value Mismatch in Parallel Reinforcement Learning
cs.LG
Submitted: 2026-08-27
Updated: 2026-08-27
License: http://creativecommons.org/licenses/by/4.0/
The gist: When a single policy is trained in parallel across multiple environments of the same task, such as procedurally generated levels, randomized dynamics, or curricula, implementations commonly use one
Terminology
Abstract
When a single policy is trained in parallel across multiple environments of the same task, such as procedurally generated levels, randomized dynamics, or curricula, implementations commonly use one critic across all sampled environments. Yet different environments can assign different expected returns to the same input visible to the critic. A critic without environment information must then reconcile distinct value targets, systematically shifting the sampled advantages within individual environments. Using illustrative bandit models with multiple environments and a common optimal arm, we characterize how this value mismatch redistributes sampled policy updates, reinforcing unhelpful actions while attenuating or even reversing useful ones. The oracle processes using no baseline, the shared value, or the value specific to the sampled environment have the same mean logit update at a fixed policy and converge to the same optimal policy, yet their realized learning paths can differ sharply. The analysis motivates a minimal intervention: give only a logged environment index to the critic so that it can separate the value targets. Controlled CartPole and MuJoCo experiments expose the predicted shifted values, advantages, and performance gaps. In the more complex BipedalWalker and Procgen settings, the same intervention yields more stable learning and higher returns. Across all 16 Procgen games, the multihead conditional critic improves aggregate normalized return on 600 unseen levels per game by 40.8%. In conclusion, the theory identifies value mismatch as a direct mechanism through which critic sharing can degrade stochastic learning dynamics, not captured by scalar estimator variance alone, and the experiments show that conditioning on an index is broadly effective in parallel reinforcement learning.
Sources
- OpenAI Gym
- Q-Prop: Sample-Efficient Policy Gradient with An Off-Policy Critic
- Action-depedent Control Variates for Policy Optimization via Stein's Identity
- Sim-to-Real Transfer of Robotic Control with Dynamics Randomization
- FiLM: Visual Reasoning with a General Conditioning Layer
- Decoupling Value and Policy for Generalization in Reinforcement Learning
- Proximal Policy Optimization Algorithms
- The Mirage of Action-Dependent Baselines in Reinforcement Learning
- Learning values across many orders of magnitude
- The Optimal Reward Baseline for Gradient-Based Reinforcement Learning
- Variance Reduction for Policy Gradient with Action-Dependent Factorized Baselines
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks