Improving Offline Goal-Conditioned Reinforcement Learning via Selective Reward Stimulation
cs.LG
Submitted: 2026-09-16
Updated: 2026-09-16
License: http://creativecommons.org/licenses/by/4.0/
The gist: Goal-conditioned reinforcement learning aims to learn policies that reach specified goals, but remains challenging in offline settings with sparse rewards and long-horizon dependencies.
Terminology
Abstract
Goal-conditioned reinforcement learning aims to learn policies that reach specified goals, but remains challenging in offline settings with sparse rewards and long-horizon dependencies. In such settings, goal-completion information can be temporally distant from the early decisions that enable success, while offline value estimation introduces additional error. We study this issue from a reward-propagation perspective and show, in a stylized delayed-goal setting, how goal-directed value separation can become small relative to local estimation error. Motivated by this analysis, we propose Reward Stimulation Implicit Q-Learning (RSIQL), a simple non-hierarchical method that introduces additional reward signals at progress-making intermediate states in offline trajectories. RSIQL uses an auxiliary goal-conditioned value function to identify intermediate states estimated to make progress toward the goal and applies reward stimulation to provide less-delayed training supervision. Unlike hierarchical methods, RSIQL does not learn a separate high-level subgoal policy. Experiments on D4RL goal-reaching benchmarks and OGBench show that RSIQL improves over goal-conditioned IQL on average and achieves performance competitive with hierarchical offline goal-conditioned methods, while retaining a simple flat policy structure.
Sources
- OPAL: Offline Primitive Discovery for Accelerating Offline Reinforcement Learning
- Stochastic Variational Video Prediction
- Pessimistic Bootstrapping for Uncertainty-Driven Offline Reinforcement Learning
- RvS: What is Essential for Offline RL via Supervised Learning?
- D4RL: Datasets for Deep Data-Driven Reinforcement Learning
- Extreme Q-Learning: MaxEnt RL without Entropy
- Learning to Reach Goals via Iterated Supervised Learning
- Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning
- IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies
- Dynamical Distance Learning for Semi-Supervised and Unsupervised Skill Discovery
- Planning with Diffusion for Flexible Behavior Synthesis
- Efficient Planning in a Compact Latent Action Space
- Model-Based Reinforcement Learning for Atari
- Offline Reinforcement Learning with Implicit Q-Learning
- Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review
- Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
- Learning Multi-Level Hierarchies with Hindsight
- Hierarchical Reinforcement Learning with Hindsight
- Hierarchical Foresight: Self-Supervised Learning of Long-Horizon Tasks via Visual Subgoal Generation
- OGBench: Benchmarking Offline Goal-Conditioned RL
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks