Temporal Consistency Improves Generalization in Contextual Offline Meta Reinforcement Learning
cs.LG
Submitted: 2026-03-03
Updated: 2026-09-07
Code: https://github.com/NM512/dreamerv3-torch
License: http://creativecommons.org/licenses/by/4.0/
The gist: Offline meta-reinforcement learning seeks to learn a policy that generalizes to new related tasks online.
Terminology
Abstract
Offline meta-reinforcement learning seeks to learn a policy that generalizes to new related tasks online. Context-based methods infer a task representation from transition histories, yet learning an effective task representation without supervision remains challenging. Existing methods relying on contrastive learning learn discriminative task representations, but fail to identify task-specific dynamics, while relying on reconstruction can be insufficient to model long-horizon dependencies, limiting generalization to new tasks. We investigate the impact of temporal consistency in latent space on task representation learning, showing that enforcing multi-step predictions in latent space encourages task representations that are able to capture task-dependent dynamics while preventing representation collapse. We provide theoretical analysis characterizing sources of error in value estimation and show through extensive experiments on MuJoCo, Contextual DeepMind Control, and MetaWorld benchmarks that temporal consistency significantly improves both zero-shot and few-shot generalization.
Sources
- Layer Normalization
- Estimating or Propagating Gradients Through Stochastic Neurons
- Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning
- iQRL -- Implicitly Quantized Representations for Sample-efficient Reinforcement Learning
- Mish: A Self Regularized Non-Monotonic Activation Function
- Simplicial Embeddings Improve Sample Efficiency in Actor-Critic Agents
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks