Characterizing Replay Retention Under Dynamics Shift in Model-Based Reinforcement Learning
cs.RO, cs.LG
Submitted: 2026-09-16
Updated: 2026-09-16
Code: https://github.com/NM512/dreamerv3-torch
License: http://creativecommons.org/licenses/by/4.0/
The gist: Adapting to changes in robot dynamics requires learning from new data without discarding experience that may still be useful.
Terminology
Abstract
Adapting to changes in robot dynamics requires learning from new data without discarding experience that may still be useful. In continual model-based reinforcement learning (RL), replay collected before a dynamics change can slow adaptation, while removing it unnecessarily reduces available training data and can be especially costly if earlier dynamics return. We study when recent transitions are preferable to the full replay history. Two quantities characterize this trade-off: change magnitude and age-staleness area under the curve (AUC), measuring how well transition age separates stale from fresh data. Forgetting stale data helps after large permanent shifts but hurts when dynamics recur and older data becomes useful again. Choosing a replay strategy therefore depends on predicting when older data will help or hurt. We test these effects across two locomotion morphologies, two model-based RL algorithms, and Real-World RL benchmark perturbations. Because ground-truth staleness labels are unavailable on deployed robots, we evaluate whether an estimator built from interaction data can still provide the quantities needed to choose a replay strategy after permanent changes. Our results show that replay retention depends on change magnitude and on how the dynamics evolve.
Sources
- Mastering Diverse Domains through World Models
- A Deeper Look at Experience Replay
- Augmenting Replay in World Models for Continual Reinforcement Learning
- DeepMind Control Suite
- An empirical investigation of the challenges of real-world reinforcement learning
- Bayesian Online Changepoint Detection
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving