Streaming Deep Reinforcement Learning Finally Works
cs.LG, cs.AI
Submitted: 2024-10-18
Updated: 2026-09-21
Code: https://github.com/mohmdelsayed/streaming-drl
License: http://creativecommons.org/licenses/by/4.0/
The gist: Learning from a stream of experience as it arrives, also known as streaming learning, is a core part of natural learning.
Terminology
Abstract
Learning from a stream of experience as it arrives, also known as streaming learning, is a core part of natural learning. However, reliable streaming learning has remained a persistent challenge in modern deep reinforcement learning (RL). Instead, most deep RL algorithms learn from old experience by storing past interactions in a buffer. We show that both classical streaming RL, such as Q-learning and actor-critic, when used with deep neural networks, and batch deep RL, such as PPO, SAC, and DQN, when adapted to the streaming setting, often fail to learn. Across 58 Atari games and 50 continuous-control tasks, we find that these methods, in aggregate, perform close to random policies despite extensive task-specific hyperparameter searches. We call this pattern stream barrier. Here, we introduce Stream-X, a shared recipe for streaming deep RL algorithms that combines signal normalization, representation stabilization, and controlled parameter updates. By applying Stream-X to several base streaming RL algorithms, we provide the first family of deep RL algorithms to overcome the stream barrier. Using one prescribed hyperparameter configuration per algorithm across tasks, Stream-X substantially improves aggregate performance, often on par with batch RL algorithms. Beyond these benchmarks, we demonstrate learning with Stream-X algorithms under nonstationarity and resource constraints. Stream-AC, one of the Stream-X algorithms, repeatedly recovers performance across alternating floor-friction regimes in simulation, outperforming the evaluated PPO and SAC baselines. It also learns a heading tracking task on a robot using proprioceptive and visual features from the on-board camera in a naturally changing laboratory environment. Stream-Q learns a Pong game from pixels directly on an ESP32-S3 microcontroller, a device with limited compute and memory.
Sources
- What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study
- Layer Normalization
- Real-Time Recurrent Learning using Trace Units in Reinforcement Learning
- Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO
- Simplifying Deep Temporal Difference Learning
- Investigating Recurrence and Eligibility Traces in Deep Q-Networks
- Mastering Diverse Domains through World Models
- Maintaining Plasticity in Continual Learning via Regenerative Regularization
- Elephant Neural Networks: Born to Be a Continual Learner
- Learning Continually by Spectral Regularization
- Directions of Curvature as an Explanation for Loss of Plasticity
- Disentangling the Causes of Plasticity Loss in Neural Networks
- Normalization and effective learning rates in reinforcement learning
- Towards model-free RL algorithms that scale well with unstructured data
- Improving Lexical Choice in Neural Machine Translation
- Is the Policy Gradient a Gradient?
- How to Make Deep RL Work in Practice
- Mastering Memory Tasks with World Models
- MetaOptimize: A Framework for Optimizing Step Sizes and Other Meta-parameters
- Proximal Policy Optimization Algorithms
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks