ELEMENT: Episodic and Lifelong Exploration via Maximum Entropy
cs.LG, cs.AI
Submitted: 2024-12-05
Updated: 2026-09-22
Code: https://github.com/giorgionicoletti/deep_Q_learning_maze
License: http://creativecommons.org/licenses/by/4.0/
The gist: Reinforcement learning agents depend on reward signals whose density is rarely under the designer's control, and when such signals are absent, an agent must generate its own drive to explore.
Terminology
Abstract
Reinforcement learning agents depend on reward signals whose density is rarely under the designer's control, and when such signals are absent, an agent must generate its own drive to explore. State entropy maximization offers a principled objective for this, but existing methods break down at scale in two ways: the intrinsic reward vanishes once a state has been visited, discouraging revisits to the very gateways that lead onward, and estimating entropy over millions of accumulated observations becomes computationally prohibitive. We address both with Episodic and Lifelong Exploration via Maximum Entropy (ELEMENT), a multiscale intrinsically motivated framework for reward-free exploration that transfers to downstream tasks. ELEMENT couples lifelong entropy maximization with a complementary episodic term acting on a faster timescale. For the episodic term, we derive average episodic state entropy, an intrinsic reward that is the exact minimizer of a tractable upper bound on the reward-decomposition objective; for the lifelong term, we propose a k NN graph-based estimator that keeps entropy tractable without forgetting. ELEMENT consistently outperforms state-of-the-art intrinsic reward baselines on state coverage and unsupervised pre-training. Videos, code, and supplementary material: https://sites.google.com/view/element-rl.
Sources
- A Survey of Exploration Methods in Reinforcement Learning
- k-Means Maximum Entropy Exploration
- Fast Rates for Maximum Entropy Exploration
- Deep Curiosity Search: Intra-Life Exploration Can Improve Performance on Challenging Deep Reinforcement Learning Problems
- Go-Explore: a New Approach for Hard-Exploration Problems
- Never Give Up: Learning Directed Exploration Strategies
- Fast Online k-nn Graph Building
- Efficient Exploration via State Marginal Matching
- Learning Long-Term Reward Redistribution via Randomized Return Decomposition
- Soft Actor-Critic Algorithms and Applications
- Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
- Behavioral Cloning from Observation
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks