EfficientTDMPC: Improved MPC Objectives for Sample-Efficient Continuous Control
cs.LG, cs.AI, cs.RO
Submitted: 2026-05-15
Updated: 2026-09-28
Code: https://github.com/wertyuilife2/bmpc
License: http://creativecommons.org/licenses/by/4.0/
The gist: We introduce EfficientTDMPC, a sample-efficient model-based reinforcement learning method for continuous control built on the TD-MPC family of algorithms.
Terminology
Abstract
We introduce EfficientTDMPC, a sample-efficient model-based reinforcement learning method for continuous control built on the TD-MPC family of algorithms. Central to this family is a planner that aims to find an action sequence that maximizes the estimated return. The return is estimated using a learned model and value networks, each of which can introduce error. EfficientTDMPC introduces three contributions that improve performance by aiming to reduce this error. First, we introduce an aggregate multi-horizon planning objective that evaluates the value at different rollout depths and averages them. Second, we introduce ensembles for state-action value estimation to value-equivalent/MuZero-style model-based RL methods. Third, we add pessimistic reanalyze, which penalizes uncertain return estimates when creating policy targets. We evaluate EfficientTDMPC on HumanoidBench and the DeepMind Control Suite, to the best of our knowledge, it is the new state of the art on both domains in terms of sample efficiency.
Sources
- Soft Actor-Critic Algorithms and Applications
- TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint
- Masked Generative Priors Improve World Models Sequence Modelling Capabilities
- Value Improved Actor Critic Algorithms
- Twice Sequential Monte Carlo for Tree Search
- Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
- DeepMind Control Suite
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks