JumpStart Your Policy Learning with Lessons from 160,000 Training Runs
cs.LG
Submitted: 2026-09-12
Updated: 2026-09-25
Code: https://github.com/huggingface/lerobot
License: http://creativecommons.org/licenses/by/4.0/
The gist: Reliable progress in offline policy learning depends on careful reporting, well-tuned baselines, and evaluation across diverse conditions.
Terminology
Abstract
Reliable progress in offline policy learning depends on careful reporting, well-tuned baselines, and evaluation across diverse conditions. Prior work has shown that results can be sensitive to reporting choices, hyperparameter tuning, and dataset properties, but these sources of variability have not been systematically investigated together at the scale needed to understand how they shape conclusions. To address this gap, we present a large-scale empirical study of offline reinforcement and imitation learning, training over 160,000 policies across 114 datasets. At this scale, no algorithm dominates: aggregate performance among the strongest methods is often close, but the leaders differ substantially across environments. We find that proper hyperparameter tuning frequently reshuffles perceived algorithm rankings and that benchmark composition can produce conflicting conclusions. We also study hyperparameter sensitivity and transfer across environments, identifying a simple strategy for deriving strong default configurations. We use our findings to develop a dataset-conditioned recommender that provides task-specific algorithm recommendations for practitioners. Finally, we release JumpStart: a resource suite containing every trained policy, per-model scores and hyperparameters, strong baselines across all environments, training and evaluation code, and an extensible website for retrieving, analyzing, and contributing results. Together, these resources aim to make offline policy-learning research more reliable and enable future work beyond the scope of this study.
Sources
- Improving and Benchmarking Offline Reinforcement Learning Algorithms
- Simple Ingredients for Offline Reinforcement Learning
- A Dataset Perspective on Offline Reinforcement Learning
- When should we prefer Decision Transformers for Offline Reinforcement Learning?
- Deep Reinforcement Learning at the Edge of the Statistical Precipice
- Offline Reinforcement Learning with Implicit Q-Learning
- Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware
- Diffusion Policy: Visuomotor Policy Learning via Action Diffusion
- Revisiting the Minimalist Approach to Offline Reinforcement Learning
- Off-Policy Deep Reinforcement Learning without Exploration
- OpenAI Gym
- Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations
- Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning
- Flow: A Modular Learning Framework for Mixed Autonomy Traffic
- Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction
- Behavior Regularized Offline Reinforcement Learning
- robosuite: A Modular Simulation Framework and Benchmark for Robot Learning
- LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning
- D5RL: Diverse Datasets for Data-Driven Deep Reinforcement Learning
- OGBench: Benchmarking Offline Goal-Conditioned RL
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks