A Balanced Data Diet: Addressing the Exploration Bottleneck in Mega-Scale RL for Robot Control
cs.RO, cs.LG
Submitted: 2026-10-08
Updated: 2026-10-08
Project page: https://sgs-rl.github.io
Terminology
Sources
- Orbit: A Unified Simulation Framework for Interactive Robot Learning Environments
- Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning
- DextrAH-RGB: Visuomotor Policies to Grasp Anything with Dexterous Hands
- Solving Rubik's Cube with a Robot Hand
- Emergent Dexterity via Diverse Resets and Large-Scale Reinforcement Learning
- VIRAL: Visual Sim-to-Real at Scale for Humanoid Loco-Manipulation
- Parkour in the Wild: Learning a General and Extensible Agile Locomotion Policy Using Multi-expert Distillation and RL Fine-tuning
- Reverse Forward Curriculum Learning for Extreme Sample and Demonstration Efficiency in Reinforcement Learning
- Combined Constrained Sampling and Reinforcement Learning for Robotic Manipulation
- RMA: Rapid Motor Adaptation for Legged Robots
- BeyondMimic: From Motion Tracking to Versatile Humanoid Control via Guided Diffusion
- BFM-Zero: A Promptable Behavioral Foundation Model for Humanoid Control Using Unsupervised Reinforcement Learning
- Proximal Policy Optimization Algorithms
- Learning Montezuma's Revenge from a Single Demonstration
- ManiSkill3: GPU Parallelized Robotics Simulation and Rendering for Generalizable Embodied AI
- DexPBT: Scaling up Dexterous Manipulation for Hand-Arm Systems with Population Based Training
- Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments
- Intrinsic Motivation and Automatic Curricula via Asymmetric Self-Play
- Towards bridging the gap: Systematic sim-to-real transfer for diverse legged robots
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving