ReBound: Reset-Free Reinforcement Learning for Agile Driving via Reset-Aware Semi-Markov Bootstrapping

arXiv:2604.07672 · cs.RO · Submitted 2026-04-09 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.

Dev: Today's paper: "ReBound: Reset-Free Reinforcement Learning for Agile Driving via Reset-Aware Semi-Markov Bootstrapping".

Rosa: Reset-free reinforcement learning for real-world agile driving addresses the practical barrier of frequent manual resets by enabling continuous, autonomous training on physical platforms.

Dev: First, who's behind it and why it matters.

Title and authors: Rosa: So we're looking at the paper 'ReBound: Reset-Free Reinforcement Learning for Agile Driving via Reset-Aware Semi-Markov Bootstrapping', and I want to start by asking what this title really suggests about what they've accomplished.

Dev: It points toward a major hurdle they tackled, which is the need for continuous learning on physical hardware without those annoying manual resets we have to do in simulations.

Taro: That reset-free aspect seems key because real-world driving involves things like unexpected events that would normally force us to stop and restart the entire training run.

Rosa: Exactly, and the way they frame it with semi-Markov bootstrapping suggests a more nuanced approach than just throwing a single policy at the problem.

Dev: It means they aren't just learning one thing; they're managing transitions between different modes of operation to keep things going even when things go wrong.

Taro: That continuous learning capability is what I'm most interested in because it opens up possibilities for training agents on physical systems that are truly complex.

The paper's summary: Rosa: Moving onto the core summary of 'ReBound: Reset-Free Reinforcement Learning for Agile Driving via Reset-Aware Semi-Markov Bootstrapping', they explain how this system manages the continuous training process after a failure by switching between a forward policy and a reset policy.

Dev: That switch is critical; it allows the agent to recover from something like a collision and immediately resume learning without human intervention.

Taro: I read that the base policy, which they use as both the reset mechanism and for residual learning, is Model Predictive Path Integral control or MPPI, which seems like a smart way to handle those complex vehicle dynamics we talked about earlier.

Rosa: Right, and they set up the task as a Markov Decision Process where that fixed base policy is actually incorporated into the environment side of the MDP so standard RL algorithms can learn the forward policy.

Dev: That MDP formulation is clever because it lets them use established RL techniques while still respecting the underlying system dynamics encoded in that base control mechanism.

Taro: The paper also breaks down how they compare different forward policies like PPO, SAC, and TD-MPC2, showing distinct behaviors during training phases after just fifteen minutes or thirty minutes of learning.

The paper's improvements: Rosa: When we look at the specific improvements they propose in 'ReBound: Reset-Free Reinforcement Learning for Agile Driving via Reset-Aware Semi-Markov Bootstrapping', they focus on using MPPI as both the reset policy and the base policy for residual learning.

Dev: That dual role for MPPI is significant because it gives them a reliable way to get the vehicle back into a safe state after an error, while also providing a stable foundation for new learning steps.

Taro: I see that they're testing this setup against three different forward policy algorithms—PPO, SAC, and TD-MPC2—both with and without residual learning to see what works best.

Rosa: The paper highlights a clear finding: SAC with residual learning gets the highest returns in simulation, but only TD-MPC2 consistently beats the MPPI baseline when tested on the physical platform.

Dev: That discrepancy between simulation and reality is something we need to focus on because it shows that simply following simulation results isn't enough for real-world deployment.

Taro: I think one of the main improvements they suggest is explicitly designing mechanisms to handle unmodeled dynamics like tire slip and actuation delays during policy learning, rather than just relying on a simple sim-to-real transfer or fixed residual learning.

Conclusion: Rosa: So wrapping up the discussion on 'ReBound: Reset-Free Reinforcement Learning for Agile Driving via Reset-Aware Semi-Markov Bootstrapping', the authors conclude that while simulation rankings can be misleading, TD-MPC2 is the only one that consistently outperforms their MPPI baseline on the actual physical platform.

Dev: They emphasize that SAC's tendency to converge to overly conservative behavior in reality shows why relying solely on simulation isn't sufficient for these kinds of control problems.

Taro: I think the real implication here is that we need algorithms specifically tailored for continuous learning in real-world settings, instead of just chasing the highest simulated returns.

Rosa: It really highlights the unique challenges posed by things like real-world noise and observation delays that simulation can't fully capture, which is what makes this whole paper so important.

Dev: That suggests future work should focus on understanding why TD-MPC2 maintains its robustness and perhaps scaling these concepts to handle higher speeds or different track geometries.

Taro: I agree; exploring the roles of latent-space planning and learned dynamics in TD-MPC2 seems like a promising direction for making these agents more capable in complex physical environments.

Department of Mechanical Systems Engineering, Nagoya University · CyberAgent AI Lab

cs.RO

Submitted: 2026-04-09

Updated: 2026-10-06

Comments: 9 pages, 6 figures,

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: Reset-free reinforcement learning for real-world agile driving addresses the practical barrier of frequent manual resets by enabling continuous, autonomous training on physical platforms.

Key concepts

Reset-Free RL
This framework allows an agent to continue learning after a failure (like a crash) without needing a human to manually reset the system. It achieves this by switching between two policies: one for driving forward and another specifically designed to return the vehicle to a safe, restartable state.
Model Predictive Path Integral control (MPPI)
MPPI is used as both the base policy and the reset controller. It predicts future trajectories based on a model of vehicle dynamics and optimizes a path integral, effectively guiding the car's movement in real-time to achieve desired goals.
Residual Learning
This technique involves training an additional policy that learns only small corrections or 'residuals' on top of the main base policy. In this study, residual learning is tested to see if it can improve performance when training on physical hardware.

Terminology

Summary

Reset-free reinforcement learning for real-world agile driving addresses the practical barrier of frequent manual resets by enabling continuous, autonomous training on physical platforms. The gist: Reset-free RL in the real world poses unique challenges absent from simulation, such as real-world noise, observation delays, and model inaccuracies interacting with algorithmic properties in ways that simulation alone cannot reveal.

The Challenge and Motivation

The central difficulty in high-speed autonomous driving lies in the complexity of vehicle dynamics where unmodeled effects like slip, suspension dynamics, actuation delay, and aerodynamic forces significantly influence behavior. These factors hinder both accurate simulation and direct sim-to-real transfer of learned policies. While model-based control methods like Model Predictive Control (MPC) are effective due to their ability to account for system dynamics and constraints, their performance is fundamentally limited by the accuracy of the predictive model, which becomes critical in aggressive driving regimes where true dynamics are highly complex.

The Reset-Free Framework

Reset-free RL addresses the impracticality of manual resets after failures (e.g., collisions or off-track excursions) by alternating between a forward policy and a reset policy. This paradigm allows training to continue autonomously after a failure, provided a predefined or co-trained reset policy returns the robot to a restartable state. In this study, Model Predictive Path Integral control (MPPI) is employed as both the base policy and the reset controller. Furthermore, MPPI serves as the baseline for residual learning.

Algorithm Comparison and Residual Learning

The study systematically compares three representative RL algorithms: Proximal Policy Optimization (PPO), Soft Actor-Critic (SAC), and Temporal Difference MPC 2 (TD-MPC2), both with and without residual learning in simulation and real-world experiments. The formulation of the task as a Markov Decision Process (MDP) incorporates the base policy into the environment side, allowing standard RL algorithms to learn the forward policy.

Key comparisons include:

  1. SAC with residual learning achieves the highest returns in simulation.

  2. Only TD-MPC2 consistently outperforms the MPPI baseline on the physical platform.

  3. Residual learning, while beneficial in simulation, fails to transfer its advantage to the real world and can even degrade performance.

Real-World Performance Discrepancy

The experiments reveal a clear gap between simulation and real-world learning. In simulation, SAC (Residual) shows the most stable reward curve and the highest sample efficiency, while TD-MPC2 reaches high peaks but exhibits greater episode-to-episode variance. In the real world, only TD-MPC2 consistently surpasses the MPPI baseline; SAC converges to overly conservative behavior, while residual learning provides little benefit because it constrains exploration to a neighborhood of the base policy output.

Conclusion and Future Directions

The findings underscore that reset-free RL in the real world is uniquely challenging due to real-world noise, observation delays, and model inaccuracies. The results call for further algorithmic development specifically tailored to real-world continuous learning rather than relying on simulation-based rankings. Future work suggests investigating why TD-MPC2 maintains robust performance—specifically the roles of latent-space planning and learned dynamics—and exploring scaling to higher speeds and transfer learning across different track geometries.

The gist: Reset-free RL in the real world poses unique challenges absent from simulation, such as real-world noise, observation delays, and model inaccuracies interacting with algorithmic properties in ways that simulation alone cannot reveal.

How it works

  1. The task is formulated as a Markov Decision Process (MDP) where the base policy is fixed and incorporated into the environment side of the MDP.

  2. The action space consists of commanded vehicle speed and steering angle, which are converted into actual control signals by an onboard motor controller.

  3. The state space combines vehicle kinematics (speed, angular velocity, acceleration), steering angle, LiDAR range measurements, and the base policy output at each time step with a history length of L steps.

  4. The reward function encourages high-speed driving while penalizing collisions based on a cost map constructed from the 2D LiDAR scan.

Base Policy Design

The base policy is implemented using Model Predictive Path Integral control (MPPI), which functions in two roles:

  1. As a reset policy: after detecting a collision caused by the forward policy, it recovers the vehicle to a restartable state.

  2. As the base policy for residual learning: the forward policy can be trained to output residual corrections on top of the base policy output, or directly output full control commands.

Forward Policy Algorithms

Three representative algorithms are tested: PPO (on-policy, model-free), SAC (off-policy, model-free), and TD-MPC2 (off-policy, model-based). TD-MPC2 is particularly notable as it "learn[s] a latent dynamics and reward model and plan[s]

Improvements for AI systems

Here are the specific improvements that can be made to AI systems based on this study, along with what those improved systems could achieve:

  1. Improve real-world deployment robustness of Reinforcement Learning (RL) controllers by integrating a hybrid architecture that combines model-based control with model-free learning. This system would utilize a learned dynamics/reward model (as in TD-MPC2) to provide robust, optimistic planning in the latent space, while employing an on-policy or off-policy RL algorithm (like SAC or PPO) to learn residual corrections.

  2. Enable autonomous, continuous training of high-speed driving policies directly on physical hardware without requiring manual resets. This system would consist of a forward policy that learns to output residual corrections atop a pre-trained, stochastic Model Predictive Path Integral control (MPPI) base policy, allowing the agent to recover from failures autonomously and continue learning.

  3. Develop RL agents capable of achieving high-speed maneuverability on slippery or unknown physical surfaces in real environments. Specifically, the improved system can drive a 1/10-scale vehicle on indoor tracks with significant tire slip (e.g., friction coefficient 0.25) and successfully navigate corners at speeds up to 3.0 m/s, surpassing the performance of purely model-based controllers like MPPI in these challenging conditions.

  4. Create RL systems that generalize better from simulation to the real world by incorporating explicit mechanisms for handling unmodeled dynamics (like tire slip and actuation delays) during policy learning, rather than relying on simple sim-to-real transfer or fixed residual learning. This improved system can maintain high performance in the real world where model mismatch is significant.

  5. Design more sample-efficient RL algorithms for continuous control tasks by leveraging off-policy methods (SAC/TD-MPC2) combined with learned latent dynamics models, which allows the agent to explore more effectively and avoid converging to overly conservative local optima observed in purely model-free approaches like SAC in real-world testing.

Sources

Related papers