Planning for Change: Reinforcement Learning Combined with Bounded Extremum Seeking for Robotic Control under Distribution Shift

summary

Video file (mp4)

The gist

Reinforcement learning policies often degrade in performance when test conditions differ from their training distribution, especially in contact-rich tasks like pushing and pick-and-place.

In short

A hybrid controller combines deep reinforcement learning (DDPG) for fast initial task entry with bounded extremum seeking (ES) for robust adaptation during manipulation. This approach addresses performance degradation when test conditions differ from training, such as varying friction or time-varying goals. The system switches between RL and ES based on contact to leverage the strengths of both methods.

Key concepts

Deep Deterministic Policy Gradient (DDPG)
DDPG is a deep reinforcement learning method used to train an agent's policy. It learns a mapping from states to actions, aiming for optimal performance in tasks like fetching and pushing objects. It is trained initially on standard conditions to learn fast manipulation behaviors.
Bounded Extremum Seeking (ES)
ES is a control technique used for online adaptation in systems with unknown or noisy dynamics. It seeks the maximum or minimum of a system's performance signal, providing robust feedback even when the underlying physical model changes unexpectedly during operation.
Hybrid Switching Architecture
This architecture uses a contact flag to dynamically switch control between two modes: RL for rapid task entry and ES for online adaptation. This switching ensures the controller maintains fast initial movement while gaining robustness against distribution shift once interaction begins.
Distribution Shift
Distribution shift occurs when the real-world conditions during testing (e.g., different friction or a moving goal) are significantly different from the conditions seen during training. The hybrid controller is designed to maintain performance even when this mismatch is severe.

Terminology used across episodes

This episode discusses

The paper

Planning for Change: Reinforcement Learning Combined with Bounded Extremum Seeking for Robotic Control under Distribution Shift · Read on arXiv

Department of Electrical and Computer Engineering, University of New Mexico · Los Alamos National Laboratory

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "Planning for Change".

Dev: Reinforcement learning policies often degrade in performance when test conditions differ from their training distribution, especially in contact-rich tasks like pushing and pick-and-place.

Rosa: First, who's behind it and why it matters.

Paper summary: Rosa: So, we're talking about this paper titled "Planning for Change: Reinforcement Learning Combined with Bounded Extremum Seeking for Robotic Control under Distribution Shift." Basically, the main idea is that reinforcement learning policies often get shaky when the test conditions don't match what they were trained on, which is a big problem in tasks like pushing or pick-and-place where things move around.

Dev: That sounds like it addresses a real practical issue we see every day on the floor; if the environment shifts even slightly from training, those learned policies can go completely off track. Rosa, what exactly is the paper proposing to fix that degradation when things change?

Taro: The authors are proposing a hybrid controller that uses deep deterministic policy gradient, or DDPG, for fast initial learning under standard conditions and then switches to bounded extremum seeking during deployment when the environment starts acting weird. This lets them use the speed of RL initially and then switch to something more robust when things get messy <ref:2604.01142#pg1>.

Rosa: Exactly, and it claims this combination helps maintain performance even when things shift, like having different friction patches or goals changing over time. It suggests that DDPG handles the initial rapid task entry well, while bounded extremum seeking provides that necessary adaptation capability at inference time <ref:2604.01142#pg0>.

Dev: From an engineering standpoint, I'm interested in the switching mechanism; how do they manage the transition between those two different control strategies without introducing latency or instability? Rosa, what are your thoughts on this hybrid approach?

Rosa: The core of it is a switching architecture governed by a contact flag, which dictates when to rely on the RL policy versus when to use bounded extremum seeking. This structure is designed to capture the complementary strengths of both methods during different phases of manipulation <ref:2604.01142#pg0>.

Taro: It’s interesting because it acknowledges that you need fast behavior for getting started, which RL is good at, but then you also need something robust for when the system deviates from expectations <ref:2604.01142#pg1>. I think this addresses the autonomy side of things by giving the system a mechanism to handle unforeseen circumstances after it has established an initial interaction.

Dev: So, when we look at how it handles those time-varying systems that are noisy and analytically unknown, is bounded extremum seeking actually providing more stability than just letting the RL policy try to follow a drifting target? That’s a critical question for the loop rate we need to maintain.

Rosa: Bounded extremum seeking is specifically used because it offers guaranteed bounds on control efforts and parameter update rates even when dealing with noisy or unknown time-varying systems <ref:2604.01142#pg1>. For pushing tasks, for instance, it drives the object toward a fixed goal by using performance signals as an objective function <ref:2604.01142#pg0>.

Paper summary: Taro: That sounds like it gives the system a predictable way to correct course when things go wrong in that contact-rich phase where distribution shift is most severe. If the RL policy starts getting erratic, ES steps in to keep things within safe limits <ref:2604.01142#pg1>.

Dev: It’s important to know what those bounds actually are; if the control effort gets too high or the adaptation rate spikes too much, we have a failure mode that we need to anticipate before deployment. Rosa, does this approach work reliably outside of controlled lab settings?

Rosa: The paper suggests that it's designed specifically to handle out-of-distribution settings, like spatially varying friction patches or time-varying goals, which are common in real-world applications <ref:2604.01142#pg0>. It’s meant to work when the conditions depart significantly from the training regime.

Taro: That’s where I see its real value for autonomy; it means a robot operating in an unknown factory floor environment could maintain its goal even if the friction changes unexpectedly mid-task <ref:2604.01142#pg0>. It extends the viability of these learned policies beyond perfect simulation.

Dev: If we consider the latency involved in switching between RL and ES, does this hybrid architecture introduce any noticeable delay that could cause a collision or an unstable grasp? I'm thinking about the actual execution time on our hardware <ref:2604.01142#pg1>.

Rosa: The design aims to preserve fast manipulation behavior from the RL component during the "rapid task entry" phase, which happens before contact is established <ref:2604.01142#pg0>. The switching logic is tied to when the end-effector first makes contact with the object, which helps minimize disruption right at that critical moment.

Taro: So, it’s about having a fast learning phase and then a robust adaptation phase that kicks in only after the system has actually engaged with its environment <ref:2604.01142#pg0>. That sequencing seems logical for complex physical tasks where initial setup and subsequent tracking require different kinds of control.

Dev: I see how it splits the responsibility, but what about the training phase itself? The paper mentions DDPG policies are trained on standard Fetch manipulation tasks, so how does that training relate to these real-world distribution shifts? Rosa, can you elaborate on the initial training setup?

Rosa: The DDPG policies are initially trained on standard Fetch manipulation tasks using environments like FetchPush and FetchPickAndPlace <ref:2604.01142#pg0>. They are trained in a goal-conditioned setting where both the initial object pose and the desired goal are randomized at the start of each episode, which helps train a policy that maps state and goal to an action <ref:2604.01142#pg0>.

Taro: That randomization during training seems smart because it encourages a more general policy that isn't overly specialized for one single starting condition, which is helpful when deployment conditions are unpredictable <ref:2604.01142#pg1>.

Paper summary: Dev: And looking at the DDPG setup, they use deep neural networks with two hidden layers of two hundred fifty-six neurons each and specific learning rates for the actor and critic—that tells me they're aiming for a certain level of complexity in the learned model <ref:2604.01142#pg2>. Are these network architectures typically computationally demanding when running on embedded systems?

Rosa: The architecture involves a fully connected multilayer perceptron actor and critic with hyperbolic tangent activation to keep actions bounded <ref:2604.01142#pg2>. While the networks are deep, they are designed to be compatible with the DDPG framework which avoids high variance in stochastic actions by using deterministic policy gradient methods <ref:2604.01142#pg2>.

Taro: The critic approximating the action-value function Q(st,at;θQ) is a key part of making sure that the RL policy learns a good value estimate for its actions, which feeds into the actor's update <ref:2604.01142#pg2>. That feedback loop is what allows it to learn how to behave correctly in those standard conditions.

Dev: It sounds like they’re balancing complexity with stability by using those stabilizing mechanisms like experience replay and slowly moving target networks for both the actor and critic parameters <ref:2604.01142#pg2>. That takes a lot of tuning on our side to get those updates right without causing instability during online learning.

Rosa: The reward design they use is dense, shaped by terms like r t = -d one - d two + two when d two delta, which encourages reaching the object and then moving toward the goal with a terminal bonus upon success <ref:2604.01142#pg0>. That dense shaping is crucial for guiding the RL agent effectively during training.

Taro: If that reward function isn't well-designed, even a good policy will fail to learn the desired behavior in complex contact scenarios, so the reward structure seems tightly coupled with their success criteria <ref:2604.01142#pg0>.

Dev: So, if we're talking about long-term deployment under distribution shift, how long do you think this system could reliably operate before requiring a manual recalibration or a full retraining cycle? Rosa, what's the limitation they admit in their study?

Rosa: They acknowledge that the bounded ES method is used for online adaptation when conditions depart from training, but they also point out that the overall robustness relies on how well the initial RL policy handles those shifts <ref:2604.01142#pg0>. The system's performance improvement is demonstrated under specific types of distribution shift, like friction patches or evolving goals <ref:2604.01142#pg0>.

Taro: The limitation they state is that the success hinges on the RL policy providing rapid control when conditions are near training data, and the ES component taking over afterward <ref:2604.01142#pg1>. If the shift is so extreme that it invalidates what RL learned initially, then even this hybrid approach might struggle to recover quickly <ref:2604.01142#pg0>.

Dev: That makes sense; if the system is operating way outside the learned manifold, neither controller performs optimally on its own. But considering the loop rate and latency we discussed earlier, how fast does that switch between RL and ES actually need to happen to be effective in a dynamic push task?

Paper summary: Rosa: The switching time t c is defined as the moment when the end-effector first comes into contact with the object <ref:2604.01142#pg0>. This timing is crucial because it defines the boundary between the fast entry phase and the adaptation phase where ES takes over <ref:2604.01142#pg0>.

Taro: So, it’s not just a fixed time delay, but a state-based switch triggered by physical interaction—that makes sense for handling contact-rich manipulation <ref:2604.01142#pg0>. It ties the control strategy directly to the physical reality of the task.

Dev: I think that state-based trigger is what gives it more control over latency compared to a fixed time switch, as you only engage ES when interaction has actually occurred, which should keep things tighter <ref:2604.01142#pg1>. We need to check if that contact detection itself introduces too much noise into the switching decision <ref:2604.01142#pg0>.

Rosa: And that's the exciting part, Dev; the implication is that for long-horizon manipulation tasks where things change after you grasp something, this combined RL and bounded ES controller can maintain substantially closer tracking when operating under distribution shift compared to using just the RL component alone <ref:2604.01142#pg0>.

Taro: That suggests a future where robots don't need perfect environmental models or perfectly known dynamics to perform complex manipulation tasks, provided they have that kind of adaptive mechanism built in <ref:2604.01142#pg1>. It really pushes the boundary on what we consider robust autonomy <ref:2604.01142#pg0>.

Dev: If this controller is effective for time-varying goals and spatially varying friction, I'm curious about the practical implications for industrial automation. Rosa, how long do you think a system built with this would need to operate in an uncontrolled industrial setting before we expect it to show significant degradation?

Rosa: The paper shows superior performance when operating conditions differ significantly from training, specifically mentioning scenarios with spatially varying friction patches <ref:2604.01142#pg0>. It demonstrates effectiveness in tracking a three dee time-varying goal where the reference evolves after grasp acquisition <ref:2604.01142#pg0>.

Taro: That means we could deploy these systems in dynamic environments, like assembly lines where material properties or tool placements might change slightly over time without needing constant reprogramming <ref:2604.01142#pg1>. It moves us closer to truly adaptable robotic agents.

Dev: So, it’s about improving the reliability of manipulation in unpredictable settings, which is exactly what control engineers are focused on; we need systems that don't fail when the environment isn't perfectly modeled <ref:2604.01142#pg1>. We have to worry about those failure modes under stress, though.

Rosa: The authors are clear that the hybrid approach provides a path to handling those stresses better by combining RL’s speed with ES’s guaranteed bounds during the critical contact phase <ref:2604.01142#pg0>. It gives us a more resilient foundation for complex physical interaction <ref:2604.01142#pg1>.

Conclusion: Rosa: So, we're wrapping up our discussion on "Planning for Change: Reinforcement Learning Combined with Bounded Extremum Seeking for Robotic Control under Distribution Shift." This paper essentially lays out a hybrid control strategy that uses deep reinforcement learning for initial rapid learning and then switches to bounded extremum seeking when the environment starts behaving unexpectedly.

Dev: That hybrid approach is certainly the core idea, Rosa; I'm thinking about how that switching mechanism manages latency in a real-time loop. The authors describe it as leveraging "the complementary strengths of the two controllers" at different times.

Taro: From an autonomy angle, what excites me most is how this system maintains tracking accuracy when goals or friction patches shift after the initial contact phase has begun. It shows a path toward systems that don't need a perfect model of their surroundings to succeed.

Rosa: I agree with Taro; the ability to adapt online during those contact-rich phases is what makes this work interesting for field robotics, especially in areas where conditions are constantly changing. The authors claim it performs better under distribution shift than RL alone in experiments involving varying friction patches and time-varying goals.

Dev: But we have to consider the practical reality of deployment, Rosa; how long can this system reliably function before we need a full recalibration? I'm concerned about the stability of that online adaptation when things get truly out of distribution.

Taro: The paper itself points out that the robustness depends heavily on how well the initial RL policy handles those shifts right after contact is established, so if the shift is too drastic, even this hybrid system might struggle to recover quickly. That’s a limitation they explicitly state.

Rosa: That makes sense; it shows that while we can get good results under distribution shift, we still need a solid foundation from the RL part of the system to make that online adaptation effective. The implications here are huge for any robot needing to work in dynamic, real-world industrial settings.

Dev: So, if this method works as described for long-horizon manipulation tasks where conditions change post-grasp, it could significantly extend the operational window for complex robotic applications outside of highly controlled lab environments. That's a big deal for deployment planning.

Taro: Indeed; this moves us closer to autonomous agents capable of handling unexpected physical interactions without constant manual intervention or system resets. It’s about building resilience into the control logic itself, which is a major step forward for autonomy research in robotics.

More episodes

← Home