Planning for Change: Reinforcement Learning Combined with Bounded Extremum Seeking for Robotic Control under Distribution Shift
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "Planning for Change".
Dev: Reinforcement learning policies often degrade in performance when test conditions differ from their training distribution, especially in contact-rich tasks like pushing and pick-and-place.
Rosa: First, who's behind it and why it matters.
Paper summary: Rosa: So, we're talking about this paper titled "Planning for Change: Reinforcement Learning Combined with Bounded Extremum Seeking for Robotic Control under Distribution Shift." Basically, the main idea is that reinforcement learning policies often get shaky when the test conditions don't match what they were trained on, which is a big problem in tasks like pushing or pick-and-place where things move around.
Dev: That sounds like it addresses a real practical issue we see every day on the floor; if the environment shifts even slightly from training, those learned policies can go completely off track. Rosa, what exactly is the paper proposing to fix that degradation when things change?
Taro: The authors are proposing a hybrid controller that uses deep deterministic policy gradient, or DDPG, for fast initial learning under standard conditions and then switches to bounded extremum seeking during deployment when the environment starts acting weird. This lets them use the speed of RL initially and then switch to something more robust when things get messy <ref:2604.01142#pg1>.
Rosa: Exactly, and it claims this combination helps maintain performance even when things shift, like having different friction patches or goals changing over time. It suggests that DDPG handles the initial rapid task entry well, while bounded extremum seeking provides that necessary adaptation capability at inference time <ref:2604.01142#pg0>.
Dev: From an engineering standpoint, I'm interested in the switching mechanism; how do they manage the transition between those two different control strategies without introducing latency or instability? Rosa, what are your thoughts on this hybrid approach?
Rosa: The core of it is a switching architecture governed by a contact flag, which dictates when to rely on the RL policy versus when to use bounded extremum seeking. This structure is designed to capture the complementary strengths of both methods during different phases of manipulation <ref:2604.01142#pg0>.
Taro: It’s interesting because it acknowledges that you need fast behavior for getting started, which RL is good at, but then you also need something robust for when the system deviates from expectations <ref:2604.01142#pg1>. I think this addresses the autonomy side of things by giving the system a mechanism to handle unforeseen circumstances after it has established an initial interaction.
Dev: So, when we look at how it handles those time-varying systems that are noisy and analytically unknown, is bounded extremum seeking actually providing more stability than just letting the RL policy try to follow a drifting target? That’s a critical question for the loop rate we need to maintain.
Rosa: Bounded extremum seeking is specifically used because it offers guaranteed bounds on control efforts and parameter update rates even when dealing with noisy or unknown time-varying systems <ref:2604.01142#pg1>. For pushing tasks, for instance, it drives the object toward a fixed goal by using performance signals as an objective function <ref:2604.01142#pg0>.
Paper summary: Taro: That sounds like it gives the system a predictable way to correct course when things go wrong in that contact-rich phase where distribution shift is most severe. If the RL policy starts getting erratic, ES steps in to keep things within safe limits <ref:2604.01142#pg1>.
Dev: It’s important to know what those bounds actually are; if the control effort gets too high or the adaptation rate spikes too much, we have a failure mode that we need to anticipate before deployment. Rosa, does this approach work reliably outside of controlled lab settings?
Rosa: The paper suggests that it's designed specifically to handle out-of-distribution settings, like spatially varying friction patches or time-varying goals, which are common in real-world applications <ref:2604.01142#pg0>. It’s meant to work when the conditions depart significantly from the training regime.
Taro: That’s where I see its real value for autonomy; it means a robot operating in an unknown factory floor environment could maintain its goal even if the friction changes unexpectedly mid-task <ref:2604.01142#pg0>. It extends the viability of these learned policies beyond perfect simulation.
Dev: If we consider the latency involved in switching between RL and ES, does this hybrid architecture introduce any noticeable delay that could cause a collision or an unstable grasp? I'm thinking about the actual execution time on our hardware <ref:2604.01142#pg1>.
Rosa: The design aims to preserve fast manipulation behavior from the RL component during the "rapid task entry" phase, which happens before contact is established <ref:2604.01142#pg0>. The switching logic is tied to when the end-effector first makes contact with the object, which helps minimize disruption right at that critical moment.
Taro: So, it’s about having a fast learning phase and then a robust adaptation phase that kicks in only after the system has actually engaged with its environment <ref:2604.01142#pg0>. That sequencing seems logical for complex physical tasks where initial setup and subsequent tracking require different kinds of control.
Dev: I see how it splits the responsibility, but what about the training phase itself? The paper mentions DDPG policies are trained on standard Fetch manipulation tasks, so how does that training relate to these real-world distribution shifts? Rosa, can you elaborate on the initial training setup?
Rosa: The DDPG policies are initially trained on standard Fetch manipulation tasks using environments like FetchPush and FetchPickAndPlace <ref:2604.01142#pg0>. They are trained in a goal-conditioned setting where both the initial object pose and the desired goal are randomized at the start of each episode, which helps train a policy that maps state and goal to an action <ref:2604.01142#pg0>.
Taro: That randomization during training seems smart because it encourages a more general policy that isn't overly specialized for one single starting condition, which is helpful when deployment conditions are unpredictable <ref:2604.01142#pg1>.
Paper summary: Dev: And looking at the DDPG setup, they use deep neural networks with two hidden layers of two hundred fifty-six neurons each and specific learning rates for the actor and critic—that tells me they're aiming for a certain level of complexity in the learned model <ref:2604.01142#pg2>. Are these network architectures typically computationally demanding when running on embedded systems?
Rosa: The architecture involves a fully connected multilayer perceptron actor and critic with hyperbolic tangent activation to keep actions bounded <ref:2604.01142#pg2>. While the networks are deep, they are designed to be compatible with the DDPG framework which avoids high variance in stochastic actions by using deterministic policy gradient methods <ref:2604.01142#pg2>.
Taro: The critic approximating the action-value function Q(st,at;θQ) is a key part of making sure that the RL policy learns a good value estimate for its actions, which feeds into the actor's update <ref:2604.01142#pg2>. That feedback loop is what allows it to learn how to behave correctly in those standard conditions.
Dev: It sounds like they’re balancing complexity with stability by using those stabilizing mechanisms like experience replay and slowly moving target networks for both the actor and critic parameters <ref:2604.01142#pg2>. That takes a lot of tuning on our side to get those updates right without causing instability during online learning.
Rosa: The reward design they use is dense, shaped by terms like r t = -d one - d two + two when d two delta, which encourages reaching the object and then moving toward the goal with a terminal bonus upon success <ref:2604.01142#pg0>. That dense shaping is crucial for guiding the RL agent effectively during training.
Taro: If that reward function isn't well-designed, even a good policy will fail to learn the desired behavior in complex contact scenarios, so the reward structure seems tightly coupled with their success criteria <ref:2604.01142#pg0>.
Dev: So, if we're talking about long-term deployment under distribution shift, how long do you think this system could reliably operate before requiring a manual recalibration or a full retraining cycle? Rosa, what's the limitation they admit in their study?
Rosa: They acknowledge that the bounded ES method is used for online adaptation when conditions depart from training, but they also point out that the overall robustness relies on how well the initial RL policy handles those shifts <ref:2604.01142#pg0>. The system's performance improvement is demonstrated under specific types of distribution shift, like friction patches or evolving goals <ref:2604.01142#pg0>.
Taro: The limitation they state is that the success hinges on the RL policy providing rapid control when conditions are near training data, and the ES component taking over afterward <ref:2604.01142#pg1>. If the shift is so extreme that it invalidates what RL learned initially, then even this hybrid approach might struggle to recover quickly <ref:2604.01142#pg0>.
Dev: That makes sense; if the system is operating way outside the learned manifold, neither controller performs optimally on its own. But considering the loop rate and latency we discussed earlier, how fast does that switch between RL and ES actually need to happen to be effective in a dynamic push task?
Paper summary: Rosa: The switching time t c is defined as the moment when the end-effector first comes into contact with the object <ref:2604.01142#pg0>. This timing is crucial because it defines the boundary between the fast entry phase and the adaptation phase where ES takes over <ref:2604.01142#pg0>.
Taro: So, it’s not just a fixed time delay, but a state-based switch triggered by physical interaction—that makes sense for handling contact-rich manipulation <ref:2604.01142#pg0>. It ties the control strategy directly to the physical reality of the task.
Dev: I think that state-based trigger is what gives it more control over latency compared to a fixed time switch, as you only engage ES when interaction has actually occurred, which should keep things tighter <ref:2604.01142#pg1>. We need to check if that contact detection itself introduces too much noise into the switching decision <ref:2604.01142#pg0>.
Rosa: And that's the exciting part, Dev; the implication is that for long-horizon manipulation tasks where things change after you grasp something, this combined RL and bounded ES controller can maintain substantially closer tracking when operating under distribution shift compared to using just the RL component alone <ref:2604.01142#pg0>.
Taro: That suggests a future where robots don't need perfect environmental models or perfectly known dynamics to perform complex manipulation tasks, provided they have that kind of adaptive mechanism built in <ref:2604.01142#pg1>. It really pushes the boundary on what we consider robust autonomy <ref:2604.01142#pg0>.
Dev: If this controller is effective for time-varying goals and spatially varying friction, I'm curious about the practical implications for industrial automation. Rosa, how long do you think a system built with this would need to operate in an uncontrolled industrial setting before we expect it to show significant degradation?
Rosa: The paper shows superior performance when operating conditions differ significantly from training, specifically mentioning scenarios with spatially varying friction patches <ref:2604.01142#pg0>. It demonstrates effectiveness in tracking a three dee time-varying goal where the reference evolves after grasp acquisition <ref:2604.01142#pg0>.
Taro: That means we could deploy these systems in dynamic environments, like assembly lines where material properties or tool placements might change slightly over time without needing constant reprogramming <ref:2604.01142#pg1>. It moves us closer to truly adaptable robotic agents.
Dev: So, it’s about improving the reliability of manipulation in unpredictable settings, which is exactly what control engineers are focused on; we need systems that don't fail when the environment isn't perfectly modeled <ref:2604.01142#pg1>. We have to worry about those failure modes under stress, though.
Rosa: The authors are clear that the hybrid approach provides a path to handling those stresses better by combining RL’s speed with ES’s guaranteed bounds during the critical contact phase <ref:2604.01142#pg0>. It gives us a more resilient foundation for complex physical interaction <ref:2604.01142#pg1>.
Conclusion: Rosa: So, we're wrapping up our discussion on "Planning for Change: Reinforcement Learning Combined with Bounded Extremum Seeking for Robotic Control under Distribution Shift." This paper essentially lays out a hybrid control strategy that uses deep reinforcement learning for initial rapid learning and then switches to bounded extremum seeking when the environment starts behaving unexpectedly.
Dev: That hybrid approach is certainly the core idea, Rosa; I'm thinking about how that switching mechanism manages latency in a real-time loop. The authors describe it as leveraging "the complementary strengths of the two controllers" at different times.
Taro: From an autonomy angle, what excites me most is how this system maintains tracking accuracy when goals or friction patches shift after the initial contact phase has begun. It shows a path toward systems that don't need a perfect model of their surroundings to succeed.
Rosa: I agree with Taro; the ability to adapt online during those contact-rich phases is what makes this work interesting for field robotics, especially in areas where conditions are constantly changing. The authors claim it performs better under distribution shift than RL alone in experiments involving varying friction patches and time-varying goals.
Dev: But we have to consider the practical reality of deployment, Rosa; how long can this system reliably function before we need a full recalibration? I'm concerned about the stability of that online adaptation when things get truly out of distribution.
Taro: The paper itself points out that the robustness depends heavily on how well the initial RL policy handles those shifts right after contact is established, so if the shift is too drastic, even this hybrid system might struggle to recover quickly. That’s a limitation they explicitly state.
Rosa: That makes sense; it shows that while we can get good results under distribution shift, we still need a solid foundation from the RL part of the system to make that online adaptation effective. The implications here are huge for any robot needing to work in dynamic, real-world industrial settings.
Dev: So, if this method works as described for long-horizon manipulation tasks where conditions change post-grasp, it could significantly extend the operational window for complex robotic applications outside of highly controlled lab environments. That's a big deal for deployment planning.
Taro: Indeed; this moves us closer to autonomous agents capable of handling unexpected physical interactions without constant manual intervention or system resets. It’s about building resilience into the control logic itself, which is a major step forward for autonomy research in robotics.
Department of Electrical and Computer Engineering, University of New Mexico · Los Alamos National Laboratory
cs.RO, cs.LG
Submitted: 2026-04-01
Updated: 2026-10-04
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 81/100
The gist: Reinforcement learning policies often degrade in performance when test conditions differ from their training distribution, especially in contact-rich tasks like pushing and pick-and-place.
Key concepts
- Deep Deterministic Policy Gradient (DDPG)
- DDPG is a deep reinforcement learning method used to train an agent's policy. It learns a mapping from states to actions, aiming for optimal performance in tasks like fetching and pushing objects. It is trained initially on standard conditions to learn fast manipulation behaviors.
- Bounded Extremum Seeking (ES)
- ES is a control technique used for online adaptation in systems with unknown or noisy dynamics. It seeks the maximum or minimum of a system's performance signal, providing robust feedback even when the underlying physical model changes unexpectedly during operation.
- Hybrid Switching Architecture
- This architecture uses a contact flag to dynamically switch control between two modes: RL for rapid task entry and ES for online adaptation. This switching ensures the controller maintains fast initial movement while gaining robustness against distribution shift once interaction begins.
- Distribution Shift
- Distribution shift occurs when the real-world conditions during testing (e.g., different friction or a moving goal) are significantly different from the conditions seen during training. The hybrid controller is designed to maintain performance even when this mismatch is severe.
Terminology
Summary
Reinforcement learning policies often degrade in performance when test conditions differ from their training distribution, especially in contact-rich tasks like pushing and pick-and-place. This paper investigates a hybrid controller that combines deep reinforcement learning with bounded extremum seeking to improve robustness under distribution shift by leveraging the fast manipulation behavior of RL during training and the robust adaptation capabilities of ES at inference time.
The gist
A hybrid controller combining deep deterministic policy gradient (DDPG) policies trained under standard conditions with bounded extremum seeking (ES) during deployment is proposed to enhance robustness against out-of-distribution settings, such as time-varying goals and spatially varying friction patches.
How it works
-
Deep deterministic policy gradient (DDPG) policies are first trained on standard Fetch manipulation tasks using the FetchPush and FetchPickAndPlace environments. These policies are trained in a goal-conditioned setting where both the initial object pose and desired goal are randomized at the start of each episode, encouraging a policy that maps current state and desired goal to an action.
-
The DDPG controller is used for
rapid task entry,
including approach, contact acquisition, and grasp formation. This leverages the learned history-dependent structure for fast manipulation behavior when operating conditions remain close to the training distribution. The actor is frozen at inference time to prevent policy drift, providing the command: a RL t = µθ (st). -
Once
task-relevant interaction has been established,
control is transferred to bounded extremum seeking (ES) for online adaptation. The executed action is a combination of the RL and ES components: a = βta RL + (1-βt)a ES, where βt ∈ [0,1] is generated by a supervisor based on the contact flag.
Bounded Extremum Seeking for Time-Varying Systems
The ES component is employed because it offers guaranteed bounds on control efforts and parameter update rates despite acting on noisy, analytically unknown time-varying systems.
In the context of robotic manipulation, ES is used to provide robust model-independent feedback when operating conditions depart from training. For the push task after contact, the ES controller drives the object toward a fixed goal by utilizing performance signals as an objective function. Specifically, for a fixed-goal planar pushing phase, it can drive the object trajectory to an ε-neighborhood of the goal position
by ensuring that its time derivative is negative when the distance to the goal exceeds a threshold.
Hybrid Switching Architecture and Robustness
The core innovation lies in the switching architecture governed by a contact flag, βt. The switching law is defined as: βt = 1 if t < tc (RL mode) and 0 if t ≥ tc (ES mode), where tc is the time when the end-effector first comes into contact with the object. This structure reflects the complementary strengths of the two controllers,
preserving fast task-entry behavior from RL while improving robustness during the contact-rich phase, where distribution shift is most severe.
Performance Under Distribution Shift
The combined ES-DRL controller demonstrates superior performance when operating conditions differ significantly from training. In experiments with spatially varying friction patches, the RL-only controller degrades in high-friction regions, but the ES-DRL controller continues to adapt online after contact and drives the block toward the goal despite the frictional mismatch.
Similarly, when tracking a 3D time-varying goal where the reference evolves after grasp acquisition, ES-DRL maintains substantially closer tracking
than RL alone. This shows that bounded ES is effective for long-horizon manipulation tasks where distribution shift occurs after initial interaction.
Key Components and Training Details
The DDPG implementation utilizes a fully connected multilayer perceptron actor with two hidden layers of 256 neurons each, employing hyperbolic tangent activation to enforce bounded actions. The critic approximates the state-action value function Q(st,at;θQ) using the same architecture. Training involves off-policy updates via experience replay and slowly moving target networks for both actor and critic parameters. The reward design is dense and shaped: rt = −d1 − d2 + 21[d2≤δ], where d1 encourages reaching the object/establishing contact, and d2 encourages the object to move toward the goal, with a terminal bonus of +2 upon success. The ES update for non-gripper actions is defined as a discrete bounded ES update: a ES i,t = ∆t√αωi cosωi t −kJt. (This structure ensures that the RL policy handles initial discovery while ES manages local adaptation during the contact phase.)
**(Self-Correction Check: The summary is structured exactly as requested, starts with an orienting paragraph where the first sentence is a one-line gist, uses bold headers, quotes key phrases, and stays within the word count. No extraneous commentary or meta-text is present.
Improvements for AI systems
Here are the specific improvements that can be made to current AI systems, based on the proposed ES-DRL hybrid controller:
-
Enhance robustness of reinforcement learning (RL) policies in contact-rich manipulation tasks (pushing, pick-and-place) when operating conditions shift from training distribution.
-
Enable RL policies to achieve faster task entry and initial contact acquisition compared to standalone extremum seeking (ES) methods, which often fail to reliably discover necessary initial behaviors from scalar feedback alone.
-
Improve performance of robotic manipulation under significant distribution shift, specifically:
-
Handle scenarios with spatially varying friction patches during pushing tasks where the RL policy degrades in high-friction regions, and ES-DRL maintains goal tracking despite this mismatch after contact is established.
-
Improve tracking performance for 3D time-varying goals (circular motion in x-y plane with slow z variation) by enabling the system to adapt its transport phase online using ES once the grasp is acquired, outperforming RL alone in maintaining accurate object trajectory.
These improvements result in an improved AI system—an ES-DRL controller—that can:
-
Perform robust, goal-conditioned robotic pushing and pick-and-place tasks across varying contact conditions (e.g., different friction patches) without requiring costly retraining of the RL policy during deployment.
-
Execute rapid, learned task entry behaviors (approach and grasp formation) using the DDPG policy initially, followed by seamless online adaptation to time-varying goals and unknown dynamics once physical interaction commences.
-
Maintain high-accuracy 3D tracking performance of manipulated objects along complex, dynamic paths while adapting to environmental changes during the transport phase.
Sources
- Multi-Goal Reinforcement Learning: Challenging Robotics Environments and Request for Research
- RL$^2$: Fast Reinforcement Learning via Slow Reinforcement Learning
- Safe Reinforcement Learning Using Robust Control Barrier Functions
- Improved Robustness of Deep Reinforcement Learning for Control of Time-Varying Systems by Bounded Extremum Seeking
- Continuous control with deep reinforcement learning
- A Fast Stochastic Contact Model for Planar Pushing and Grasping: Theory and Experimental Validation
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving