Learning On The Job: Zero-Shot Task Execution under Parametric Uncertainty via Trajectory-Parametrized Dual Control

arXiv:2510.20483 · cs.RO, cs.IT, math.IT · Submitted 2025-10-23 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "Learning On The Job".

Dev: This work addresses "the problem of robot manipulation tasks under unknown dynamics, such as pick-and-place tasks under payload uncertainty,

Rosa: First, who's behind it and why it matters.

Title and authors: Rosa: So we’re diving into "Learning On The Job: Zero-Shot Task Execution under Parametric Uncertainty via Trajectory-Parametrized Dual Control." This paper tackles that tricky problem of robots performing tasks, like pick-and-place, when they don't know the exact mass or dynamics of what they are handling.

Dev: It sounds intense, Rosa; it's about getting the robot to figure out its own dynamics while simultaneously trying to complete the actual movement.

Taro: Exactly! What I find interesting is how it frames this as a dual control problem seeking a closed-loop optimal control problem that handles that parameter uncertainty directly. It moves beyond just trying to identify parameters separately from controlling the task.

Rosa: That's the core idea, Taro; they simplify it by setting up the feedback policy structure beforehand to include an explicit adaptation mechanism. Dev, from your side of things, how does this move away from traditional methods that might decouple identification and control?

Dev: Well, usually you have separate phases for learning parameters and then running a fixed controller based on those estimates; this paper seems to propose something different where the reference trajectory generation is inherently tied to the control design itself. It’s about designing the motion in a way that forces the system to provide useful information about those unknown parameters during execution, which addresses some of those issues with decoupling.

Rosa: That makes sense; so instead of guessing the parameters first, you design a path that is specifically probing the uncertainty in a way that helps the adaptation law work better. Taro, does this active exploration during execution change how we think about system robustness?

Taro: It definitely shifts our thinking toward active exploration during task execution rather than just pre-task identification, which I think is more relevant for real-world deployment where you can't always stop to learn parameters beforehand. If the world misbehaves mid-task, this framework suggests the robot has a better chance of maintaining accuracy because it’s constantly gathering data that actually matters for its control performance.

Dev: From a control engineering standpoint, I'm curious about how this affects our loop rates and latency. The paper mentions that both their trajectory generation methods reason over the Fisher information as a natural side effect of their formulations while simultaneously pursuing optimal task execution. That implies the computational overhead of generating these dual-control trajectories must be manageable within our real-time constraints.

Rosa: That’s a valid concern, Dev; if the generation process is too slow, the benefits of that optimized trajectory are lost because we miss the moment where parameter estimation and task execution need to happen together. What about those two methods they propose for reference trajectory generation?

Title and authors: Dev: Method one directly embeds uncertainty into robust optimal control methods that minimize the expected task cost, which sounds like a direct way to keep things tightly coupled within the optimization process itself. Whereas method two looks at minimizing an optimality loss, which measures how sensitive the parameter-relevant information is with respect to task performance.

Taro: The idea that both approaches reason over the Fisher information simultaneously while optimizing for the task is compelling because it suggests a unified objective function rather than two separate, potentially conflicting goals. That’s where I see the real potential for handling unpredictable situations in dynamic environments.

Rosa: And looking at those methods, they demonstrate effectiveness on a pick-and-place manipulation task under payload uncertainty. It seems they’ve shown a concrete result on a fundamental robotic challenge. Dev, what does this mean for the kind of model-based control we’re using?

Dev: It means that even if you're using adaptive controllers, like Natural Adaptive Controller, this framework can guide the system toward more accurate parameter updates because it provides a trajectory that is specifically designed to be informative. It improves the consistency of those estimates compared to what I see in some of the literature.

Rosa: That speaks to how much better we can tune our control policies when we have this kind of informed reference input. Taro, if a robot encounters an unexpected external force during transport, does this trajectory-parametrized dual control system handle that deviation well?

Taro: It should be more resilient than a purely nominal trajectory because the framework is designed to account for the uncertainty in dynamics and adjust its internal model based on what it’s sensing during the movement. It suggests a level of active adaptation that makes the system behave more predictably when faced with unexpected external disturbances.

Dev: I worry about stability, though; since we're dealing with closed-loop optimal control problems, any aggressive exploration strategy needs careful tuning to ensure we don't introduce oscillations or instability in the actual hardware loop. The latency of generating that reference trajectory has to be very low for this to translate into real-time performance.

Rosa: That’s the practical hurdle, isn't it; translating a complex mathematical formulation into a stable, fast piece of software running on physical hardware. We need to see if this works outside the lab environment for extended periods, not just in controlled simulations.

Taro: I think its application outside the lab could be huge because it addresses the fundamental issue of model mismatch in unstructured environments. Imagine a warehouse robot picking up items whose properties change constantly; this approach seems built for that kind of messy reality.

Title and authors: Dev: I agree with Taro on the real-world relevance, but we have to be careful about the failure modes. If the parameter adaptation law itself gets stuck in a local minimum because of noisy data, or if the reference trajectory generation becomes overly sensitive to noise, that’s where we could see a hard failure in operation.

Rosa: So, we've got this dual control formulation based on optimality loss minimization and the direct expected task cost optimization. It seems like the methodology is quite powerful for tackling uncertainty in manipulation tasks. Dev, what's your take on the practical implications of deploying this kind of integrated learning and control?

Dev: The implication is that we might move toward systems where the planning stage isn't just about finding a path, but actively shaping that path to improve our understanding of the system while completing the mission. This reduces reliance on perfect prior models for every single deployment scenario.

Taro: I think this could lead to robots that are much more adaptable to novel scenarios because they aren't stuck following a pre-programmed plan when things go wrong; they’re actively learning how to move better under the conditions they encounter.

Rosa: It certainly sounds like a significant step forward in how we design autonomous agents for physical tasks. Dev, looking ahead, what's the next logical step for this research? Are there limitations they acknowledge that need addressing?

Dev: They do mention that their proposed structure is simplified by predefining the feedback policy, and they also show that omitting a sensitivity-based variant of their framework provides an alternative approach. That suggests there's room for further refinement in tailoring the reference generation to specific controller dynamics.

Taro: I think future work should focus on making this framework even more general so it doesn't rely so heavily on the pre-defined structure of the feedback policy, allowing it to be truly autonomous in its exploration strategy. That would make it far more robust across different robot types and tasks.

Rosa: It sounds like the trajectory-parametrized dual control approach presented in "Learning On The Job: Zero-Shot Task Execution under Parametric Uncertainty via Trajectory-Parametrized Dual Control" offers a solid, integrated way to handle uncertainty during manipulation. Dev, Taro, that’s a lot to digest before we move on to what this means for the broader field.

Dev: I think the key is that it bridges the gap between theoretical optimal control and practical task execution by embedding parameter estimation directly into trajectory generation in a structured way.

Taro: And for me, it’s about moving from reactive problem-solving to proactive, information-gathering control during operation, which is what we need for truly autonomous agents.

Rosa: Well, that’s our rundown on this fascinating paper today. We'll be keeping an eye on these kinds of integrated approaches as we continue to explore how AI can help robots handle the messy reality of physical tasks.

The paper's summary: Rosa: So, essentially, this paper is proposing a way for robots to perform tasks like pick-and-place when they have no prior knowledge about things like payload mass or exact friction dynamics by designing the motion itself to gather that information while simultaneously hitting the target.

Dev: That’s right, Rosa; they're taking that notoriously hard dual control problem and simplifying it by focusing on how the reference trajectory is generated to include both task completion and parameter adaptation. It seems like a structured way to handle that uncertainty we usually treat with separate identification steps.

Taro: I’m really interested in what this means for real-world autonomy; if the robot can actively explore its environment and learn about the object's properties while it's moving, that makes it much more adaptable to unexpected situations during execution.

Rosa: Exactly, Taro; they show that by designing those reference trajectories thoughtfully, we get faster and more accurate task performance because the control is already aware of what information it needs. It’s not just about reaching the end point; it’s about learning along the way.

Dev: From an engineering standpoint, I gotta ask how this plays out in terms of loop rates; since they're optimizing a dual control problem, generating those trajectories can be computationally heavy, so we need to make sure that this whole system runs fast enough for real-time control.

Taro: And that’s where the active exploration aspect becomes crucial; if the robot encounters something unexpected mid-move, this framework suggests it has a built-in mechanism to adjust its behavior based on what it's sensing in real time, rather than failing because its model is wrong.

Rosa: It really shows how these two things—optimizing the task and acquiring new information—don’t have to be separate concerns; they can be integrated into one cohesive optimization problem that drives the entire process.

Dev: That unified approach sounds promising for stability, but I still need assurance on the failure modes; if the parameter estimation gets biased by noisy sensor data or if our chosen adaptation law isn't robust, could we end up with a system that performs poorly instead of better?

Taro: The paper addresses that by considering how the Fisher information guides their formulation, which suggests a principled way to balance exploration versus exploitation without just blindly chasing every piece of noisy data.

Rosa: That’s exactly what excites me about it; this feels like it moves us closer to agents that can actually operate in messy, unstructured environments without needing a perfect map beforehand.

Dev: I still see the practical challenge in deployment duration; we need to know if this level of active exploration and parameter adaptation can sustain itself over long operational periods without degrading performance or introducing instability over time.

Taro: The implication for autonomy is that robots won't be limited to pre-programmed paths when they encounter a novel object; they can dynamically adapt their control strategy on the fly based on what they learn about that object during the manipulation.

Rosa: It really feels like we’re looking at a significant step toward systems that are truly capable of zero-shot execution in physical tasks, and I wonder how long this kind of performance can be maintained outside of a perfectly controlled lab setting.

The paper's improvements: Rosa: So, this paper goes beyond just presenting two methods for trajectory generation; it actually suggests refining how we structure the entire feedback policy to explicitly include an adaptation mechanism that’s already aware of the task uncertainty.

Dev: That's a crucial point, Rosa; it means they aren't just tweaking the path in isolation; they are designing the path based on what kind of controller and adaptation law we plan to use, which gives us more control over how parameter estimation actually behaves during execution.

Taro: I see that as a major improvement because it connects the planning phase directly to the learning phase; it’s not just about finding a good trajectory but designing one that is optimized for both doing the task and gathering high-quality data simultaneously.

Rosa: Exactly, Taro; this moves us toward a more sophisticated kind of active exploration where the robot doesn't just randomly probe; it probes in a way that maximizes its information gain relative to its need to complete the specific manipulation task at hand.

Dev: From my side, I’m looking at the implications for system robustness again; if we can structure this feedback policy better, we might find ways to mitigate those stability issues that arise when adaptation and control are happening in a tight loop under uncertainty.

Taro: And that leads me to thinking about when this stuff will actually be usable; I mean, if the robot's exploration strategy is informed by the task objective in this way, it should handle misbehavior much more gracefully than systems that just rely on pre-set behaviors.

Rosa: It seems like a really strong direction for future work because they've laid out a framework that balances information acquisition and control performance systematically within one optimization problem, which is something we’ve been chasing.

Dev: I agree with Rosa; the structure they propose sounds like it could lead to more predictable closed-loop behavior, provided the underlying mathematical formulation holds up under real-world noise and latency constraints.

Taro: If this framework can be generalized beyond pick-and-place, then we could see autonomous systems operating in much messier environments where dynamic changes are constant and unexpected.

Rosa: It certainly has potential for those more complex scenarios; it shows the path forward for designing reference trajectories that actively guide parameter estimation alongside task completion.

Dev: I still have to stress that the real-time computational load of generating these highly informed, dual-control trajectories is something we need to rigorously test; if the generation takes too long, we lose all the benefit of having such a smart reference input.

Taro: We need those tests, Dev; because if it runs fast enough and proves robust in simulation, then this concept could significantly impact how we build robots that can truly learn and operate on the job under uncertainty.

Conclusion: Rosa: So, to wrap up this discussion on "Learning On The Job: Zero-Shot Task Execution under Parametric Uncertainty via Trajectory-Parametrized Dual Control," we've seen how this research integrates parameter estimation directly into trajectory design through dual control optimization.

Dev: Right, it’s a solid piece of work that tackles the hard problem of uncertainty in manipulation by making the planning process inherently aware of what information it needs to gather while executing the task.

Taro: I think what really stands out is how this moves us toward truly autonomous systems that can adapt their control strategy on the fly when things get unexpected during operation.

Rosa: Exactly, Taro; it sets a clear direction for building robots that aren't just following pre-set paths but are actively learning and adjusting based on the information they encounter in real time.

Dev: From an engineering view, I’m still focused on the practical hurdles we have to jump; how long can we expect this level of active exploration to stay stable and perform consistently outside of a highly controlled lab environment?

Taro: That’s a fair concern, Dev; if the system can maintain its performance over extended periods in messy conditions, then it could fundamentally change how we deploy robots for tasks where the exact physical properties of objects are never perfectly known.

Rosa: It really feels like this paper lays out a very practical roadmap for developing more intelligent robotic agents that can handle real-world variability without constant manual reprogramming.

Dev: I'm just waiting to see if the computational overhead of generating those optimized trajectories can be managed effectively within the strict loop rate requirements we need for stable control.

Taro: If it proves robust, it suggests a new way for robots to tackle novel scenarios by prioritizing learning during execution, which is something we need to explore further in autonomy research.

Rosa: Well, that’s our rundown on "Learning On The Job: Zero-Shot Task Execution under Parametric Uncertainty via Trajectory-Parametrized Dual Control." It's a lot of exciting stuff for the field.

Dev: I think it shows the way forward for integrating identification and control in a single optimization framework, which is something we need to keep watching closely as we push our own hardware capabilities.

Taro: For me, this paper opens up possibilities for robots that are far more flexible and resilient when faced with dynamic physical realities during manipulation tasks.

Rosa: We’ll keep an eye on this kind of integrated approach as we continue to explore how AI can help robots handle the messy reality of physical tasks.

Victor Vantilborgh, Hrishikesh Sathyanarayan, Guillaume Crevecoeur, Ian Abraham, Tom Lefebvre

Department of Electromechanical, Systems and Metal Engineering, Ghent University · Department of Mechanical Engineering, Yale University

cs.RO, cs.IT, math.IT

Submitted: 2025-10-23

Updated: 2026-09-29

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

Importance score: 82/100

The gist: This work addresses "the problem of robot manipulation tasks under unknown dynamics, such as pick-and-place tasks under payload uncertainty, where active exploration and online parameter adaptation

Key concepts

Dual Control Problem
This is a problem where a system must simultaneously try to complete a task while also estimating unknown parameters about its environment or dynamics. The paper frames the solution as seeking a closed-loop optimal control problem that handles this parameter uncertainty directly, rather than separating identification and control.
Trajectory-Parametrized Dual Control
This method involves designing the reference trajectory itself to be aware of the unknown parameters. Instead of guessing parameters first, the motion is shaped to probe uncertainty in a way that helps the system adapt its internal model during execution.
Active Exploration
This refers to a robot actively gathering data about its environment and object properties while it is performing a task, rather than just following a pre-programmed path. The framework suggests this proactive learning makes the robot more adaptable to unexpected situations mid-task.
Fisher Information
This concept is used in the paper to guide the formulation of trajectory generation methods. It suggests a principled way to balance exploration (gathering new information) against exploitation (achieving task performance) without blindly chasing noisy data.

Terminology

Summary

This work addresses "the problem of robot manipulation tasks under unknown dynamics, such as pick-and-place tasks under payload uncertainty, where active exploration and online parameter adaptation during task execution are essential to enable accurate model-based control. The problem is framed as dual control seeking a closedloop optimal control problem that accounts for parameter uncertainty. The authors simplify the dual control problem by predefining the structure of the feedback policy to include an explicit adaptation mechanism." They propose two methods for reference trajectory generation:

  1. The first directly embeds parameter uncertainty in robust optimal control methods that minimize the expected task cost.

  2. The second method considers minimizing the so-called optimality loss, which measures the sensitivity of parameter-relevant information with respect to task performance. The authors observe that both approaches reason over the Fisher information as a natural side effect of their formulations, simultaneously pursuing optimal task execution.

They demonstrate the effectiveness of our approaches for a pick-and-place manipulation task. They show that designing the reference trajectories whilst taking into account the control enables faster and more accurate task performance and system identification while ensuring stable and efficient control.

The paper makes the following contributions:

A dual control formulation based on optimality-loss minimization, designing the reference trajectory, incorporating closed-loop controller dynamics and parameter adaptation alongside task performance,

"an alternative dual control formulation that directly optimizes for the expected task cost under parameter uncertainty. Moreover, we show that a simplified, sensitivity-based variant of our framework - omitting the arXiv:2510.20483v1 [cs.RO] 23 Oct 2025"

The paper is structured as follows: "Section II provides a discussion on related work, Section III introduces preliminaries and Section IV formalizes the problem statement. Section V presents the proposed methodology. Results and conclusion are presented in Section VI and VII."

In related work, the authors situate their work at the intersection of optimal experiment design, adaptive and robust control, and learning-based approaches for robotic systems under uncertainty. They discuss prior work in Optimal Experiment Design (OED), Adaptive and Robust Control, Dual Control and Exploration. The paper notes that prior work either focuses solely on information gain, remains task-agnostic, or designs trajectories without considering the sensitivity of the chosen controller and adaptation law. The authors state that their work addresses this gap by proposing a framework for task- and control aware reference design, systematically balancing information acquisition and control performance within a unified optimization problem.

In Section IV (Problem Formulation), they consider the system dynamics as x˙ = f(x, u, θ) (11) where x is the state, u is the control input, and θ is an unknown set of parameters. Their goal is to solve min u Eθ∼P [J(u)] s.t. (11) (12) where the cost function is defined as J(u) = m(x(T)) + Z T 0 l(t, x(t), u(t))dt (13) and requires the controller to achieve two objectives simultaneously: (i) execute the task optimally, and (ii) identify parameters, but only insofar as this improves task performance. They identify this formulation as a dual control problem [17]. They note that the separation principle does not hold, rendering the dual control problem notoriously unsolvable [18], as is emphasized by the following cost decomposition which splits the cost into past and future terms.

The authors propose a modified control system structure:

"x˙ = f(x, u, θ) (16a)

"u = π(t, x, d, ˆθ) (16b)

˙ˆθ = ρ(t, x, d, ˆθ) (16c)

where d ∈ R nd represents a deterministic control design signal. They state that "The design, d influences the closed-loop behavior directly by shaping the state trajectory through the control policy, π, and indirectly by affecting the informativeness of the data available for parameter adaptation via ρ, and thus the quality of the estimate, ˆθ, regardless of the parameter’s true value, θ."

In Section V (Dual Control Reference Generation), they consider two active learning objectives to determine the control design, d:

"A. Robust optimization: In our first approach we simply minimize the expected value of (12) subject to (16) instead of (11). In that sense the optimization problem is simply made aware of the parameter estimation scheme inherent to the control."

The objective is "min d Eθ∼p(θ¯) [J(d)] (17).

Improvements for AI systems

Here are the specific improvements that can be made to AI systems based on this research, along with what those improved systems could achieve:


) Improved AI Systems: Dual-Control Reference Generation for Robust Manipulation

The core improvements stem from moving beyond decoupled identification/control strategies toward a unified, dual-control optimization framework. The resulting system will exhibit superior performance in dynamic, uncertain environments by simultaneously optimizing task execution and parameter estimation.

The system will utilize a reference trajectory generation module that explicitly incorporates both the expected task cost and an exploration term derived from the Fisher Information Matrix (FI) or optimality loss minimization (JOED). This moves beyond simple nominal path planning or purely information-maximizing trajectories.

This dual-control reference generator will be informed by a structured feedback policy, allowing it to design trajectories that are not just task-relevant but also sensitive to the specific closed-loop controller and adaptation law being used (e.g., Natural Adaptive Controller vs. Computed Torque Control).

The resulting AI system will be capable of performing active exploration during execution rather than relying on a separate, pre-task identification phase followed by control design. The controller will actively excite the system dynamics in ways that yield high-information data about unknown parameters (like payload mass or inertia) while simultaneously striving to reach the desired pick-and-place target accurately.

The AI system will exhibit enhanced passive robustness. When deployed with a fixed, non-adaptive controller, it can generate motion plans that are inherently robust against model mismatch by considering the full distribution of possible payloads (using methods like Gaussian Mixture Models to represent parameter priors). This allows the system to avoid configurations or accelerations that would amplify errors caused by unknown parameters.

The improved AI system will demonstrate superior performance across different control paradigms:

  • It can achieve lower and more consistent final pose errors compared to nominal task trajectories, even when using adaptive controllers (like NAC) that might otherwise struggle with parameter estimation consistency.

  • It provides a better balance than purely information-maximizing trajectories (like FIM-based ones), which often sacrifice final tracking accuracy for excessive data collection.

) What the Improved AI System Can Do: Specific Applications

The improved AI system can be specifically applied to:

Manufacturing and Logistics Robots: Perform high-precision pick-and-place tasks on objects with unknown or varying payloads (e.g., grasping items of unknown mass, shape, or inertia). The system will not only place the object correctly but will also rapidly and accurately estimate the payload's physical properties during the movement itself, leading to significantly reduced errors in subsequent operations.

Autonomous Inspection and Assembly: Employ this system for tasks where precise manipulation is required on objects whose physical characteristics are not fully known beforehand (e.g., assembling components with variable mass or handling irregularly shaped parts). The dual control nature ensures that the robot collects the necessary data to build a reliable internal model of the object while executing its primary assembly goal.

Adaptive Control Systems: Enhance existing adaptive robotic systems by providing them with dynamically informed reference trajectories that guide their exploration phase, leading to faster convergence and more accurate parameter updates compared to using static or task-agnostic excitation signals.

Model-Based Control in Uncertain Environments: Enable model-based controllers (like Computed Torque Control) to operate effectively under significant physical uncertainty by generating trajectories that are explicitly designed to minimize the impact of that uncertainty on the closed-loop performance, achieving a level of robustness previously only achievable through worst-case design.

Abstract

Model-based control can achieve reliable task performance, but its effectiveness depends on the accuracy of the underlying model. Robots operating under model uncertainty must often adapt to previously unseen payloads, objects, and interaction dynamics to complete a task successfully. Conventional approaches typically rely on a dedicated task- and control-agnostic excitation phase for estimating physics parameters, delaying task execution and collecting potentially irrelevant data. In this paper, we instead consider zero-shot task execution under parametric uncertainty, where online model learning and control proceed concurrently during task execution. Our approach formulates reference generation within a dual control framework to produce task-relevant, informative trajectories that reduce parameter uncertainty in directions critical to task success. We predefine a feedback policy with an explicit parameter adaptation law and optimize the reference through the resulting adaptive closed-loop dynamics. We propose two formulations: 1) minimizing expected task cost under parameter uncertainty, and 2) minimizing optimality loss, which quantifies the degradation in task performance caused by planning with an incorrect parameter estimate. By actively optimizing for informative trajectories via a natural Fisher information measure, we tightly approach a lower bound on task-relevant parameter uncertainty while simultaneously achieving reliable task execution. Across diverse tasks, controllers, and adaptation laws, targeted exploration enables the identification required for successful task execution, allowing robots to learn relevant physical parameters while completing the task in a single attempt.

Sources

Related papers