Dense Temporal Motion Retargeting for Legged Robots

summary

Video file (mp4)

The gist

Legged robots can learn expressive whole-body skills from human and animal motions, but adapting these motions to a robot's specific dynamics requires careful adjustment of timing and control,

In short

Dense Temporal Motion Retargeting (DTMR) allows legged robots to accurately copy human or animal movements by adjusting the timing of those motions. It solves this by optimizing timing and control simultaneously for every step, enabling precise retargeting of dynamic actions like jumps while maintaining computational efficiency through parallel GPU processing.

Key concepts

Dense Temporal Motion Retargeting (DTMR)
This method jointly optimizes the robot's motor controls and the timing of those controls across every single control step. This dense approach means it can change the timing at every frame, allowing for highly precise adaptation of source motions, especially during fast or dynamic movements.
Phase Variable ($\phi_k$)
The phase variable is a decision variable that dictates how fast or slow the robot should execute a specific part of the motion compared to the original source. It allows the robot to stretch or compress time during execution, enabling it to match the timing characteristics of different reference motions.
Deformation Rate ($r_k$)
The deformation rate measures how much the timing is being changed at a specific step. A positive rate means the robot is speeding up its movement relative to the source, while a negative rate indicates it is slowing down, providing a mathematical way to quantify temporal adjustments.

Terminology used across episodes

This episode discusses

The paper

Dense Temporal Motion Retargeting for Legged Robots · Read on arXiv

Jaeryeong Kim, *Taerim Yoon*, *Jin Cheng*, *Sungjoon Choi*, *Stelian Coros*

Department of Computer Science, ETH Zurich · Department of Artificial Intelligence, Korea University

Legged robots can learn expressive whole-body skills from the motions of humans and animals. Due to the morphology gap between the source and the robot, however, the motion must be tailored to the dynamic properties of the robot. In particular, dynamic motions such as a jump require careful adjustment, since their timing and control are interdependent. We propose dense temporal motion retargeting (DTMR), which jointly optimizes timing and control within a single optimization, where dense means that the timing is adjusted for every control step. This dense formulation enables DTMR to deform only the parts of the motion that need a change in timing. The problem is solved with sampling-based model predictive control (MPC) in parallel on a GPU. We evaluate DTMR against baselines on two hours of human motion with four humanoid robots, where the results show that DTMR outperforms baseline methods, particularly on dynamic motions. We also show that allowing more temporal deformation yields more precise retargeting. We further compare DTMR with a baseline that optimizes the temporal dimension, where the result shows that DTMR retargets more precisely under the same deformation budget while being 19x faster. Lastly, policies trained on our references transfer to a real humanoid robot.

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "Dense Temporal Motion Retargeting for Legged Robots".

Dev: Legged robots can learn expressive whole-body skills from human and animal motions, but adapting these motions to a robot's specific dynamics requires careful adjustment of timing and control,

Rosa: First, who's behind it and why it matters.

Paper summary: Rosa: So, we're talking about this paper by Jaeryeong Kim et al., "Dense Temporal Motion Retargeting for Legged Robots," which is really focusing on how robots can learn complex movements from video. Rosa here, I want to start with the main idea—what’s the core thesis of this work?

Dev: Exactly, Rosa; basically, the authors are tackling the problem that when you try to teach a legged robot a human jump or some dynamic action using recorded motions, you have this big gap between what the source motion looks like and what your robot can physically do. So, they propose dense temporal motion retargeting as a way to bridge that gap by jointly optimizing timing and control in one single optimization process.

Taro: From an autonomy standpoint, I’m interested in how this dense formulation handles the interdependence of timing and control when dealing with dynamic movements like jumps or sudden changes in speed. If the robot is trying to mimic a rapid deceleration, how does this joint optimization help it manage those interdependent variables effectively?

Rosa: That's a good place to start, Taro; the paper claims that this dense formulation allows them to deform only the parts of the motion that actually need a change in timing, which they describe as adjusting the timing at every control step. This is presented as a way to maintain efficiency while achieving precise retargeting.

Dev: And from an engineering loop rate perspective, Rosa, that dense formulation means the timing adjustment—what they call the phase trajectory phi —is not just a separate calculation but is generated directly by a control input d phi k, which helps them manage those dynamics more tightly within the optimization framework.

Taro: I wonder about the practical application when things go wrong; if the world misbehaves, does this DTMR system have any inherent mechanism to handle unexpected external disturbances while still trying to follow the retargeted motion?

Rosa: The paper suggests that by allowing more temporal deformation, they found it leads to more precise retargeting, which means the robot can reproduce those dynamic motions with greater accuracy than before. This is a significant finding because it shows how fine-grained control over timing affects the final output of the imitation learning process.

Dev: That precision comes at a cost, and I want to talk about that trade-off; the authors show that precision and timing preservation are controlled by a phase-cost weight w phi, where a large w phi keeps the timing close to the source motion, while a small one permits more deformation.

Paper summary: Taro: So, if we're pushing for high fidelity in dynamic tasks, like hopping or landing gracefully, should we be favoring that larger timing constraint or are there situations where allowing that greater temporal deformation is beneficial?

Rosa: The paper points out that allowing more temporal deformation actually yields more precise retargeting when compared to baseline methods under the same deformation budget. This suggests we might need to tune that weight based on whether we prioritize strict timing adherence or maximizing kinematic accuracy.

Dev: And looking at the efficiency, they achieved results where DTMR outperformed baseline temporal optimization by about nineteen times faster when using a similar deformation budget. That speedup is pretty substantial for real-time applications on hardware.

Taro: A nineteen times speedup is notable, but I'm curious about the real-world deployment; Rosa, how long do you think this kind of retargeting system can reliably operate outside of a controlled lab environment before we start seeing significant degradation in its performance?

Rosa: The paper does show that policies trained using these references transfer to a real humanoid robot, and they constructed a two-hour dataset of dynamically feasible motions on which policies learn dynamic motions better. This indicates promise for real-world applicability, but the authors are focused on the transferability to those specific embodiments.

Dev: From my side, I'm concerned about the implementation details; they solve this using sampling-based model predictive control in parallel on a GPU, which is computationally intensive, so we have to consider that loop rate and any potential failure modes in that MPC setup.

Taro: If we look at the cost terms they defined in Table I, specifically the tracking term c track, it penalizes errors for root position, link position, and link orientation, which is good for kinematic accuracy in dynamic motions. What about the regularization terms that control torque energy and action rate?

Rosa: The paper includes a regularization term c reg which penalizes joint torque energy and action rate; they set weights of zero point zero five for those terms. This helps keep the generated motions physically plausible, ensuring the robot doesn't require impossible forces to execute the retargeted trajectory.

Dev: And I see their control input d phi k is directly linked to the deformation rate r k, which is defined as two(d phi k/d phi src), where a negative rate means slowing down and a positive rate means speeding up. That linkage between the control input and the timing adjustment is key to their method.

Paper summary: Taro: That direct control over the deformation rate sounds like it gives us some agency in shaping *how* the robot adapts its timing, rather than just passively following a pre-set schedule derived from human motion. What happens when we push this concept further into scenarios where the source motion itself is highly ambiguous or poorly defined?

Rosa: The paper seems to focus on motions that are already dynamically feasible, and they mention that policies trained on these references transfer well to a real humanoid robot. It suggests that the success hinges heavily on the quality of the reference data used for training.

Dev: So, to summarize what we've heard about "Dense Temporal Motion Retargeting for Legged Robots," it’s a method that jointly optimizes timing and control in a dense formulation to precisely adapt source motions to robot dynamics, showing performance gains over baselines while maintaining speed.

Taro: I think the implication here is that we might move closer to systems that can generalize motion skills across different robot morphologies, not just specific ones, because the method focuses on the underlying dynamic properties rather than just pose matching.

Rosa: That sounds like a big step toward truly adaptable robotic skills. So, moving into the conclusion of this discussion for now, what are your thoughts on where this research is headed?

Dev: I see the implication as making imitation learning more robust to physical constraints, which is critical for deploying robots in unstructured environments where dynamics are constantly changing.

Taro: I think we're looking at a path where autonomy systems can adapt their movement strategies on the fly based on real-time physical feedback and environmental cues, rather than just executing pre-programmed sequences.

Rosa: We've covered the core concept of DTMR and its performance metrics in relation to baselines, looking at how it handles timing versus control constraints.

Dev: I'm still thinking about the computational demands; solving this with parallel GPU MPC is impressive, but scaling this up for complex, long-horizon tasks remains a practical hurdle we need to address.

Taro: It seems like the future involves integrating these types of dynamic motion adaptation techniques directly into the perception and planning layers of autonomous systems, giving them a richer understanding of physical interaction.

Rosa: And that's where we wrap up our discussion on this paper; "Dense Temporal Motion Retargeting for Legged Robots" is clearly pushing the boundaries of how we translate human expertise into robotic action, and its ability to handle dynamic movements efficiently is certainly worth sharing with listeners.

Conclusion: Rosa: So, we've been digging into Dense Temporal Motion Retargeting for Legged Robots, and now it’s time to wrap up by looking at what this paper is actually about and why it matters to us as field roboticists.

Dev: I think the title itself tells a lot; "Dense Temporal Motion Retargeting" suggests they aren't just matching poses, but they’re getting down into the timing adjustments at every single step of the control sequence.

Taro: I agree with Dev; that density is what lets them handle those dynamic movements much better than traditional methods that treat timing as a separate, fixed parameter.

Rosa: And looking at the authors, we see a team focused on bridging the gap between observed human motion and real robot dynamics, which is exactly where we need to be if we want robots to perform complex tasks in messy environments.

Dev: They’re clearly targeting the control side of things because they are optimizing timing directly within the control loop via that phase variable, which makes sense for a systems engineer concerned with latency and loop rates.

Taro: And from an autonomy angle, this implies a system that can adapt its entire movement strategy in real-time based on how the robot is physically reacting to its environment, not just following a pre-recorded script.

Rosa: Exactly; if we can get robots to learn these fine temporal adjustments, it opens up possibilities for them to handle unpredictable physical interactions with far more grace than current systems allow.

Dev: I wonder about deployment outside of the lab; Rosa mentioned that they did train policies on humanoid robots, but how long can we expect this level of precision to hold up when things get really rough and unexpected?

Taro: That’s a fair question for field work; the success seems heavily tied to the quality and feasibility of those initial reference motions they used for training.

Rosa: The paper shows promise in that transferability, but we still need more data on how robust these retargeted policies are when faced with novel physical disturbances outside of a controlled setting.

Dev: We need to keep an eye on the computational cost too; running this kind of dense optimization in real-time means we have to be very careful about the processing power needed for that parallel GPU setup.

Taro: That computational aspect is something we can definitely push further, but I think the fundamental impact here is moving imitation learning closer to true physical generalization.

Rosa: It really is; this research suggests a path where robots can learn motion not just as a sequence of positions, but as a continuous dance of timing and control that respects their own physical limitations.

Dev: So, we’re looking at systems that are smarter about *when* to move, which directly addresses some of the biggest hurdles in making real-world robotics functional.

More episodes

← Home