Dense Temporal Motion Retargeting for Legged Robots
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "Dense Temporal Motion Retargeting for Legged Robots".
Dev: Legged robots can learn expressive whole-body skills from human and animal motions, but adapting these motions to a robot's specific dynamics requires careful adjustment of timing and control,
Rosa: First, who's behind it and why it matters.
Paper summary: Rosa: So, we're talking about this paper by Jaeryeong Kim et al., "Dense Temporal Motion Retargeting for Legged Robots," which is really focusing on how robots can learn complex movements from video. Rosa here, I want to start with the main idea—what’s the core thesis of this work?
Dev: Exactly, Rosa; basically, the authors are tackling the problem that when you try to teach a legged robot a human jump or some dynamic action using recorded motions, you have this big gap between what the source motion looks like and what your robot can physically do. So, they propose dense temporal motion retargeting as a way to bridge that gap by jointly optimizing timing and control in one single optimization process.
Taro: From an autonomy standpoint, I’m interested in how this dense formulation handles the interdependence of timing and control when dealing with dynamic movements like jumps or sudden changes in speed. If the robot is trying to mimic a rapid deceleration, how does this joint optimization help it manage those interdependent variables effectively?
Rosa: That's a good place to start, Taro; the paper claims that this dense formulation allows them to deform only the parts of the motion that actually need a change in timing, which they describe as adjusting the timing at every control step. This is presented as a way to maintain efficiency while achieving precise retargeting.
Dev: And from an engineering loop rate perspective, Rosa, that dense formulation means the timing adjustment—what they call the phase trajectory phi —is not just a separate calculation but is generated directly by a control input d phi k, which helps them manage those dynamics more tightly within the optimization framework.
Taro: I wonder about the practical application when things go wrong; if the world misbehaves, does this DTMR system have any inherent mechanism to handle unexpected external disturbances while still trying to follow the retargeted motion?
Rosa: The paper suggests that by allowing more temporal deformation, they found it leads to more precise retargeting, which means the robot can reproduce those dynamic motions with greater accuracy than before. This is a significant finding because it shows how fine-grained control over timing affects the final output of the imitation learning process.
Dev: That precision comes at a cost, and I want to talk about that trade-off; the authors show that precision and timing preservation are controlled by a phase-cost weight w phi, where a large w phi keeps the timing close to the source motion, while a small one permits more deformation.
Paper summary: Taro: So, if we're pushing for high fidelity in dynamic tasks, like hopping or landing gracefully, should we be favoring that larger timing constraint or are there situations where allowing that greater temporal deformation is beneficial?
Rosa: The paper points out that allowing more temporal deformation actually yields more precise retargeting when compared to baseline methods under the same deformation budget. This suggests we might need to tune that weight based on whether we prioritize strict timing adherence or maximizing kinematic accuracy.
Dev: And looking at the efficiency, they achieved results where DTMR outperformed baseline temporal optimization by about nineteen times faster when using a similar deformation budget. That speedup is pretty substantial for real-time applications on hardware.
Taro: A nineteen times speedup is notable, but I'm curious about the real-world deployment; Rosa, how long do you think this kind of retargeting system can reliably operate outside of a controlled lab environment before we start seeing significant degradation in its performance?
Rosa: The paper does show that policies trained using these references transfer to a real humanoid robot, and they constructed a two-hour dataset of dynamically feasible motions on which policies learn dynamic motions better. This indicates promise for real-world applicability, but the authors are focused on the transferability to those specific embodiments.
Dev: From my side, I'm concerned about the implementation details; they solve this using sampling-based model predictive control in parallel on a GPU, which is computationally intensive, so we have to consider that loop rate and any potential failure modes in that MPC setup.
Taro: If we look at the cost terms they defined in Table I, specifically the tracking term c track, it penalizes errors for root position, link position, and link orientation, which is good for kinematic accuracy in dynamic motions. What about the regularization terms that control torque energy and action rate?
Rosa: The paper includes a regularization term c reg which penalizes joint torque energy and action rate; they set weights of zero point zero five for those terms. This helps keep the generated motions physically plausible, ensuring the robot doesn't require impossible forces to execute the retargeted trajectory.
Dev: And I see their control input d phi k is directly linked to the deformation rate r k, which is defined as two(d phi k/d phi src), where a negative rate means slowing down and a positive rate means speeding up. That linkage between the control input and the timing adjustment is key to their method.
Paper summary: Taro: That direct control over the deformation rate sounds like it gives us some agency in shaping *how* the robot adapts its timing, rather than just passively following a pre-set schedule derived from human motion. What happens when we push this concept further into scenarios where the source motion itself is highly ambiguous or poorly defined?
Rosa: The paper seems to focus on motions that are already dynamically feasible, and they mention that policies trained on these references transfer well to a real humanoid robot. It suggests that the success hinges heavily on the quality of the reference data used for training.
Dev: So, to summarize what we've heard about "Dense Temporal Motion Retargeting for Legged Robots," it’s a method that jointly optimizes timing and control in a dense formulation to precisely adapt source motions to robot dynamics, showing performance gains over baselines while maintaining speed.
Taro: I think the implication here is that we might move closer to systems that can generalize motion skills across different robot morphologies, not just specific ones, because the method focuses on the underlying dynamic properties rather than just pose matching.
Rosa: That sounds like a big step toward truly adaptable robotic skills. So, moving into the conclusion of this discussion for now, what are your thoughts on where this research is headed?
Dev: I see the implication as making imitation learning more robust to physical constraints, which is critical for deploying robots in unstructured environments where dynamics are constantly changing.
Taro: I think we're looking at a path where autonomy systems can adapt their movement strategies on the fly based on real-time physical feedback and environmental cues, rather than just executing pre-programmed sequences.
Rosa: We've covered the core concept of DTMR and its performance metrics in relation to baselines, looking at how it handles timing versus control constraints.
Dev: I'm still thinking about the computational demands; solving this with parallel GPU MPC is impressive, but scaling this up for complex, long-horizon tasks remains a practical hurdle we need to address.
Taro: It seems like the future involves integrating these types of dynamic motion adaptation techniques directly into the perception and planning layers of autonomous systems, giving them a richer understanding of physical interaction.
Rosa: And that's where we wrap up our discussion on this paper; "Dense Temporal Motion Retargeting for Legged Robots" is clearly pushing the boundaries of how we translate human expertise into robotic action, and its ability to handle dynamic movements efficiently is certainly worth sharing with listeners.
Conclusion: Rosa: So, we've been digging into Dense Temporal Motion Retargeting for Legged Robots, and now it’s time to wrap up by looking at what this paper is actually about and why it matters to us as field roboticists.
Dev: I think the title itself tells a lot; "Dense Temporal Motion Retargeting" suggests they aren't just matching poses, but they’re getting down into the timing adjustments at every single step of the control sequence.
Taro: I agree with Dev; that density is what lets them handle those dynamic movements much better than traditional methods that treat timing as a separate, fixed parameter.
Rosa: And looking at the authors, we see a team focused on bridging the gap between observed human motion and real robot dynamics, which is exactly where we need to be if we want robots to perform complex tasks in messy environments.
Dev: They’re clearly targeting the control side of things because they are optimizing timing directly within the control loop via that phase variable, which makes sense for a systems engineer concerned with latency and loop rates.
Taro: And from an autonomy angle, this implies a system that can adapt its entire movement strategy in real-time based on how the robot is physically reacting to its environment, not just following a pre-recorded script.
Rosa: Exactly; if we can get robots to learn these fine temporal adjustments, it opens up possibilities for them to handle unpredictable physical interactions with far more grace than current systems allow.
Dev: I wonder about deployment outside of the lab; Rosa mentioned that they did train policies on humanoid robots, but how long can we expect this level of precision to hold up when things get really rough and unexpected?
Taro: That’s a fair question for field work; the success seems heavily tied to the quality and feasibility of those initial reference motions they used for training.
Rosa: The paper shows promise in that transferability, but we still need more data on how robust these retargeted policies are when faced with novel physical disturbances outside of a controlled setting.
Dev: We need to keep an eye on the computational cost too; running this kind of dense optimization in real-time means we have to be very careful about the processing power needed for that parallel GPU setup.
Taro: That computational aspect is something we can definitely push further, but I think the fundamental impact here is moving imitation learning closer to true physical generalization.
Rosa: It really is; this research suggests a path where robots can learn motion not just as a sequence of positions, but as a continuous dance of timing and control that respects their own physical limitations.
Dev: So, we’re looking at systems that are smarter about *when* to move, which directly addresses some of the biggest hurdles in making real-world robotics functional.
Jaeryeong Kim, *Taerim Yoon*, *Jin Cheng*, *Sungjoon Choi*, *Stelian Coros*
Department of Computer Science, ETH Zurich · Department of Artificial Intelligence, Korea University
cs.RO
Submitted: 2026-09-29
Updated: 2026-09-29
Comments: 8 pages, 7 figures, 6 tables. Project page: https://jaeryeongnicolekim.com/Dense-Temporal-Motion-Retargeting-For-Legged-Robots/
Code: https://github.com/googledeepmind/mujoco
Project page: https://jaeryeongnicolekim.com/Dense-Temporal-Motion-Retargeting-For-Legged-Robots
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 86/100
The gist: Legged robots can learn expressive whole-body skills from human and animal motions, but adapting these motions to a robot's specific dynamics requires careful adjustment of timing and control,
Key concepts
- Dense Temporal Motion Retargeting (DTMR)
- This method jointly optimizes the robot's motor controls and the timing of those controls across every single control step. This dense approach means it can change the timing at every frame, allowing for highly precise adaptation of source motions, especially during fast or dynamic movements.
- Phase Variable ($\phi_k$)
- The phase variable is a decision variable that dictates how fast or slow the robot should execute a specific part of the motion compared to the original source. It allows the robot to stretch or compress time during execution, enabling it to match the timing characteristics of different reference motions.
- Deformation Rate ($r_k$)
- The deformation rate measures how much the timing is being changed at a specific step. A positive rate means the robot is speeding up its movement relative to the source, while a negative rate indicates it is slowing down, providing a mathematical way to quantify temporal adjustments.
Terminology
Summary
Legged robots can learn expressive whole-body skills from human and animal motions, but adapting these motions to a robot's specific dynamics requires careful adjustment of timing and control, especially for dynamic movements like jumps. This paper proposes Dense Temporal Motion Retargeting (DTMR), a method that jointly optimizes timing and control within a single optimization problem by adjusting the timing at every frame, enabling precise retargeting of dynamic motions while maintaining efficiency.
The gist
Dense temporal motion retargeting (DTMR) jointly optimizes timing and control within a single optimization, where dense means that the timing is adjusted for every control step. This dense formulation enables DTMR to deform only the parts of the motion that need a change in timing.
Problem Formulation
The core problem seeks motor controls under which a legged robot reproduces a source motion while allowing the timing to change. The formulation is described by minimizing a cost function subject to dynamic constraints:
min u,ϕ 1/2Tϕ X Tϕ k=0 FK(xk) − LI(ϕk; p¯) 2 Q s.t. xk+1 = f(xk, uk), ϕ0 = 0, ϕk ≤ ϕk+1, ϕTϕ = 1.
The phase variable, defined as a decision variable for the timing at each control step k, allows the robot to imitate the source motion faster or slower than the source itself. The phase trajectory is generated by the control sequence through a control input, where the deformation rate is defined as: the deformation rate rk = log2(dϕk/dϕsrc) is negative where the robot slows down and positive where it speeds up.
Methodology: Dense Temporal Motion Retargeting (DTMR)
The method defines a phase-augmented state, control, and dynamics: x˜k = (xk, ϕk), u˜k = (uk, dϕk), x˜k+1 = ˜f(x˜k, u˜k).
The optimization is formulated as:
min ũ 1 Tϕ X Tϕ k=0 c(x̃ k, ũ k) s.t. x̃ k+1 = ˜f(x̃ k, ũ k), where Tϕ = min(k: ϕk = 1).
The running cost is defined as: c(x˜k, u˜k) = ctrack + wdϕ cϕ + creg.
The tracking term, ctrack, penalizes errors between the robot and the reference keypoints at phase ϕk. The regularization term, creg, penalizes torque energy and action rate. Crucially, the phase increment dϕk is a control input; The phase trajectory ϕ of Eq. (2) is no longer a separate variable but is generated by the control sequence through dϕk.
Optimization and Efficiency
The problem is solved using sampling-based model predictive control (MPC) in parallel on a GPU. The algorithm involves iterative updates using an annealing schedule of DIAL-MPC [21]. Key steps include:
-
Sampling noise sequences [W˜]1:NW ∼ N (0, Σ˜i k:k+H).
-
Rolling out the augmented dynamics,
advancing ϕ before each physics step.
-
Keeping the elite fraction κ of rollouts with the lowest cost to update the plan using Eq. (1).
This approach allows for efficient parallelization on a GPU because the update is efficiently parallelized on a GPU.
The method utilizes Hnode control nodes
from which the control signal is generated by linear interpolation, reducing the dimension of sampled noise.
Performance and Trade-offs
The paper evaluates DTMR against baselines like STMR [8] and shows significant improvements.
DTMR retargets more precisely than baseline temporal optimization [8] under a similar deformation budget while being ∼19× faster.
The trade-off between precision and timing preservation is controlled by the phase-cost weight wdϕ. A large wdϕ keeps the timing close to the source motion, and a small one allows more deformation.
Allowing more temporal deformation improves precision: allowing more temporal deformation yields more precise retargeting.
Furthermore, DTMR is shown to be superior in dynamic motions: temporal optimization is especially effective on dynamic motions.
The method also demonstrates that policies trained on references transfer to a real humanoid robot,
and the wall-clock time shows significant speedup compared to STMR.
Limitations and Future Work
The paper notes several limitations. First, DTMR fits timing to dynamics, not meaning: "DTMR fits the timing to the dynamic properties of the robot, not to the meaning of the motion.
Improvements for AI systems
Here are the specific improvements that can be made to AI systems by implementing Dense Temporal Motion Retargeting (DTMR), along with what these improved systems will be able to do:
-
A robot system will be able to execute complex, dynamic human skills (like jumps, flips, or fast running) learned from source motions with high fidelity and physical feasibility.
-
The system will achieve superior performance in imitation learning tasks compared to existing motion retargeting methods (GMR, PHUMA, OmniRetarget) by generating motions that are demonstrably more trackable by downstream policies.
-
The robot will demonstrate a significant speedup in the motion retargeting process (up to 19x on G1 and 21x on Go1 compared to STMR), allowing for faster data generation and training cycles.
-
The system will be able to achieve higher precision in motion retargeting by effectively balancing timing preservation against kinematic accuracy, as controlled by the temporal deformation weight parameter (wdϕ).
-
The AI system will produce temporally optimized motions where only the necessary parts of a motion require timing adjustments, preserving the overall structure and meaning of non-dynamic segments.
-
The resulting policies trained on this high-quality, dynamically feasible motion dataset will transfer effectively to real-world humanoid robots, leading to better final tracking rewards in deployment scenarios (e.g., contact-rich dances or fast center-of-mass movements).
-
The system can handle the interdependence of control and timing simultaneously by solving them within a single optimization framework (MPPI), ensuring that the robot's control inputs are physically executable given the adapted timing.
Abstract
Legged robots can learn expressive whole-body skills from the motions of humans and animals. Due to the morphology gap between the source and the robot, however, the motion must be tailored to the dynamic properties of the robot. In particular, dynamic motions such as a jump require careful adjustment, since their timing and control are interdependent. We propose dense temporal motion retargeting (DTMR), which jointly optimizes timing and control within a single optimization, where dense means that the timing is adjusted for every control step. This dense formulation enables DTMR to deform only the parts of the motion that need a change in timing. The problem is solved with sampling-based model predictive control (MPC) in parallel on a GPU. We evaluate DTMR against baselines on two hours of human motion with four humanoid robots, where the results show that DTMR outperforms baseline methods, particularly on dynamic motions. We also show that allowing more temporal deformation yields more precise retargeting. We further compare DTMR with a baseline that optimizes the temporal dimension, where the result shows that DTMR retargets more precisely under the same deformation budget while being 19x faster. Lastly, policies trained on our references transfer to a real humanoid robot.
Sources
- TWIST: Teleoperated Whole-Body Imitation System
- SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control
- Retargeting Matters: General Motion Retargeting for Humanoid Motion Tracking
- PHUMA: Physically Reliable Humanoid Locomotion Dataset
- OmniRetarget: Interaction-Preserving Data Generation for Humanoid Whole-Body Loco-Manipulation and Scene Interaction
- GMT: General Motion Tracking for Humanoid Whole-Body Control
- Full-Order Sampling-Based MPC for Torque-Level Locomotion Control via Diffusion-Style Annealing
- SOMA: Unifying Parametric Human Body Models
- BeyondMimic: From Motion Tracking to Versatile Humanoid Control via Guided Diffusion
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving