PhysMoDPO: Physically-Plausible Humanoid Motion with Preference Optimization

arXiv:2603.13228 · cs.LG, cs.AI, cs.CV, cs.RO · Submitted 2026-03-13 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "PhysMoDPO: Physically-Plausible Humanoid Motion with Preference Optimization".

Tom: The gist:

Jane: First, who's behind it and why it matters.

Paper summary: Tom: We've been talking about PhysMoDPO, and to recap its main thesis: diffusion models generate motions that look good based on text instructions but often fail when put through a Whole-Body Controller because those controllers fix the motion in ways that destroy the original intent.

Jane: The paper proposes PhysMoDPO as a Direct Preference Optimization framework to solve this discrepancy. They integrate the Whole-Body Controller into the training pipeline itself to optimize the motion generator directly against both physics and task requirements at once.

Lu: The core claim is that by using this approach, they can produce motions that remain stable and physically realistic when deployed on robots like the Unitree G1, which is a kinematic space evaluation benchmark they compared against prior methods.

Meng: They are essentially targeting the evaluation mismatch directly by computing both physics-based rewards for trackability and task-specific rewards to measure condition faithfulness at the same time.

Tom: That’s right. The optimization objective is formulated as L=L DPO(X win,X lose)+λSFT L SFT(X win), which progressively refreshes the preference data toward the model’s current transfer failure modes under a certain timeframe T.

Jane: It matters because it moves away from just using hand-crafted physics heuristics, like those penalties for foot sliding, and integrates the controller into the training process itself to learn what is truly feasible.

Lalam: This level of integration into the training loop shows that we can teach motion generators not just how to follow a command, but how to follow a command *while respecting* the underlying physical laws of movement.

Tom: And they showed consistent gains in both physical realism and task-related metrics across text-to-motion and spatial control tasks when testing on simulated robots.

Jane: But the real impact comes from their zero-shot generalization, meaning they can deploy these optimized motions directly to a real robot without needing extra motion refinement steps beforehand.

Meng: So, for someone focused on practical application, this means the system is designed to be ready for deployment out of the box on hardware like a humanoid robot.

Lu: The authors show that post-training a generator with these physics-guided preferences can produce motions that transfer beyond just kinematic benchmarks, which is what makes it useful in robotics.

Conclusion: Tom: So we wrap up with the title PhysMoDPO, by Yangsong Zhang, Anujith Muraleedharan, Rikhat Akizhanov, Abdul Ahad Butt, Gül Varol, Pascal Fua, Fabio Pizzati—it really summarizes the whole approach: making humanoid motion physically plausible through preference optimization.

Jane: In simple terms for someone listening who isn't deep in diffusion models or robotics: this paper shows how to train an AI generator so that when you use it to drive a robot, the resulting movement is not just aesthetically pleasing but also actually safe and executable by the hardware.

Lu: The implication here is that we are moving toward motion generation that inherently respects dynamics and contacts, rather than having to patch up physical impossibilities after the fact with external controllers.

Meng: For an engineer looking at this, it means less time spent debugging execution failures on the robot side because the generator has already been trained to produce trajectories that respect those constraints.

Tom: It suggests a future where generating complex human motion for robotics becomes more reliable right out of the gate, moving beyond just matching data points to actually creating functional physical actions.

Jane: The paper opens up a path where we can rely on AI-generated motions for tasks that require high degrees of physical coordination, because the system is optimized for real-world execution.

Lalam: From a cultural view, this pushes the boundary on what we can expect from generative AI in embodied systems; it validates the idea that grounding generation in physical constraints leads to more robust and useful outputs.

Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) · LIGM, École des Ponts, IP Paris, Univ Gustave Eiffel, CNRS · École Polytechnique Fédérale de Lausanne (EPFL)

cs.LG, cs.AI, cs.CV, cs.RO

Submitted: 2026-03-13

Updated: 2026-10-08

Comments: Project page: https://mael-zys.github.io/PhysMoDPO/

Project page: https://mael-zys.github.io/PhysMoDPO

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 92/100

The gist: The gist: PhysMoDPO proposes a Direct Preference Optimization framework that integrates Whole-Body Control into the training pipeline to optimize diffusion motion generators such that their outputs

Key concepts

Whole-Body Control (WBC)
WBC is a mechanism used to convert motion generated by diffusion models into executable trajectories for robots. It ensures that the proposed motions adhere to physical constraints, making them suitable for real-world application on hardware like humanoid robots.
Direct Preference Optimization (DPO)
DPO is a training technique used to fine-tune generative models based on human preferences. Instead of traditional reinforcement learning, DPO directly optimizes the model's output by comparing preferred and dispreferred motion samples, guiding the generator toward better results.
Composite Preference Reward R(X',C)
This reward is a multi-faceted signal used to judge motion quality. It combines several components: tracking rewards measure how closely the realized motion matches the original input, while sliding rewards penalize undesirable foot micro-sliding, ensuring both fidelity and physical realism.
Physics-Grounded Rewards
These are rewards calculated based on established physics principles applied to the generated motion. They specifically assess aspects like trackability and contact realism, providing a strong signal that the generated motions behave realistically in a physical environment.

Terminology

Summary

The gist: PhysMoDPO proposes a Direct Preference Optimization framework that integrates Whole-Body Control into the training pipeline to optimize diffusion motion generators such that their outputs are compliant with both physics and original text instructions.

Introduction and Problem

Recent progress in text-conditioned human motion generation is largely driven by diffusion models trained on large-scale human motion data. Building on this progress, recent methods attempt to transfer such models for character animation and real robot control by applying a Whole-Body Controller (WBC) that converts diffusion-generated motions into executable trajectories. While WBC trajectories become compliant with physics, they may expose substantial deviations from original motion. This introduces a discrepancy between what the generator produces and what is actually realized in the simulation.

PhysMoDPO Framework

The core idea of PhysMoDPO is to post-train a motion diffusion model with a physics-grounded reward signal. The procedure involves an extensive generation–finetuning loop. For each trajectory, the framework computes physics-based rewards and task-specific rewards and uses them to construct preference pairs for DPO finetuning. The training objective is formulated as L=L DPO(X win,X lose)+λSFT L SFT(X win). This procedure is iterated to progressively refresh the preference data toward the model’s current transfer failure modes under T.

Reward Construction

The composite preference reward R(X',C) is defined implicitly by declaring that a realized motion X'k is preferred to X'l under the same condition C if it improves every reward term. The set of rewards S(C) includes tracking reward Rtrack, sliding reward Rslide, and task rewards. The tracking reward Rtrack explicitly minimizes the difference between X and X'. The sliding reward Rslide penalizes foot micro-sliding. Task rewards include the text adherence reward R M2T and spatial control reward R control. The method demonstrates zero-shot generalization to simulation and real-world deployment on a G1 humanoid robot. In the zero-shot transfer to Unitree G1 robot evaluation, PhysMoDPO achieves the best performance across spatial controllability metrics. Experiments on text-to-motion and spatial control tasks demonstrate consistent gains of PhysMoDPO in physical realism and task metrics in simulation. Moreover, zero-shot transfer to Unitree G1 robot indicates the potential of PhysMoDPO for robotics-oriented motion generation. The limitation noted is that the current setup primarily considers locomotion on flat ground. Future work could incorporate human-validated models to reduce evaluator bias.

How it works

  1. Data Generation and Tracking: Given a conditioning signal (text and optional joint controls), the framework samples multiple motions X from a pretrained generator. A fixed tracking policy then projects each sample into a simulated trajectory X'.

  2. Reward Calculation: Physics-based rewards measure trackability and contact realism, while task-specific rewards measure whether the tracked motion still matches the input condition. The composite preference R(X',C) is defined implicitly by declaring that a realized motion X'k is preferred to X'l under the same condition C if it improves every reward term.

  3. Preference Optimization: The generator G is finetuned using Direct Preference Optimization (DPO).

Improvements for AI systems

  1. Bold header: Post-training for Robotics Plausibility

The improved system can generate motions that remain stable and physically realistic when deployed on the Unitree G1 robot, moving beyond kinematic feasibility to ensure dynamics are respected, as shown by the results demonstrating zero-shot generalization to simulation and real-world deployment.

  1. Bold header: Physics-Guided Preference Optimization (PhysMoDPO)

The system will optimize diffusion models using a pipeline that integrates WBC into our training pipeline to measure how close the motion is to transfer to executable trajectory, directly targeting the evaluation mismatch by optimizing for motions that remain both physically feasible and condition-faithful.

  1. Bold header: Multi-Objective Reward Construction

The system can construct preference pairs based on a composite reward structure, explicitly combining physics rewards (like minimizing tracking distortion) with task rewards such as Rslide to penalize foot micro-sliding, ensuring that the optimization is guided by both physical realism and fidelity to text instructions.

  1. Bold header: Zero-Shot Embodied Transfer

The improved system can perform zero-shot motion transfer to novel robot embodiments, specifically demonstrating the ability to move generated motions from SMPL representations to a G1 humanoid robot without additional motion refinement, proving the effectiveness of post-training in real-world deployment.

  1. Bold header: Robust Preference Data Selection

The preference construction mechanism employs a strict dominance-based selection, which ensures that the winning sample achieves a better R than Xlose across all reward terms, avoiding the sensitivity of score fusion by normalizing rewards and using weighted summation and leading to more stable training.

  1. Bold header: Hyperparameter Tuning for Stability

The system is configured with a default Diffusion-DPO temperature of β=20 to balance preference learning against distributional quality, as performance degrades significantly when β becomes too large (e.g., β=50), ensuring overly aggressive preference updates can harm distributional quality and motion smoothness.

  1. Bold header: Data Scale and Representation Scaling

The system utilizes an ablation study showing that increasing the data scale up to 100% consistently improves generation quality (FID) and text-motion alignment, confirming that the preference construction is sample-efficient even with smaller scales, while using SMPL representation for consistency in downstream tasks.

  1. Bold header: Controllability Enhancement via SFT

The system incorporates a supervised fine-tuning term on winning samples only, denoted LSFT(Xwin), which is shown to consistently improve spatial controllability, text consistency, as well as FID and jerk, providing a necessary regularization step to preserve generative capabilities.

Abstract

Recent progress in text-conditioned human motion generation has been largely driven by diffusion models trained on large-scale human motion data. Building on this progress, recent methods attempt to transfer such models for character animation and real robot control by applying a Whole-Body Controller (WBC) that converts diffusion-generated motions into executable trajectories. While WBC trajectories become compliant with physics, they may expose substantial deviations from original motion. To address this issue, we here propose PhysMoDPO, a Direct Preference Optimization framework. Unlike prior work that relies on hand-crafted physics-aware heuristics such as foot-sliding penalties, we integrate WBC into our training pipeline and optimize diffusion model such that the output of WBC becomes compliant both with physics and original text instructions. To train PhysMoDPO we deploy physics-based and task-specific rewards and use them to assign preference to synthesized trajectories. Our extensive experiments on text-to-motion and spatial control tasks demonstrate consistent improvements of PhysMoDPO in both physical realism and task-related metrics on simulated robots. Moreover, we demonstrate that PhysMoDPO results in significant improvements when applied to zero-shot motion transfer in simulation and for real-world deployment on a G1 humanoid robot.

Sources

Related papers