PROMO: Preference-conditioned Multi-Objective Reinforcement Learning for Quadrupedal Robots
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.
Dev: Today's paper: "PROMO: Preference-conditioned Multi-Objective Reinforcement Learning for Quadrupedal Robots".
Rosa: PROMO introduces Preference-Conditioned Multi-Objective Reinforcement Learning (MORL) for quadrupedal robots, addressing the limitation of fixed scalar rewards by allowing operator intent to explicitly condition locomotion trade-offs at runtime.
Dev: First, who's behind it and why it matters.
Paper summary: Dev: So, wrapping up on this paper, "PROMO: Preference-conditioned Multi-Objective Reinforcement Learning for Quadrupedal Robots," it seems the core contribution is successfully formalizing control by separating task objectives from fixed embodiment requirements through semantic objective/prior factorization.
Rosa: That separation is what allows a single policy to span diverse behaviors by letting the operator specify which trade-off—tracking, stability, or efficiency—the robot should prioritize at runtime. It’s about making those deployment-facing preferences an explicit input to the policy instead of burying them deep inside a reward function.
Taro: I think the real implication here is that it gives us a way to manage flexibility without needing different controllers for every operational mode; we can command a continuum of desired behaviors through this preference vector.
Dev: From an engineering standpoint, it’s about achieving objective specialization and robustness from just one deployable policy, which simplifies the hardware and software architecture significantly compared to having multiple specialized controllers ready to switch between.
Rosa: It really suggests that the interface is interpretable because we can specify *why* behavior should change—for example, by prioritizing stability over tracking—and the policy learns how to execute that specific trade-off effectively.
Taro: If this holds up under rigorous testing in real-world scenarios, it means autonomous systems will be much more adaptable and less brittle when faced with unpredictable conditions than what we see in current fixed-priority control methods.
Dev: I just hope the implementation details on the loop rate and latency hold up when we move this from simulation to actual hardware deployment, because those real-time constraints are always where these kinds of conditioning mechanisms can introduce failure modes if they aren't tuned properly.
Rosa: We saw that in simulation, for instance, changing just the preference alone reduced specific energy by up to thirty point four percent and position error by thirty-eight point seven percent on the Unitree Go2 hardware, which shows how impactful this interface can be on physical performance right away.
Taro: That measurable impact under those conditions really validates the idea that operator intent can systematically modulate hardware behavior with a fixed policy, proving that this is a meaningful control interface for autonomous systems.
Conclusion: Rosa: It seems like the title itself pretty much sums up what this paper is doing: giving a quadruped robot a way to listen to your goals and adjust its movement priorities dynamically while still operating within its physical limits.
Dev: Yeah, I think the authors are really smart for tackling that trade-off between needing fast computation and needing accurate, real-time decision-making; it's a tough spot for control engineers.
Taro: From an autonomy side, the paper suggests we move toward systems where the robot doesn't just follow a pre-programmed path but actively chooses its operational strategy based on the immediate environment or user command.
Rosa: Exactly, and I'm really curious about the long-term impact; if this works reliably outside of a controlled lab setting, how quickly could we see these robots deployed in genuinely unpredictable environments?
Dev: That’s the big question for me—the deployment longevity. We need to know if this preference interface survives real-world noise and unexpected sensor glitches without causing catastrophic failure modes at the loop rate required for locomotion.
Taro: If the system can handle misbehavior in the world, like sudden obstacles or unexpected terrain changes, then we could envision robots that adapt their movement strategy instantly instead of just crashing or freezing.
Rosa: It feels like this research opens up a whole new category of robot control where the operational goal isn't just "walk here," but rather "walk here *in a stable and efficient way*."
Dev: And from an engineering standpoint, the authors' focus on factorization between the semantic objectives and the fixed locomotion priors is really telling; it seems like they found a way to keep things predictable while still allowing that flexibility.
Taro: That factorization is what makes it promising because it ensures that even when we shift priorities, there's still a solid foundation of embodied physics guiding the decisions.
Rosa: It really shifts the focus from designing one perfect controller for every single scenario to designing one flexible controller that can handle a wide range of mission types.
Dev: And if we look at the results, it shows that these preference changes actually translate into measurable physical improvements in terms of energy use and error reduction on hardware like the Unitree Go2.
Taro: That tangible performance data is what will really convince other researchers that this isn't just theoretical work but something with real utility for complex autonomous missions.
Rosa: So, we've seen the mechanics of how it works, now we need to think about where this technology actually lands in the practical deployment pipeline.
Dev: Right, and that leads perfectly into our next topic: we need to look at how robust this entire framework is when faced with the messy realities of real-world operation.
Amr Mousa, Rifny Rachman, Neil Karavis, Michele Caprio, Richard Allmendinger
University of Manchester
cs.RO, cs.AI, cs.HC, cs.LG, cs.SY, eess.SY
Submitted: 2026-10-01
Updated: 2026-10-01
Comments: Submitted to IEEE Transactions on Robotics. Project website, code, and videos: https://amrmousa.com/promo/
Project page: https://amrmousa.com/promo
License: http://creativecommons.org/licenses/by-sa/4.0/
Importance score: 86/100
The gist: PROMO introduces Preference-Conditioned Multi-Objective Reinforcement Learning (MORL) for quadrupedal robots, addressing the limitation of fixed scalar rewards by allowing operator intent to
Key concepts
- Preference-Conditioned MORL
- This framework conditions a single reinforcement learning policy on a preference vector that defines trade-offs between multiple objectives (tracking, stability, efficiency). Instead of having different policies for each goal, one policy learns how to adjust its actions based on the desired balance specified by the operator.
- Semantic Reward Vector
- This is a set of objectives—like tracking or stability—represented as a vector. The actual utility received by the robot is calculated by taking the dot product of this objective vector and a preference weight vector provided at runtime, allowing dynamic weighting of goals.
- Objective/Prior Factorization
- This design separates the desired high-level semantic objectives from fixed, embodiment-specific locomotion priors (like contact timing). This factorization is crucial because it keeps the core physical requirements internal to the controller, making the preference vector a meaningful interface for changing behavior.
- Action Diversity Regularizer
- A small regularization term added to the actor's objective ensures that even when observations are similar, different preference settings lead to distinct policies. This prevents the policy from converging on a single behavior regardless of the input weights.
Terminology
Summary
PROMO introduces Preference-Conditioned Multi-Objective Reinforcement Learning (MORL) for quadrupedal robots, addressing the limitation of fixed scalar rewards by allowing operator intent to explicitly condition locomotion trade-offs at runtime. This framework matters because it enables a single policy to adapt its behavior dynamically based on deployment preferences—such as prioritizing tracking, stability, or efficiency—without requiring retraining or switching between specialized controllers.
The gist: A single policy is conditioned on preferences over three deployment-facing objectives—tracking, stability, and efficiency—while embodiment-specific locomotion priors remain fixed.
How it works
PROMO frames locomotion as a partially observed multiobjective decision process by separating semantic objectives from fixed locomotion priors. The core mechanism involves conditioning the policy on the preference vector, denoted as the runtime input to a single policy: A single policy is conditioned on preferences over three deployment-facing objectives—tracking, stability, and efficiency—while embodiment-specific locomotion priors remain fixed.
This separation keeps embodiment-specific requirements internal to the controller and makes the preference vector a meaningful runtime interface rather than a reformulation of the full reward function.
The framework utilizes a semantic reward vector where the semantic objective index set is I = [track, stab, eff], with K = I = 3.
The scalar semantic utility is defined as U(r sem t, w t) = w T t r sem t,
where the preference vector defines the weights applied to these objectives. During training, this utility is augmented by a fixed locomotion prior term: During training, semantic utility is augmented by a fixed locomotion prior, u loc(w t) = w T t r sem t + λpriorr prior t,
where λprior is a fixed coefficient.
Architecture and Factorization
PROMO extends the TeacherAligned Representations (TAR) architecture to incorporate preference conditioning through specialized encoders. The architecture features distinct latent representations: zP t = EP(s priv t, w t), zH t = EH(ot-H:t, w t),
where both encoders are preference-conditioned.
During training, the critic receives inputs including the preference-conditioned privileged latent zP t. The actor receives inputs that include the history latent and velocity estimator output: The actor receives (ot, zH t, v hat b t, w t) and outputs 12 joint-position targets.
A critical design choice is the Objective/prior factorization
to address the mismatch between vector-valued MORL objectives and dense reward design. This involves: separating operator-facing semantic objectives from fixed embodiment-level locomotion priors.
The fixed prior terms, such as contact timing and posture constraints, are represented by r prior t and are added to the actor's utility but do not define critic heads. This factorization is deemed critical for controllability, as the paper notes that the fixed-prior formulation attains a 6% fall rate under challenging evaluation conditions,
whereas alternative designs can increase it significantly.
Optimization and Deployment
The training objective for preference wt is defined as: Jtrain(θ, w t) = w T t Jsem(πθ, w t) + λpriorJprior(πθ, w t),
where λprior = 1 in all experiments. The actor optimization uses a scalarized reward: For actor optimization, the semantic reward is scalarized as r sem t(w) = w T t r sem t.
Furthermore, the policy includes a small action-diversity regularizer to ensure preference responsiveness: The resulting actor objective is Lactor = LPPO + λdivLact div,
which penalizes differences between policies conditioned on similar observations but different preferences.
Evaluation and Results
PROMO is evaluated across simulation and real-robot deployment using four main evaluation criteria: training performance, common semantic evaluation, preference-interface structure, and hardware deployment. In simulation, the dense 100-preference study shows that 67 behaviors are non-dominated under exact Pareto dominance,
with a mean preference–objective correlation of 0.843.
On the Unitree Go2 hardware, changing only the preference produces significant changes: preference changes alone reduce specific energy by up to 30.4%, position error by 38.7%, and peak body-attitude deviation by 59.0% relative to the balanced preference.
The results establish that preference changes systematically modulate hardware behavior with a fixed policy,
demonstrating that the interface is interpretable
as it allows operators to specify why behavior should change—for example, by prioritizing stability over tracking—and the policy learns how to realize that trade-off.
Core Contributions
The paper outlines four core contributions.
Improvements for AI systems
Here are specific improvements to AI systems derived from the PROMO framework, along with what those improved systems can achieve:
)1. Implementation of a Semantic Preference Interface for Adaptive Control:
Instead of hard-coding fixed reward weights or relying on complex, opaque latent gait codes, the system will be conditioned on an explicit semantic preference vector (tracking, stability, efficiency). This allows for runtime switching of locomotion behavior by simply inputting a preference weight vector.
- Decoupling Task Objectives from Embodiment Priors:
The architecture will maintain a fixed set of embodiment-specific locomotion priors
(e.g., contact timing, posture constraints) that are optimized during training but remain constant during deployment. The semantic objectives (tracking, stability, efficiency) will be treated as the sole runtime control inputs. This prevents low-level shaping terms from becoming artificial objective dimensions and ensures that operator intent maps directly to high-level behavioral goals.
- Robust Sim-to-Real Transfer via Preference Modulation:
The system can be deployed on physical hardware (e.g., Unitree Go2) without retraining or controller switching. The learned policy is inherently robust because it has been trained across a dense set of 100 preferences in simulation, and the hardware performance is shown to scale predictably based on the preference input alone (e.g., reducing specific energy by up to 30.4% when prioritizing efficiency).
- Multi-Objective Specialization from a Single Policy:
The system can achieve objective specialization
from one deployable policy. Rather than maintaining three separate controllers (one for tracking, one for stability, etc.), the single policy learns the ability to realize any desired trade-off on demand by modulating its internal objective weights.
- Interpretable Control and Behavioral Intent:
The preference vector acts as an interpretable representation of behavioral intent. An operator can explicitly specify prioritize stability over tracking,
and the policy is trained to realize that trade-off, rather than manipulating a latent gait code or reward function coefficients. This provides a transparent interface for mission planning and human oversight.
)What the Improved AI System Can Do:
The resulting AI system will be capable of performing complex quadrupedal tasks in dynamic environments where the required locomotion strategy shifts based on mission needs, terrain changes, or operator demands. Specifically:
-
Maximize Mission Flexibility: The robot can seamlessly transition between high-speed tracking (for following a fast target), stable maneuvering (for navigating rough terrain or recovering from disturbances), and energy-efficient cruising (for long-duration patrols) simply by changing the semantic preference input.
-
Achieve Optimized Performance on Demand: It can execute a
best-of
gait for any given situation—such as maximizing tracking accuracy while maintaining acceptable stability, or minimizing energy expenditure while ensuring the robot remains upright—without needing to be pre-trained for those specific conditions. -
Ensure Robust Deployment: The robot can be deployed rapidly in the field, and its performance will be predictably modulated by the preference input, offering high reliability that surpasses fixed-objective controllers and specialized experts trained offline on only one objective.
-
Enable Semantic Command Execution: It moves beyond simple trajectory following to executing
semantic commands
(e.g.,Be stable,
Be efficient
), allowing for higher-level task execution in autonomous agents where the desired outcome is defined by a trade-off, not a fixed trajectory.
Abstract
Quadrupedal locomotion requires balancing conflicting objectives such as command tracking, stability, and energy efficiency, yet conventional reinforcement learning (RL) hardcodes these priorities into a fixed scalar reward at training time. We present PROMO (Preference-Conditioned Multi-Objective Reinforcement Learning), a semantic multi-objective approach that makes this trade-off an explicit runtime input to a single locomotion policy. PROMO conditions the policy on deployment facing preferences while keeping embodiment-specific locomotion priors fixed, thereby separating operator intent from reward shaping terms required for viable gait generation. Compared with fixed-objective controllers, multi-objective baselines, and independently trained specialists, PROMO achieves objective specialization and robustness from a single deployable policy. Across 100 sampled preferences in simulation, 67 behaviors are non-dominated under exact Pareto dominance, with a mean preference-objective correlation of 0.843, demonstrating broad Pareto coverage and predictable preference response. The same policy transfers zero-shot to a Unitree Go2, where preference changes alone reduce specific energy by up to 30.4%, position error by 38.7%, and peak body-attitude deviation by 59.0% relative to the balanced preference. These results establish preference-conditioned multi-objective RL as a practical runtime interface for adaptive legged locomotion, extending its role beyond offline Pareto-set construction. Open-source code and videos are available at https://amrmousa.com/promo/.
Sources
- Learning Multiple Gaits within Latent Space for Quadruped Robots
- Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization
- Reward-Conditioned Policies
- In Search for Architectures and Loss Functions in Multi-Objective Reinforcement Learning
- Scalable Multi-Objective Robot Reinforcement Learning through Gradient Conflict Resolution
- Controllability in preference-conditioned multi-objective reinforcement learning
- GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving