PROMO: Preference-conditioned Multi-Objective Reinforcement Learning for Quadrupedal Robots
summary
The gist
PROMO introduces Preference-Conditioned Multi-Objective Reinforcement Learning (MORL) for quadrupedal robots, addressing the limitation of fixed scalar rewards by allowing operator intent to
In short
PROMO introduces Preference-Conditioned Multi-Objective Reinforcement Learning (MORL) for quadrupedal robots. It allows a single policy to adapt its behavior dynamically based on operator preferences—like prioritizing tracking, stability, or efficiency—at runtime. This enables robots to trade off locomotion objectives without needing separate controllers or retraining.
Key concepts
- Preference-Conditioned MORL
- This framework conditions a single reinforcement learning policy on a preference vector that defines trade-offs between multiple objectives (tracking, stability, efficiency). Instead of having different policies for each goal, one policy learns how to adjust its actions based on the desired balance specified by the operator.
- Semantic Reward Vector
- This is a set of objectives—like tracking or stability—represented as a vector. The actual utility received by the robot is calculated by taking the dot product of this objective vector and a preference weight vector provided at runtime, allowing dynamic weighting of goals.
- Objective/Prior Factorization
- This design separates the desired high-level semantic objectives from fixed, embodiment-specific locomotion priors (like contact timing). This factorization is crucial because it keeps the core physical requirements internal to the controller, making the preference vector a meaningful interface for changing behavior.
- Action Diversity Regularizer
- A small regularization term added to the actor's objective ensures that even when observations are similar, different preference settings lead to distinct policies. This prevents the policy from converging on a single behavior regardless of the input weights.
Terminology used across episodes
This episode discusses
- PROMO: Preference-conditioned Multi-Objective Reinforcement Learning for Quadrupedal Robots · Paper Radio
- Learning Multiple Gaits within Latent Space for Quadruped Robots
- Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization
- Reward-Conditioned Policies
- In Search for Architectures and Loss Functions in Multi-Objective Reinforcement Learning
- Scalable Multi-Objective Robot Reinforcement Learning through Gradient Conflict Resolution
- Controllability in preference-conditioned multi-objective reinforcement learning
- GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning
The paper
PROMO: Preference-conditioned Multi-Objective Reinforcement Learning for Quadrupedal Robots · Read on arXiv
Amr Mousa, Rifny Rachman, Neil Karavis, Michele Caprio, Richard Allmendinger
University of Manchester
Quadrupedal locomotion requires balancing conflicting objectives such as command tracking, stability, and energy efficiency, yet conventional reinforcement learning (RL) hardcodes these priorities into a fixed scalar reward at training time. We present PROMO (Preference-Conditioned Multi-Objective Reinforcement Learning), a semantic multi-objective approach that makes this trade-off an explicit runtime input to a single locomotion policy. PROMO conditions the policy on deployment facing preferences while keeping embodiment-specific locomotion priors fixed, thereby separating operator intent from reward shaping terms required for viable gait generation. Compared with fixed-objective controllers, multi-objective baselines, and independently trained specialists, PROMO achieves objective specialization and robustness from a single deployable policy. Across 100 sampled preferences in simulation, 67 behaviors are non-dominated under exact Pareto dominance, with a mean preference-objective correlation of 0.843, demonstrating broad Pareto coverage and predictable preference response. The same policy transfers zero-shot to a Unitree Go2, where preference changes alone reduce specific energy by up to 30.4%, position error by 38.7%, and peak body-attitude deviation by 59.0% relative to the balanced preference. These results establish preference-conditioned multi-objective RL as a practical runtime interface for adaptive legged locomotion, extending its role beyond offline Pareto-set construction. Open-source code and videos are available at https://amrmousa.com/promo/.
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.
Dev: Today's paper: "PROMO: Preference-conditioned Multi-Objective Reinforcement Learning for Quadrupedal Robots".
Rosa: PROMO introduces Preference-Conditioned Multi-Objective Reinforcement Learning (MORL) for quadrupedal robots, addressing the limitation of fixed scalar rewards by allowing operator intent to explicitly condition locomotion trade-offs at runtime.
Dev: First, who's behind it and why it matters.
Paper summary: Dev: So, wrapping up on this paper, "PROMO: Preference-conditioned Multi-Objective Reinforcement Learning for Quadrupedal Robots," it seems the core contribution is successfully formalizing control by separating task objectives from fixed embodiment requirements through semantic objective/prior factorization.
Rosa: That separation is what allows a single policy to span diverse behaviors by letting the operator specify which trade-off—tracking, stability, or efficiency—the robot should prioritize at runtime. It’s about making those deployment-facing preferences an explicit input to the policy instead of burying them deep inside a reward function.
Taro: I think the real implication here is that it gives us a way to manage flexibility without needing different controllers for every operational mode; we can command a continuum of desired behaviors through this preference vector.
Dev: From an engineering standpoint, it’s about achieving objective specialization and robustness from just one deployable policy, which simplifies the hardware and software architecture significantly compared to having multiple specialized controllers ready to switch between.
Rosa: It really suggests that the interface is interpretable because we can specify *why* behavior should change—for example, by prioritizing stability over tracking—and the policy learns how to execute that specific trade-off effectively.
Taro: If this holds up under rigorous testing in real-world scenarios, it means autonomous systems will be much more adaptable and less brittle when faced with unpredictable conditions than what we see in current fixed-priority control methods.
Dev: I just hope the implementation details on the loop rate and latency hold up when we move this from simulation to actual hardware deployment, because those real-time constraints are always where these kinds of conditioning mechanisms can introduce failure modes if they aren't tuned properly.
Rosa: We saw that in simulation, for instance, changing just the preference alone reduced specific energy by up to thirty point four percent and position error by thirty-eight point seven percent on the Unitree Go2 hardware, which shows how impactful this interface can be on physical performance right away.
Taro: That measurable impact under those conditions really validates the idea that operator intent can systematically modulate hardware behavior with a fixed policy, proving that this is a meaningful control interface for autonomous systems.
Conclusion: Rosa: It seems like the title itself pretty much sums up what this paper is doing: giving a quadruped robot a way to listen to your goals and adjust its movement priorities dynamically while still operating within its physical limits.
Dev: Yeah, I think the authors are really smart for tackling that trade-off between needing fast computation and needing accurate, real-time decision-making; it's a tough spot for control engineers.
Taro: From an autonomy side, the paper suggests we move toward systems where the robot doesn't just follow a pre-programmed path but actively chooses its operational strategy based on the immediate environment or user command.
Rosa: Exactly, and I'm really curious about the long-term impact; if this works reliably outside of a controlled lab setting, how quickly could we see these robots deployed in genuinely unpredictable environments?
Dev: That’s the big question for me—the deployment longevity. We need to know if this preference interface survives real-world noise and unexpected sensor glitches without causing catastrophic failure modes at the loop rate required for locomotion.
Taro: If the system can handle misbehavior in the world, like sudden obstacles or unexpected terrain changes, then we could envision robots that adapt their movement strategy instantly instead of just crashing or freezing.
Rosa: It feels like this research opens up a whole new category of robot control where the operational goal isn't just "walk here," but rather "walk here *in a stable and efficient way*."
Dev: And from an engineering standpoint, the authors' focus on factorization between the semantic objectives and the fixed locomotion priors is really telling; it seems like they found a way to keep things predictable while still allowing that flexibility.
Taro: That factorization is what makes it promising because it ensures that even when we shift priorities, there's still a solid foundation of embodied physics guiding the decisions.
Rosa: It really shifts the focus from designing one perfect controller for every single scenario to designing one flexible controller that can handle a wide range of mission types.
Dev: And if we look at the results, it shows that these preference changes actually translate into measurable physical improvements in terms of energy use and error reduction on hardware like the Unitree Go2.
Taro: That tangible performance data is what will really convince other researchers that this isn't just theoretical work but something with real utility for complex autonomous missions.
Rosa: So, we've seen the mechanics of how it works, now we need to think about where this technology actually lands in the practical deployment pipeline.
Dev: Right, and that leads perfectly into our next topic: we need to look at how robust this entire framework is when faced with the messy realities of real-world operation.
More episodes
- 2610.10846-Cross-Embodiment Robot Foundation World Models with Latent Actions
- 2610.10601-Teaching a Robot Dog New Tricks: Diverse Quadruped Skills via Combined Reinforcement and Imitation Learning with Adversarial Task Selection
- 2610.10637-TacHair: Tactile Contact-Distribution Guided Online Correction for Robotic Hair Stroking and Perception
- 2610.10646-Masked Generative Motion Planning with Geometry-Guided Token Search
- 2610.10812-Skill-SLM: Agent Skill-driven Small Language Models for Reliable Robot Operation
- 2610.10801-Same Action, Different Outcome: Variability in Dynamic Cloth Manipulation
- 2610.10810-Diagnosing and Recovering from Observation-Space Shift at Long-Horizon Skill Seams
- 2610.10748-TAPNAV: Humanoid Navigation through Tactile Active Perception
- 2610.10855-OmniHOI: Dexterous Hand-Object Interaction from Monocular Human Video
- 2610.11003-ActiveReg: Information-Driven Active Regional Probing for Partial-to-Full Bone Registration