Toward Hardware-Agnostic Quadrupedal World Models via Morphology Conditioning

arXiv:2604.08780 · cs.RO, cs.LG · Submitted 2026-08-17 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Toward Hardware-Agnostic Quadrupedal World Models via Morphology Conditioning".

Jane: The paper was written by Mohamad H. Danesh, Chenhao Li, Amin Abyaneh, Anas Houssaini, Kirsty Ellis et al. from McGill University and Mila - Quebec AI Institute and ETH Zurich and Universite de Montreal.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Welcome back to the arXiv Review, everyone! I'm your host Tom, and with me as always is the brilliant Jane. Jane, we've got a paper today that's got me genuinely excited — it's called "Toward Hardware-Agnostic Quadrupedal World Models via Morphology Conditioning."

Jane: Tom, I'm thrilled you brought this one in. The title alone tells you something big is happening. We're talking about robot dogs — quadrupeds — and the idea that one single AI brain could control all of them, regardless of whether it's a tiny Unitree Go1 or a massive ANYmal.

Tom: And that's the key word in the title, right? "Hardware-agnostic." That means the model doesn't care what hardware it's running on. It's like if you had one universal driver's license that worked for a go-kart, a sedan, and a semi-truck without any retraining.

Jane: Exactly. And the authors are from McGill, Mila, ETH Zurich, and Université de Montréal. That's a powerhouse lineup. The first author, Mohamad Danesh, has been working on this for a while. They're really pushing the boundaries of what we call "world models."

Tom: So for our listeners who might be new to this — what exactly is a world model? I know we've talked about it before, but let's keep it simple.

Jane: Think of it as the robot's imagination. The AI learns how the world works — gravity, friction, momentum — and then it can simulate future outcomes in its head. Instead of needing to physically try something a million times, it can imagine what would happen and learn from that imagination.

Tom: And that's the game-changer here. Previously, these world models were "hardware-locked specialists." You train one on a Spot, and it's useless on a Go1. The physics are different, the limb lengths are different, the mass is different. It's like learning to ride a bicycle and then being asked to ride a unicycle — the balance rules are completely different.

Jane: Right. And the old approach would be to just train a new model from scratch for each robot. That's expensive, it takes millions of samples, and it's just wasteful. This paper says, "What if we teach the model the universal physics of walking, and then just tell it what robot it's on?"

Tom: That's the "morphology conditioning" part. They extract the robot's physical specs — limb lengths, mass, joint configuration — from its design file, and they feed that into the model as context. So the model knows, "I'm on a heavy robot with long legs" versus "I'm on a light robot with short legs."

Jane: And the beauty is, the physics of walking is the same. Gravity pulls down, feet push against the ground, momentum carries you forward. The robot's specific body just changes how those universal laws play out.

Tom: So they're not learning eight different walking styles. They're learning one walking physics, and then adapting it to eight different bodies. That's the vision, and it's a big one.

Jane: It really is. And I think the implications go way beyond robot dogs. If this works, it could apply to any robot — bipeds, manipulators, even drones. But let's not get ahead of ourselves. We need to talk about how they actually pulled this off.

Tom: Absolutely. And that's exactly what we're going to dig into next. Stay tuned.

Summary: Tom: Welcome back to the arXiv Review. We're diving into "Toward Hardware-Agnostic Quadrupedal World Models via Morphology Conditioning." Jane, we set the stage — now let's talk about the actual meat of the paper. How did they build this thing?

Jane: So they built on an existing architecture called DreamerV3, which is a well-known world model framework. But they added three key innovations. First, they created a "Physical Morphology Encoder" — that's the part that reads the robot's design file and extracts the important physical features.

Tom: And what kind of features are we talking about? I mean, a robot design file is a huge mess of links, joints, and masses.

Jane: Right, it's a lot of data. But they distilled it down to the essentials. Things like the length of the thigh and shank, the hip offset, the stance width, the total mass, and even the knee configuration — whether the knees bend inward like an ANYmal or backward like a dog.

Tom: And that last one is crucial. You can't just scale up a dog-like robot to get an X-configuration robot. The kinematics are fundamentally different. The joints move in different directions.

Jane: Exactly. And they normalized all these features so they're on a comparable scale. So instead of saying "this robot weighs eighty kilograms," they use a log scale and ratios. That way, the model isn't confused by the raw numbers being orders of magnitude apart.

Tom: Then they inject this morphology vector into the world model in two places — the encoder and the recurrent dynamics. Can you explain why that matters?

Jane: Think of the recurrent state as the robot's short-term memory. It tracks what's been happening — velocities, positions, contacts. If you don't tell it what robot it's on, it has to guess from history. That's called "implicit system identification," and it's slow and dangerous. The robot has to move around and bump into things before it figures out its own body.

Tom: And this paper says, "No, we have the design file. We know exactly what the robot is. Why make it guess?"

Jane: Precisely. By injecting the morphology explicitly at every time step, the recurrent state is freed up to focus on the dynamic state — how fast am I moving, when did my foot touch the ground — rather than wasting capacity on "what am I?"

Tom: And then there's the third piece — the reward normalizer. I imagine training across robots with vastly different reward scales is a nightmare.

Jane: Oh, absolutely. The Boston Dynamics Spot gets rewards around three hundred fifty per episode, while the ANYmal-D gets around twenty-five. If you train on raw rewards, the Spot's signal dominates everything, and the model just ignores the smaller robots.

Jane: So they built an adaptive normalizer that tracks each robot's reward distribution and scales it. It's like turning down the volume on the loud robot so everyone can be heard.

Tom: And the results? I saw the learning curves in the paper — they're pretty dramatic.

Jane: They are. The full QWM framework converges fast and stable. But if you remove the reward normalizer, the model just flatlines — it never learns anything. If you remove the morphology encoder, it learns to walk but plateaus at a lower reward. Every component matters.

Tom: So it's a well-engineered system. But the real test — and I think this is what everyone wants to know — is whether it can handle a robot it's never seen before. That's the zero-shot generalization question.

Jane: And that's exactly what we're going to talk about next. The results there are genuinely surprising.

Improvements: Tom: Welcome back. We're still on "Toward Hardware-Agnostic Quadrupedal World Models via Morphology Conditioning." Jane, we talked about the architecture — now let's get to the exciting part. Zero-shot generalization. They trained on seven robots and held out three for testing.

Jane: Right. And the results are a mix of triumph and a really honest look at limitations. They held out the ANYmal-D and the Unitree Go1 — robots that are structurally similar to ones in the training set. And the model just... worked. It walked on both of them, first try, no fine-tuning, no adaptation period.

Tom: And not just barely walking. The episode lengths were nine hundred forty-eight and nine hundred seventy-four out of one thousand steps. That's essentially stable locomotion for the entire episode. The specialist PPO baseline, which trained specifically on those robots, got nine hundred eighty-one and nine hundred ninety-six. So QWM is within a few percent of a specialist that had millions of interactions.

Jane: And that's the key comparison. They also tested a model-free policy that was conditioned on the same morphology features — PME-PPO. It did much worse, with episode lengths around five hundred thirty and six hundred two. So it's not just about feeding the morphology to a policy. The world model's learned latent dynamics are what make the difference.

Tom: So the world model isn't just a simulator — it's a physics adapter. It takes the raw observations from an unseen robot and translates them into a latent space the policy already understands.

Jane: That's exactly the right way to think about it. The policy never sees raw observations. It only sees the latent states produced by the world model. So when you swap in a new robot's morphology, the world model adjusts the latent space to match, and the policy just does its thing.

Tom: But then there's the Unitree B2 — the outlier. That's where things get interesting.

Jane: Yeah, and I love that they included this. The B2 is nearly four times heavier than the Go1, with a much larger stance. It's a genuine outlier in the morphology space. And the zero-shot performance on it was poor — episode length around four hundred five reward near zero.

Tom: So the model failed on it. But they're upfront about why.

Jane: They are. It's a distribution-bounded interpolator, not a universal physics engine. The model can interpolate between known morphologies, but it can't extrapolate to something that's far outside the training distribution. The B2 is the most distant robot in the entire cohort, so the model has no reference points to work with.

Tom: That's actually a really important finding. It tells us the limits of this approach. You need a training set that spans the "physics basis" of mass and geometry.

Jane: And they say that explicitly in the paper. To get true universality, you need a diverse enough training cohort. But here's the thing — they also deployed the zero-shot agents on real hardware. The held-out Go1 and ANYmal-D, with no real-world fine-tuning at all.

Tom: And it worked? On actual physical robots?

Jane: It worked. The Go1 produced a high-frequency trot appropriate for its lightweight body, and the ANYmal-D produced a slower, more grounded gait. The tracking error was within about ten percent of a specialist baseline. Zero falls across twenty trials.

Tom: That's the part that gets me. It's one thing to work in simulation, but real hardware has backlash, friction variability, all sorts of unmodeled effects. And the model just handled it.

Jane: Because the morphology conditioning captures the fundamental physics, not just simulation artifacts. The model knows the Go1 is light and agile, so it produces a light and agile gait. The ANYmal-D is heavy, so it produces a heavy, stable gait.

Tom: So where does this go next? What's the big vision?

Jane: That's what we're going to wrap up with. The implications go way beyond robot dogs.

Conclusion: Tom: We're back for the final segment on "Toward Hardware-Agnostic Quadrupedal World Models via Morphology Conditioning." Jane, let's bring in the rest of the team to talk about what this means for the future.

Jane: Absolutely. Lu, you've been listening — what's your take on the big picture?

Lu: I think this is a foundational step toward what the paper calls "Type III: Robotic Agents." We're moving from specialists to generalists. The idea that you can train one world model and deploy it on any quadruped — that's the beginning of something much bigger. I'm thinking about humanoids, manipulators, even drones.

Meng: But I want to push back on that a little. The paper is honest about the limitations — it's an interpolator, not a universal physics engine. The Unitree B2 failure shows that. So what's the practical path to scaling this up?

Jane: That's a fair question. The authors suggest expanding the training cohort to span more of the morphology space. And they mention moving from fixed vectors to graph neural networks or transformers to handle variable kinematic trees.

Lu: And that's the key. If you can handle arbitrary kinematic trees, you're not just doing quadrupeds. You're doing any articulated body. That's the path to a truly universal world model.

Meng: But what about the engineering side? The paper uses a custom simulation framework called Hetero-Isaac to train on multiple robots simultaneously. That's a significant piece of infrastructure. Is that going to be accessible to the broader community?

Jane: They've open-sourced it. The code is available. And the fact that they built it on Isaac Lab, which is already widely used, lowers the barrier significantly.

Lalam: I want to add a cultural perspective here. When I look at this work, I see more than just robotics. I see a shift in how we think about AI systems — from brittle, task-specific tools to adaptable, embodied agents. The ability to transfer knowledge across physical forms is a step toward AI that can genuinely operate in the world, not just in a simulation.

Tom: That's a beautiful way to put it, Lalam. And I think it connects to something deeper. When we see a robot dog walk, we understand it because we share the same physics. This model is learning that shared physics, not just memorizing a particular body.

Jane: And that's the real achievement here. The model learned the physics of walking — the universal laws — and then applied them to specific bodies. That's a different kind of intelligence than what we usually see.

Tom: So to wrap up — "Toward Hardware-Agnostic Quadrupedal World Models via Morphology Conditioning" is a significant step toward generalizable world models. It works on unseen robots, it transfers to real hardware, and it honestly identifies its limitations.

Jane: And the limitations are just as important as the successes. Knowing that the model can't extrapolate to extreme outliers tells us where the research needs to go next.

Lu: I'd say the next big milestone is a model that can handle a humanoid. The physics are more complex, the balance problem is harder, but the same principles apply.

Meng: And I'd like to see the simulation infrastructure mature so that training on hundreds of morphologies is feasible, not just eight.

Lalam: And I'd like to see this approach extended to tasks beyond locomotion. If the world model can generalize across bodies, can it generalize across tasks? That's the ultimate question.

Tom: Well, we'll be watching for that. Thanks to everyone for joining us on this deep dive. Jane, always a pleasure.

Jane: Likewise, Tom. And to our listeners — keep an eye on this research. It's a glimpse of where embodied AI is heading.

Tom: That's all for today's episode of the arXiv Review. We've been discussing "Toward Hardware-Agnostic Quadrupedal World Models via Morphology Conditioning." Next time, we'll be looking at a paper on safe exploration in reinforcement learning. Until then, keep learning, keep questioning, and keep imagining. Goodbye, everyone!

Mohamad H. Danesh, Chenhao Li, Amin Abyaneh, Anas Houssaini, Kirsty Ellis, Glen Berseth, Marco Hutter, Hsiu-Chin Lin

McGill University · Mila - Quebec AI Institute · ETH Zurich · Universite de Montreal

cs.RO, cs.LG

Submitted: 2026-08-17

Updated: 2026-08-18

Code: https://github.com/Genesis-Embodied-AI/Genesis

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 90/100

Key concepts

World Model
The AI's imagination. It learns how the world works, such as gravity and friction, allowing it to simulate future outcomes in its head instead of needing millions of physical trials. This enables the robot to learn from its imagination.
Morphology Conditioning
The process where the AI extracts a robot's specific physical characteristics—like limb lengths, mass, and joint configuration—from its design file. This information is fed into the world model as context so it can adapt universal walking physics to that specific body.
Zero-Shot Generalization
The ability of the model to perform a task on a robot it has never seen before, without any prior fine-tuning. The paper showed success on structurally similar robots but failed when tested on a significantly different robot, indicating limits in its universality.
Reward Normalizer
A component that tracks and scales the reward signals from different robots. This prevents models trained on high-reward robots from being dominated by low-reward ones, ensuring the model learns effectively across varied physical systems.

Terminology

Summary

Summary

This paper introduces the Quadrupedal World Model (QWM), a framework for training a single, generalizable world model capable of controlling a heterogeneous fleet of quadrupedal robots. The authors address the Hardware Lottery problem, where standard policies and world models are overfit to specific robot kinematics and dynamics, requiring retraining for any hardware variation. The core contribution is a method to disentangle environmental dynamics from robot morphology by explicitly conditioning the generative dynamics on the robot's engineering specifications, rather than relying on implicit system identification.

The paper states: "We address the limitations of implicit system identification, where treating static physical properties (like mass or limb length) as latent variables to be inferred from motion history creates an adaptation lag that can compromise zero-shot safety and efficiency. Instead, we explicitly condition the generative dynamics on the robot's engineering specifications." This explicit conditioning is achieved through three key architectural innovations built upon the DreamerV3 backbone.

First, the Physical Morphology Encoder (PME) derives a static, scale-invariant feature vector µ directly from the robot's USD file. This vector encapsulates kinematic, geometric, dynamic, and actuation properties: µkin = [lhip, lthigh, lshank, kcfg], µgeo = [lstance, wstance, lstance /wstance], µdyn = [log(1 + Mtotal), mtrunk /Mtotal], and µact = (1/(Mtotal · g0 · Nj)) · Σ τmax. These features are min-max normalized to [−1, 1] for numerical stability.

Second, the Morphology-Conditioned World Model integrates µ into the RSSM architecture via a dual-tower encoder and direct injection into the recurrent dynamics. The encoder uses a dynamic tower for proprioceptive data and a static tower for µ, which are fused before being passed to the model. Crucially, the transition function is augmented to ht = fϕ (ht−1, zt−1, at−1, µ), which relieves the RNN of the burden of memorizing static physical properties in its short-term memory ht.

Third, the Adaptive Reward Normalization (ARN) addresses the challenge of heterogeneous reward scales across different robots. The authors note that a torque penalty of 100 Nm is standard for a heavy ANYmal-D robot but would indicate catastrophic failure for a lightweight Unitree A1. ARN uses an Exponential Moving Average to track the 5th and 95th percentiles of returns per robot, normalizing the reward signal to prevent any single embodiment from dominating the learning process.

The experimental setup uses a custom Hetero-Isaac environment built on NVIDIA Isaac Lab, which manages a heterogeneous batch of robots in parallel. The training cohort includes ANYmal (B, C, D), Unitree (Go1, Go2, A1, B2), and Boston Dynamics Spot, spanning diverse kinematic configurations and physical scales.

The paper validates three central hypotheses. First, Universal Mastery: QWM successfully learns locomotion across the heterogeneous fleet with a single set of weights, outperforming baselines like DreamerV3, PWM, and TWISTER, which struggle with convergence due to implicit system identification. The authors state: "Without explicit conditioning, these models must treat morphology as a latent variable inferred solely from history. This results in a 'mean-dynamics' collapse, where the model approximates the average physics of the fleet rather than the specific dynamics of the current robot."

Second, Zero-Shot Generalization: QWM demonstrates successful transfer to unseen morphologies within the training distribution's support. On held-out robots (Unitree Go1 and ANYmal-D), the agent achieves performance competitive with a specialist PPO baseline trained exclusively on that single robot. However, the paper carefully delineates the limits: QWM operates as a distribution-bounded interpolator within the quadrupedal morphology family rather than a universal physics engine. For the Unitree B2, a significant extrapolation outlier, zero-shot performance degrades substantially, confirming the model's reliance on the support of the training distribution.

Third, Physical Transfer: The frozen zero-shot agents were deployed on real hardware (Unitree Go1 and ANYmal-D) without any fine-tuning. The paper reports: "On the Unitree Go1, the agent generated a high-frequency trot appropriate for the lightweight chassis... On the ANYmal-D, the agent automatically adapted to the increased mass and rotational inertia, producing a slower, more grounded gait to maintain stability." Quantitative results show tracking errors within 10% of a specialist baseline, with zero falls across 20 trials.

The paper also includes extensive ablation studies demonstrating the necessity of each component. Removing ARN completely prevents learning, while removing explicit morphology conditioning (PME) leads to suboptimal asymptotic reward. The authors also provide latent state visualizations (PCA and t-SNE) showing that the recurrent state ht encodes morphological identity, while the stochastic state zt is morphology-agnostic, confirming the disentanglement achieved by the architecture.

The authors conclude: "Future work will extend QWM beyond 'blind' walking on quadrupeds. We aim to integrate visual observations for geometry-aware planning and disentangle task representations to support diverse skills... paving the way for a Universal WM capable of controlling any articulated rigid body, from bipeds to manipulators."

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in AI systems, along with the resulting capabilities.

Improvement: I will build a world model (a neural network that predicts the future states of an environment) that is explicitly conditioned on the robot's physical structure (morphology). Instead of treating the robot's dimensions, mass, and joint configuration as a hidden variable to be inferred from past interactions, I will extract these features directly from the robot's engineering specification (e.g., URDF/USD file) and feed them into the model as a conditioning vector.

  • Specific Implementation: I will create a Physical Morphology Encoder (PME) that processes a vector of normalized, scale-invariant features (limb lengths, mass ratios, stance geometry, torque density) from the robot's spec file. This vector µ will be injected into two critical points of the world model's recurrent architecture (the RSSM): (1) the observation encoder, and (2) the recurrent state update. This ensures the model's latent state is explicitly steered toward the correct dynamics manifold for that specific robot.

Improved AI System Capability: The resulting AI system can control a robot it has never seen before, without any fine-tuning, adaptation, or warm-up period. For example, after training on a fleet of seven different quadrupeds, the system can be given the spec file for a new, unseen quadruped (like a Unitree Go1 or ANYmal-D) and immediately generate stable, effective locomotion gaits (walking, trotting) in simulation and on real hardware. This eliminates the dangerous and time-consuming process of online system identification.

Abstract

World models promise a paradigm shift in robotics, where an agent learns the underlying physics of its environment once to enable efficient planning and behavior learning. However, current world models are often hardware-locked specialists: a model trained on a Boston Dynamics Spot robot fails catastrophically on a Unitree Go1 due to the mismatch in kinematic and dynamic properties, as the model overfits to specific embodiment constraints rather than capturing the universal locomotion dynamics. Consequently, a slight change in actuator dynamics or limb length necessitates training a new model from scratch. In this work, we take a step towards a framework for training a generalizable Quadrupedal World Model (QWM) that disentangles environmental dynamics from robot morphology. We address the limitations of implicit system identification, where treating static physical properties (like mass or limb length) as latent variables to be inferred from motion history creates an adaptation lag that can compromise zero-shot safety and efficiency. Instead, we explicitly condition the generative dynamics on the robot's engineering specifications. By integrating a physical morphology encoder and a reward normalizer, we enable the model to serve as a neural simulator capable of generalizing across morphologies. This capability unlocks zero-shot control across a range of embodiments. We introduce, for the first time, a world model that enables zero-shot generalization to new morphologies for locomotion. While we carefully study the limitations of our method, QWM operates as a distribution-bounded interpolator within the quadrupedal morphology family rather than a universal physics engine, this work represents a significant step toward morphology-conditioned world models for legged locomotion.

Sources

Related papers