Robust and Efficient MuJoCo-based Model Predictive Control via Web of Affine Spaces Derivatives
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "Robust and Efficient MuJoCo-based Model Predictive Control via Web of Affine Spaces Derivatives".
Dev: This paper introduces a method to accelerate model derivative computations within MuJoCo-based Model Predictive Control (MPC) by replacing finite differencing (FD) with Web of Affine Spaces (WASP) derivatives,
Rosa: First, who's behind it and why it matters.
Title and authors: Rosa: Now let's talk about the specific improvements they propose for this paper, focusing on how they made the WASP method more usable for practitioners through those fraction and tolerance parameters.
Dev: That’s smart engineering because it means control engineers like me can quickly dial in how much approximation we need based on the specific dynamics of the system we are modeling, which should help us optimize for our required loop rate. It gives us more fine-grained control over the execution time versus precision.
Taro: I’m interested in the implication that WASP is designed to function as a drop-in replacement, meaning researchers don't have to rewrite their entire MPC framework just to try out this derivative method. It makes adoption much smoother for new research groups.
Rosa: That’s true; they want it to be integrated directly into the MuJoCo source code so that existing applications can immediately see the speedup without needing massive architectural changes. It focuses on practical integration over theoretical purity in this step.
Dev: The implication for latency is significant because since it scales naturally with MJPC’s parallel execution model, we should see those speedups translate directly into lower end-to-end planning times, which is critical for high-DOF systems. We need to keep an eye on how that scaling plays out under heavy load.
Taro: And I'm really interested in the fact that they showed WASP can significantly outperform sampling-based planners on contact tasks; that suggests a more reliable method for handling the messy physics of real interaction.
Rosa: That reliability is what field robotics demands, Taro; it means when we’re trying to deploy this in the field, we have a better baseline for performance than relying solely on stochastic methods which might get stuck in poor local minima.
Dev: I'm focused on the robustness analysis they ran with parameter variations; that suggests the method isn't overly sensitive to small shifts in the model, which is a huge plus when dealing with imperfect simulations or real-world sensor noise. That resilience is important for deployment stability.
Taro: So if we can trust these approximated derivatives across different robot types—from quadrotors to quadruped climbers—that means we could generalize this for a wider variety of autonomous agents. We’re moving toward more universal control solutions.
Rosa: That generalization is exactly what field robotics is all about; the idea is that once you have a fast, reliable MPC core that isn't bottlenecked by derivative computation, you can focus on designing better high-level behaviors.
Dev: The practical implication for me is that we can push the complexity of the control laws we use, knowing that the underlying math won't crush our execution time during operation. That freedom to be complex without crippling latency is a big deal for my work.
Taro: I think this work moves us closer to having control systems that are not just theoretical models but actually perform better in scenarios with complex physical constraints.
Rosa: Exactly; it’s about creating a control system that is both computationally lean and capable of handling the intricate dynamics we see in the real world.
The paper's summary: Rosa: So, wrapping up our discussion on "Robust and Efficient MuJoCo-based Model Predictive Control via Web of Affine Spaces Derivatives," we’ve seen how this WASP method provides a faster way to compute model derivatives by reusing prior evaluations instead of using brute-force finite differencing.
Dev: That’s the core mechanism, Rosa; it essentially creates a more stable mathematical structure for estimating those necessary derivatives, which is crucial when you're worried about loop rate and how fast the control system can actually react.
Taro: From my view, this paper shows that we can get better performance ratios on contact-rich tasks because WASP handles the dynamics more reliably than some of the other methods we’ve looked at.
Rosa: It really does show that coherence-based derivative approximations offer a compelling balance between efficiency and robustness in iterative control settings for complex robotic systems.
Dev: The implication for us is that we can push the complexity of our control laws because they won't crush our execution time during operation, provided we manage those tunable parameters correctly.
Taro: I think this technology opens up possibilities for deploying much more capable robotic agents in environments that demand quick, dynamic responses outside of a controlled lab setting.
Rosa: That’s right; it suggests that field robotics can move toward systems that are both computationally lean and highly capable of handling intricate physical interactions in real-time.
Dev: We’re really looking forward to seeing how this method holds up when we put these policies into systems that encounter noisy sensor data or unexpected model inaccuracies during prolonged operation.
Taro: That's the next big test; verifying the reliability under continuous, messy real-world conditions is what separates a promising method from one that truly changes how we build autonomous systems.
Rosa: Well, it’s been fascinating looking at this paper on "Robust and Efficient MuJoCo-based Model Predictive Control via Web of Affine Spaces Derivatives."
Dev: I agree; the speedup figures are impressive, and the drop-in replacement aspect makes it very practical for existing systems.
Taro: I just want to keep pushing on how this reliability scales when we move away from perfect simulation environments.
The paper's improvements: Rosa: So, we’ve covered how this paper on "Robust and Efficient MuJoCo-based Model Predictive Control via Web of Affine Spaces Derivatives" shows that replacing finite differencing with WASP derivatives lets us compute model derivatives much faster while keeping performance ratios decent across various robot tasks.
Dev: Exactly; the speedup figures, especially those up to four point zero times for contact dynamics, are significant because they mean we can push the complexity of our control laws because they won't crush our execution time during operation if we manage those tunable parameters correctly.
Taro: I think this technology opens up possibilities for deploying much more capable robotic agents in environments that demand quick, dynamic responses outside of a controlled lab setting, which is really exciting for autonomy.
Rosa: That’s right; it suggests that field robotics can move toward systems that are both computationally lean and highly capable of handling intricate physical interactions in real-time.
Dev: And from an engineering standpoint, the integration into MuJoCo MPC as a drop-in replacement means we don't have to rewrite our entire control stack just to get this speed benefit, which is a practical improvement for existing systems.
Taro: I just want to keep pushing on how this reliability scales when we move away from perfect simulation environments and into genuinely messy real-world conditions.
Rosa: That’s where we need to focus our attention moving forward; the paper gives us a solid foundation showing that coherence-based derivative approximations can offer a balance between efficiency and robustness in iterative control settings.
Dev: I hope they publish more work focusing specifically on the robustness of those derivative approximations when faced with noisy sensor data or unexpected model inaccuracies during prolonged operation.
Taro: Verifying the reliability under continuous, messy real-world conditions is what separates a promising method from one that truly changes how we build autonomous systems.
Rosa: Well, it’s been fascinating looking at this paper on "Robust and Efficient MuJoCo-based Model Predictive Control via Web of Affine Spaces Derivatives."
Dev: I agree; the speedup figures are impressive, and the drop-in replacement aspect makes it very practical for existing systems.
Taro: I just want to keep pushing on how this reliability scales when we move away from perfect simulation environments.
Conclusion: Rosa: So we've seen how "Robust and Efficient MuJoCo-based Model Predictive Control via Web of Affine Spaces Derivatives" shows replacing finite differencing with WASP derivatives lets us compute model derivatives much faster while keeping performance ratios decent across various robot tasks.
Dev: Exactly; the speedup figures, especially those up to four point zero times for contact dynamics, are significant because they mean we can push the complexity of our control laws because they won't crush our execution time during operation if we manage those tunable parameters correctly.
Taro: I think this technology opens up possibilities for deploying much more capable robotic agents in environments that demand quick, dynamic responses outside of a controlled lab setting, which is really exciting for autonomy.
Rosa: That’s right; it suggests that field robotics can move toward systems that are both computationally lean and highly capable of handling intricate physical interactions in real-time.
Dev: And from an engineering standpoint, the integration into MuJoCo MPC as a drop-in replacement means we don't have to rewrite our entire control stack just to get this speed benefit, which is a practical improvement for existing systems.
Taro: I just want to keep pushing on how this reliability scales when we move away from perfect simulation environments and into genuinely messy real-world conditions.
Rosa: That’s where we need to focus our attention moving forward; the paper gives us a solid foundation showing that coherence-based derivative approximations can offer a balance between efficiency and robustness in iterative control settings.
Dev: I hope they publish more work focusing specifically on the robustness of those derivative approximations when faced with noisy sensor data or unexpected model inaccuracies during prolonged operation.
Taro: Verifying the reliability under continuous, messy real-world conditions is what separates a promising method from one that truly changes how we build autonomous systems.
Rosa: Well, it’s been fascinating looking at "Robust and Efficient MuJoCo-based Model Predictive Control via Web of Affine Spaces Derivatives."
Dev: I agree; the speedup figures are impressive, and the drop-in replacement aspect makes it very practical for existing systems.
Taro: I just want to keep pushing on how this reliability scales when we move away from perfect simulation environments.
Yale University
cs.RO
Submitted: 2025-12-24
Updated: 2026-09-28
Comments: Accepted to 2026 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2026)
Code: https://github.com/chen-dylan-liang/mujoco
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 92/100
The gist: This paper introduces a method to accelerate model derivative computations within MuJoCo-based Model Predictive Control (MPC) by replacing finite differencing (FD) with Web of Affine Spaces (WASP)
Key concepts
- Web of Affine Spaces (WASP)
- WASP derivatives are used to replace finite differencing when computing model derivatives. This method creates a more stable mathematical structure for estimating these derivatives, which is crucial for achieving faster computation in Model Predictive Control.
- MuJoCo-based MPC
- This refers to Model Predictive Control implemented within the MuJoCo simulation environment. The paper focuses on improving the efficiency of this control method by speeding up necessary model derivative computations.
- Drop-in Replacement
- The WASP method is designed to function as a drop-in replacement for finite differencing within the MuJoCo MPC framework. This means existing applications can immediately see speedups without requiring massive architectural changes.
- Robustness and Reliability
- The discussion emphasizes the method's robustness, noting that it is not overly sensitive to small shifts in the model. The focus is on verifying its reliability under continuous, messy real-world conditions like noisy sensor data.
Terminology
Summary
This paper introduces a method to accelerate model derivative computations within MuJoCo-based Model Predictive Control (MPC) by replacing finite differencing (FD) with Web of Affine Spaces (WASP) derivatives, aiming to address computational bottlenecks in high-DOF systems.
The core problem addressed is the inefficiency of FD in MPC, which relies on repeated rollouts and requires a simulator call for each input dimension perturbation. This leads to a computational burden that can dominate the control loop, making real-time performance challenging for high-DOF systems or complex scenes. While Automatic Differentiation (AD) tools like XLA and Warp exist, they are often too narrow in scope for MPC's iterative nature, leading to excessively sharp or ill-conditioned derivatives.
The proposed solution is WASP derivatives, which is described as a recently developed method for efficiently computing sequences of approximate derivatives by reusing information from prior, related evaluations, enabling faster computation and improved numerical stability.
WASP formulates derivative estimation as a constrained least-squares problem:
"Each iteration of the algorithm requires only a single Jacobian-vector product (JVP), which defines an affine subspace guaranteed to contain the true derivative. The optimization then seeks the transpose of an approximate derivative that lies within this affine subspace, enforced via a hard constraint, while simultaneously aligning with prior, related computations encoded in the objective function."
The approach allows for a trade-off between accuracy and efficiency controlled by four parameters: "Four parameters are used to control the number of JVPs used per solution, denoted as pmax, pmin, pθ, pn. Here, pmax and pmin set a maximum and minimum on the number of JVPs, respectively. The pθ and pn stop the JVP calculations and return a result when the current approximate derivative elicits approximate JVPs that sufficiently match the angle and norm of ground truth JVPs, respectively. The paper further simplifies this interface for users by exposing two parameters:
a fraction parameter, denoted as frac, and a tolerance parameter, denoted as tol, where
frac = pmin / pmax and
tol = pθ = pn." At maximum settings (frac = 1, tol = 0), WASP reduces to become equivalent to FD.
The integration into the MuJoCo MPC (MJPC) framework is designed to be a drop-in replacement: "Finally, we directly incorporate these changes into the MuJoCo source code alongside the existing FD implementation in the C programming language. This integration allows WASP to function as a true drop-in replacement for all downstream MuJoCo-based applications. The design ensures that
the benefits of WASP scale naturally with MJPC’s parallel execution model."
The evaluation across a diverse suite of MJPC tasks spanning nine robot embodiments—including Quadrotor, Swimmer, Quadruped Climb, and Humanoid Walk—demonstrates the efficacy of WASP. In Experiment 1 (WASP vs. FD), results showed that WASP achieves speedups ranging from 1.26× to 2.08× compared to FD for model derivative computation across all tasks while maintaining at least sufficient task performance (defined here as Performance Ratio ≥ 0.7).
Notably, for some tasks such as Quadrotor, Swimmer, and Quadruped Stand, WASP achieves performance ratios greater than 1, indicating improved task performance despite using approximated derivatives.
In Experiment 2 (WASP-based iLQG vs. Sampling-Based Planners), WASP-based iLQG significantly outperforms sampling-based planners on these contact-rich tasks,
with results showing WASP achieving up to 4.0× speedups
for locomotion tasks with complex contact dynamics.
The parameter robustness analysis (Experiment 3) showed that WASP is generally robust to parameter variations, with accuracy in state transitions being more critical than control accuracy.
The paper concludes that WASP consistently reduced the computational burden of model derivative evaluations while maintaining, and in some cases even improving, task performance,
suggesting that coherence-based derivative approximations can offer a compelling balance between efficiency and robustness in iterative control settings.
The work provides an open-source implementation of MJPC with WASP derivatives to support adoption. The limitations noted are that experiments were conducted in simulation, and structural limitations of short-horizon gradient-based MPC on contact-rich manipulation tasks remain an area for future architectural changes. (See Table I and Figure 3 for detailed performance comparisons).
The paper's primary contribution is the demonstration that WASP derivatives can serve as a robust and efficient drop-in replacement for finite differencing in MJPC, achieving speedups of up to 2x while maintaining or improving task performance. The open-source implementation is released to facilitate immediate experimentation. (See Section IV: Implementation Details). (See Table II for a comprehensive comparison against stochastic sampling-based planners in Experiment 2). (See Figure 3 for parameter sensitivity analysis in Experiment 3).
Improvements for AI systems
Here are the specific improvements to AI systems based on the provided research, along with what those improved systems can achieve:
-
The integration of a Web of Affine Spaces (WASP) derivative backend as a drop-in replacement for finite differencing (FD) within MuJoCo Model Predictive Control (MJPC).
-
A 2x speedup in the computation time for model derivatives compared to FD, while maintaining robust performance across diverse robot embodiments.
-
The ability of WASP-based MPC to outperform stochastic sampling-based planners on contact-rich tasks (like quadruped locomotion) due to greater efficiency and reliability.
-
The implementation of tunable parameters (via a fraction parameter, 'frac', and a tolerance parameter, 'tol') that allow practitioners to dynamically balance the trade-off between derivative approximation accuracy and computational speed.
This improved AI system can achieve the following:
-
It can perform high-speed, real-time model predictive control for complex robotic systems (e.g., quadrotors, bipeds, and quadruped robots) directly within physics simulators like MuJoCo or similar environments without requiring source code modifications to the simulator itself.
-
It can execute sophisticated locomotion behaviors (like walking gaits or climbing) with significantly reduced planning latency, enabling faster decision-making and more responsive control in dynamic physical environments.
-
It can reliably handle complex contact dynamics (e.g., those found in quadruped tasks) where traditional gradient-based methods struggle, by leveraging the stability provided by WASP derivatives, leading to better task success rates compared to purely stochastic sampling methods.
-
It allows researchers and practitioners to rapidly prototype and benchmark different derivative approximations against the standard FD method simply by adjusting the 'frac' and 'tol' parameters in a user-friendly interface, accelerating the development of more efficient control policies.
Abstract
MuJoCo is a powerful and efficient physics simulator widely used in robotics. One common way it is applied in practice is through Model Predictive Control (MPC), which uses repeated rollouts of the simulator to optimize future actions and generate responsive control policies in real time. To make this process more accessible, the open source library MuJoCo MPC (MJPC) provides ready-to-use MPC algorithms and implementations built directly on top of the MuJoCo simulator. However, MJPC relies on finite differencing (FD) to compute derivatives through the underlying MuJoCo simulator, which is often a key bottleneck that can make it prohibitively costly for time-sensitive tasks, especially in high-DOF systems or complex scenes. In this paper, we introduce the use of Web of Affine Spaces (WASP) derivatives within MJPC as a drop-in replacement for FD. WASP is a recently developed approach for efficiently computing sequences of accurate derivative approximations. By reusing information from prior, related derivative calculations, WASP accelerates and stabilizes the computation of new derivatives, making it especially well suited for MPC's iterative, fine-grained updates over time. We evaluate WASP across a diverse suite of MJPC tasks spanning multiple robot embodiments. Our results suggest that WASP derivatives are particularly effective in MJPC: it integrates seamlessly across tasks, delivers consistently robust performance, and achieves up to a 2 speedup compared to an FD backend when used with derivative-based planners, such as iLQG. In addition, WASP-based MPC outperforms MJPC's stochastic sampling-based planners on our evaluation tasks, offering both greater efficiency and reliability. To support adoption and future research, we release an open-source implementation of MJPC with WASP derivatives fully integrated.
Sources
- Predictive Sampling: Real-time Behaviour Synthesis with MuJoCo
- ad-trait: A Fast and Flexible Automatic Differentiation Library in Rust
- Coherence-based Approximate Derivatives via Web of Affine Spaces Optimization
- Whole-Body Model-Predictive Control of Legged Robots with MuJoCo
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving