SurGE: Surrogate Gradient-guided Evolution for Co-design of Legged Robots with Parallel Elasticity

arXiv:2606.21866 · cs.RO · Submitted 2026-06-20 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "SurGE: Surrogate Gradient-guided Evolution for Co-design of Legged Robots with Parallel Elasticity".

Dev: Co-design of legged robots with elastic elements is challenging due to the non-differentiability of contact dynamics and mechanism engagement.

Rosa: First, who's behind it and why it matters.

Title and authors: Rosa: Let's talk about the title and who came up with this work, SurGE: Surrogate Gradient-guided Evolution for Co-design of Legged Robots with Parallel Elasticity. It’s pretty descriptive, highlighting the core mechanism they developed to solve that non-differentiability issue we talked about earlier.

Dev: The authors listed are Zhuang, Qin, Lu, Shen, Wang, He and Ding; it looks like a solid group coming from different areas of robotics and control engineering. I'm curious if their collective expertise really covers both the control policy side and the dynamics modeling aspect well enough to tackle this specific co-design problem.

Taro: I’m more interested in what that title implies about the scope; "Co-design" suggests they aren't just optimizing a controller or just a robot body, but linking those two parts together through this surrogate gradient technique.

Rosa: Exactly, Taro; it points to the integration aspect—they are trying to optimize the physical structure and the control policy simultaneously in one go, which is inherently more complex than doing them separately.

Dev: From an engineering standpoint, that co-design challenge is exactly why we need this; treating it as a single optimization problem where design parameters map directly to control policies via a differentiable surrogate model seems like a very efficient way to approach the computational expense.

Taro: I think the implication here is that they are addressing the practical bottleneck in legged robotics, which is often how you get an elastic structure to move reliably without breaking or losing stability during locomotion.

Rosa: That’s right; it addresses that physical reality where elasticity introduces complexities into the contact and mechanism engagement models that standard optimization tools struggle with because they aren't smooth.

Dev: The potential impact on the world, if this works reliably outside of a highly controlled lab setting, is huge for applications requiring compliant locomotion, like search and rescue or navigating unstructured terrain.

Taro: If this framework can consistently transfer its learning to physical hardware as the paper suggests, it moves us closer to deploying robots in real-world scenarios where the environment is messy and unpredictable.

Rosa: It certainly points toward a future where complex physical systems are designed using AI methods that can handle the inherent complexities of their own physics rather than requiring manual, iterative tuning for every single design iteration.

Dev: I just hope they’ve accounted for the latency issues we always worry about in these closed-loop systems; if the surrogate model is too slow, that whole gradient injection pipeline becomes useless regardless of how good the theory is.

Taro: That’s a fair point on execution speed; the theoretical efficiency must translate into a practical execution speed that isn't bottlenecked by computation when we need rapid adaptation during operation.

Rosa: So, while they solved the design optimization problem in simulation, the big question for us is how long these designs remain stable and performant when faced with real-world wear and tear.

Dev: That’s what I want to know; we need to see if this framework can maintain those tight concentration metrics after several cycles of simulated or actual deployment.

The paper's summary: Rosa: The paper summarizes SurGE by explaining that it creates a differentiable pipeline using a kinodynamic single-rigid-body model and a design-aware control policy to compute surrogate gradients, which are then injected into CMA-ES via mean shift with cosine-annealed step decay.

Dev: Essentially, they're replacing the expensive inner loop optimization—optimizing the control policy for every candidate design—with a single model that predicts that optimal solution beforehand, which is a smart way to save computation time.

Taro: That replacement of the per-design inner optimization with an amortized training model is what I find most interesting from an autonomy perspective; it simplifies the overall process dramatically.

Rosa: It does simplify things by allowing the outer design search to proceed with a policy that’s held fixed, instead of retraining a controller for every new physical configuration we test.

Dev: The methodology also involves aligning states at every time step using a non-differentiable simulator to compute those surrogate gradients, which is necessary because the full dynamics aren't differentiable.

Taro: That state alignment step is the bridge between the smooth surrogate model and the hard reality of the simulator, and it seems crucial for getting meaningful information from those gradients despite the underlying physics being discontinuous.

Rosa: And they use this information to guide CMA-ES through a mean shift process, using a cosine-annealed schedule to manage how aggressively that guidance is applied during different generations of the search.

Dev: That specific mechanism for injecting the gradient into CMA-ES is sophisticated; it’s not just dumping the gradient in and hoping for the best, but scaling it based on both the covariance matrix and that annealing schedule.

Taro: This structured approach to gradient injection suggests a very careful balance between exploration—letting CMA-ES explore widely—and exploitation—using the guidance to focus on promising areas.

Rosa: Overall, they've managed to tackle the non-differentiability of legged robot dynamics by creating a differentiable pipeline for gradient computation and then feeding that information into an evolutionary search algorithm.

Dev: So, they’ve successfully created a hybrid framework where the simulation provides the necessary differentiability for gradient calculation, while CMA-ES handles the broader design search structure.

The paper's improvements: Rosa: The main improvements discussed are that SurGE moves beyond just using standard black-box methods like vanilla CMA-ES by incorporating surrogate gradients to guide the mean shift in the search distribution.

Dev: Specifically, they achieve six times lower cross-seed standard deviation and eighteen percent tighter population concentration compared to vanilla CMA-ES on their four-DOF hopping robot design space.

Taro: The quantitative results on those metrics are what really sell this framework; those numbers show a tangible improvement in the quality of the designs being proposed, not just theoretical correctness.

Rosa: Beyond the simulation, they also demonstrate that starting from a hand-tuned initial design, SurGE reduces the design objective by thirty-seven point six five percent on hardware experiments in a two-dimensional design subspace <ref:2606.21866#pg0,that starting from a hand-tuned initial design, SurGE reduces the design>.

Dev: That hardware result is what really validates the transfer of optimization trends from simulation to physical deployment; it proves the simulation isn't just generating plausible but ultimately useless designs.

Taro: The consistency between the simulated guidance and the physical performance makes this approach much more robust because it overcomes a major hurdle in robotics: ensuring that what works in a simulator actually works on the robot.

Rosa: Furthermore, they incorporate a design-aware control policy trained via Reinforcement Learning and adapted through a teacher-student architecture with regularized online adaptation to help bridge the sim-to-real gap.

Dev: That teacher-student architecture for the policy is clever because it allows us to generate controllers that are inherently optimized for specific physical designs, which should amortize the cost of needing a separate inner loop optimization later.

Taro: Amortizing that per-design optimization cost is a massive win if we consider the entire design space; it means we can explore more configurations without having to solve an inner optimization problem every single time.

Rosa: So, in essence, the improvements are about using gradient information from a differentiable surrogate model to steer an evolutionary search toward better designs that perform better both in simulation and on physical hardware.

Dev: It’s a solid methodology because it combines the gradient guidance with the annealing schedule to manage exploration and exploitation effectively during the search process, which is something vanilla CMA-ES just can't do on its own.

Conclusion: Rosa: To wrap up, SurGE provides a method for co-designing legged robots with parallel elasticity by using a differentiable pipeline to compute surrogate gradients and injecting them into CMA-ES via mean shift with cosine-annealed step decay.

Dev: The key takeaways are that this technique yields six times lower cross-seed standard deviation and eighteen percent tighter population concentration compared to vanilla CMA-ES, and it shows a thirty-seven point six five percent reduction in the design objective on hardware testing.

Taro: The implication is that we have a framework that can more effectively search for complex physical designs by leveraging gradient information derived from differentiable models, which is essential when dealing with the non-smooth nature of real robot dynamics.

Rosa: It really points toward a future where we can autonomously discover and validate optimal physical hardware configurations for legged robots in a computationally efficient way, relying on the SurGE approach to handle the challenges of elasticity.

Dev: I just have to stress that while they show significant simulation-to-real transfer, we still need more data on how long these designs maintain their performance when faced with real-world wear and tear during extended operation.

Taro: That's a critical consideration; if we want this to be truly useful for field deployment, the durability of the designs needs to be rigorously tested over long operational timescales.

Rosa: So, that’s our discussion on SurGE, and it’s been really insightful looking at how they combine surrogate gradients with CMA-ES for robot co-design.

University of Michigan

cs.RO

Submitted: 2026-06-20

Updated: 2026-10-07

Comments: 8 pages, 7 figures. Accepted for publication at IROS 2026. Website at https://arcad-lab-um.github.io/surge-codesign/

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 84/100

The gist: Co-design of legged robots with elastic elements is challenging due to the non-differentiability of contact dynamics and mechanism engagement.

Key concepts

Surrogate Gradient
A mathematical approximation of the true gradient of the robot's performance objective. Since the actual dynamics are hard to differentiate, this model creates a smooth, differentiable path from design parameters to performance metrics, allowing gradients to be calculated even in complex systems.
Kino-SRB Model
A kinodynamic single-rigid-body model used as the surrogate dynamics. This model captures the essential movement of the robot's center of mass and incorporates elastic forces smoothly based on design parameters, providing a differentiable way to simulate how design changes affect motion.
CMA-ES with Mean Shift
CMA-ES is an evolutionary algorithm that searches for optimal robot designs. SurGE enhances it by using 'mean shift,' which nudges the search distribution in the direction of the calculated surrogate gradient. This guides CMA-ES to promising design regions much faster than standard evolution.
Design-Aware Control Policy
A neural network policy that controls the robot's movement, designed to be aware of its physical design parameters. By encoding normalized design parameters into the policy input, this controller ensures that the learned control strategy adapts effectively to different robot configurations.

Terminology

Summary

Co-design of legged robots with elastic elements is challenging due to the non-differentiability of contact dynamics and mechanism engagement. SurGE presents a framework that computes surrogate gradients of the design objective through a differentiable pipeline consisting of a kinodynamic single-rigid-body (Kino-SRB) model and a designaware control policy, and injects them into CMA-ES via mean shift with cosine-annealed step decay. This method achieves 6× lower cross-seed standard deviation and 18% tighter population concentration compared to vanilla CMA-ES, while matching or improving the best objective.

The gist

SurGE is a hybrid framework for the legged robot codesign problem that integrates the surrogate gradient with CMA-ES to achieve faster convergence and reduced variance by using a differentiable surrogate dynamics model to estimate gradient information along the dynamics rollout, which aids CMA-EA in mean value shifting.

Problem Formulation and Co-design Strategy

The co-design problem is formulated as a standard Optimal Control Problem (OCP) where decision variables include states, control inputs, and design parameters. A common approach involves a bilevel optimization strategy: the outer level optimizes the design parameters with a fixed control policy, while the inner level optimizes the control policy for each candidate design. To address computational efficiency, this is often reformulated as a single-stage optimization problem by assuming that a design-aware optimal controller approximates the solution of the inner optimization by mapping design parameters to control policies. This leads to a closed-loop system where dynamics are simplified as a transition function that only depends on the last state and design parameter: xk+1 = f(xk,π∗(xk, d), d).

SurGE Framework Components

The SurGE framework integrates three main components to handle the non-differentiability of legged robot dynamics:

  1. A differentiable surrogate dynamics model: This is a kinodynamic single-rigid-body (Kino-SRB) model that captures the essential centroidal dynamics and incorporates the UPS torque as a smooth function of design parameters, providing a differentiable path from d to L.

  2. A design-aware control policy: A neural network policy, often based on Reinforcement Learning (RL), is used as the controller. To make it design-aware, normalized design parameters are encoded via a multi-layer perceptron (MLP) and concatenated with the observation to form the policy input, utilizing a teacher-student architecture with regularized online adaptation (ROA) for sim-to-real transfer.

  3. Gradient injection into CMA-ES: The surrogate gradient is computed through this differentiable pipeline and injected into CMA-ES via mean shift with cosine-annealed step decay. The mean shift direction is calculated using the preconditioned covariance matrix and the surrogate gradient to guide the search toward promising regions, while a cosine-annealed schedule scales this injection rate to emphasize guidance in early generations.

Surrogate Gradient Computation and State Alignment

To compute the gradient of total cost with respect to design parameters, a chain rule is applied through time, yielding an expression involving Jacobians of the closed-loop dynamics and the control policy. Since the full dynamics are non-differentiable, SurGE utilizes a differentiable state x diff that is realigned at every time step with the state produced by a higher-fidelity non-differentiable simulator. This ensures that each one-step prediction starts from a realistic state, limiting gradient accumulation to single-step horizons and preventing drift from model mismatch while still providing informative directional signals. The surrogate model itself is often either a subset or function of the full state, ensuring only the dominant dynamics are propagated through differentiable dynamics for the gradient computation.

Evolutionary Search and Performance

The framework employs CMA-ES as the outer loop search algorithm. At each generation, candidate solutions are sampled from a multivariate normal distribution over the design space, and their fitness is evaluated on a non-differentiable simulator to obtain the true objective L. The core innovation is incorporating the surrogate gradient ∇dL into this process via mean shift: the CMA-ES mean is shifted in the negative gradient direction before sampling. This shift is preconditioned by the current covariance matrix and scaled to match the expected Mahalanobis step length. To mitigate bias accumulation, a cosine-annealed injection rate is used, which scales the mean shift ηg = η0 / (2(1 + cos(πg/G))), emphasizing gradient guidance early on and gradually switching control to CMA-ES adaptation as the distribution contracts around a local optimum. Hardware experiments on the MUPS v2 robot confirmed that SurGE reduces the design objective by 37.65% compared to a hand-tuned initial design, demonstrating that the optimization direction identified in simulation transfers consistently to the physical system.

Improvements for AI systems

Here are specific improvements for AI systems based on the SurGE framework, detailing what these improved systems can achieve:

  1. Improved Co-design of Legged Robots with Elasticity: The system will be capable of optimizing complex physical designs (like parallel elastic actuators) and control policies simultaneously to maximize energy efficiency during locomotion.

  2. Accelerated Design Optimization for Non-Differentiable Problems: The system will use surrogate gradients to navigate non-differentiable design landscapes, allowing it to find near-optimal physical configurations much faster than traditional black-box evolutionary methods like vanilla CMA-ES.

  3. Enhanced Robustness and Reproducibility in Search: By incorporating cosine annealing for gradient injection, the AI search process will be steered toward promising regions early on and then gradually transitions to robust population adaptation, resulting in significantly lower cross-seed standard deviation (e.g., 6x lower) and more reproducible hardware designs.

  4. Hardware-Aware Design Transfer: The system will refine initial hand-tuned designs into high-performing physical prototypes. It can achieve significant objective reduction on real hardware (e.g., 37% improvement in the example) by consistently transferring optimization trends from simulation to physical deployment, overcoming the sim-to-real gap.

  5. Optimized Control Policy Generation: The system leverages a design-aware control policy, pre-trained via RL and adapted via a teacher-student architecture (ROA), allowing it to generate locomotion controllers that are inherently optimized for specific physical designs across the entire design space, amortizing the cost of per-design inner loop optimization.

  6. Efficient Gradient Computation for Complex Dynamics: The system utilizes a differentiable surrogate model (Kinodynamic SRB) coupled with real-time state alignment via non-differentiable simulators to compute useful gradient information despite the inherent discontinuities of hybrid robot dynamics, enabling gradient-based acceleration where it is applicable.

In summary, this improved AI system can perform high-stakes robotics engineering by autonomously discovering and validating optimal physical hardware designs and control laws for legged robots with elastic elements in a computationally efficient and robust manner.

Sources

Related papers