SurGE: Surrogate Gradient-guided Evolution for Co-design of Legged Robots with Parallel Elasticity
summary
The gist
Co-design of legged robots with elastic elements is challenging due to the non-differentiability of contact dynamics and mechanism engagement.
In short
SurGE is a framework for co-designing legged robots with elastic elements by overcoming the non-differentiability of contact dynamics. It uses a differentiable surrogate model and an RL policy to compute gradients, which are then injected into CMA-ES. This hybrid approach significantly reduces optimization variance and improves convergence speed compared to standard methods.
Key concepts
- Surrogate Gradient
- A mathematical approximation of the true gradient of the robot's performance objective. Since the actual dynamics are hard to differentiate, this model creates a smooth, differentiable path from design parameters to performance metrics, allowing gradients to be calculated even in complex systems.
- Kino-SRB Model
- A kinodynamic single-rigid-body model used as the surrogate dynamics. This model captures the essential movement of the robot's center of mass and incorporates elastic forces smoothly based on design parameters, providing a differentiable way to simulate how design changes affect motion.
- CMA-ES with Mean Shift
- CMA-ES is an evolutionary algorithm that searches for optimal robot designs. SurGE enhances it by using 'mean shift,' which nudges the search distribution in the direction of the calculated surrogate gradient. This guides CMA-ES to promising design regions much faster than standard evolution.
- Design-Aware Control Policy
- A neural network policy that controls the robot's movement, designed to be aware of its physical design parameters. By encoding normalized design parameters into the policy input, this controller ensures that the learned control strategy adapts effectively to different robot configurations.
Terminology used across episodes
This episode discusses
- SurGE: Surrogate Gradient-guided Evolution for Co-design of Legged Robots with Parallel Elasticity · Paper Radio
- An Introduction to Zero-Order Optimization Techniques for Robotics
- Efficient Adjoint-based Design Optimization with Optimal Control
- Gradients are Not All You Need
- Proximal Policy Optimization Algorithms
The paper
SurGE: Surrogate Gradient-guided Evolution for Co-design of Legged Robots with Parallel Elasticity · Read on arXiv
University of Michigan
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "SurGE: Surrogate Gradient-guided Evolution for Co-design of Legged Robots with Parallel Elasticity".
Dev: Co-design of legged robots with elastic elements is challenging due to the non-differentiability of contact dynamics and mechanism engagement.
Rosa: First, who's behind it and why it matters.
Title and authors: Rosa: Let's talk about the title and who came up with this work, SurGE: Surrogate Gradient-guided Evolution for Co-design of Legged Robots with Parallel Elasticity. It’s pretty descriptive, highlighting the core mechanism they developed to solve that non-differentiability issue we talked about earlier.
Dev: The authors listed are Zhuang, Qin, Lu, Shen, Wang, He and Ding; it looks like a solid group coming from different areas of robotics and control engineering. I'm curious if their collective expertise really covers both the control policy side and the dynamics modeling aspect well enough to tackle this specific co-design problem.
Taro: I’m more interested in what that title implies about the scope; "Co-design" suggests they aren't just optimizing a controller or just a robot body, but linking those two parts together through this surrogate gradient technique.
Rosa: Exactly, Taro; it points to the integration aspect—they are trying to optimize the physical structure and the control policy simultaneously in one go, which is inherently more complex than doing them separately.
Dev: From an engineering standpoint, that co-design challenge is exactly why we need this; treating it as a single optimization problem where design parameters map directly to control policies via a differentiable surrogate model seems like a very efficient way to approach the computational expense.
Taro: I think the implication here is that they are addressing the practical bottleneck in legged robotics, which is often how you get an elastic structure to move reliably without breaking or losing stability during locomotion.
Rosa: That’s right; it addresses that physical reality where elasticity introduces complexities into the contact and mechanism engagement models that standard optimization tools struggle with because they aren't smooth.
Dev: The potential impact on the world, if this works reliably outside of a highly controlled lab setting, is huge for applications requiring compliant locomotion, like search and rescue or navigating unstructured terrain.
Taro: If this framework can consistently transfer its learning to physical hardware as the paper suggests, it moves us closer to deploying robots in real-world scenarios where the environment is messy and unpredictable.
Rosa: It certainly points toward a future where complex physical systems are designed using AI methods that can handle the inherent complexities of their own physics rather than requiring manual, iterative tuning for every single design iteration.
Dev: I just hope they’ve accounted for the latency issues we always worry about in these closed-loop systems; if the surrogate model is too slow, that whole gradient injection pipeline becomes useless regardless of how good the theory is.
Taro: That’s a fair point on execution speed; the theoretical efficiency must translate into a practical execution speed that isn't bottlenecked by computation when we need rapid adaptation during operation.
Rosa: So, while they solved the design optimization problem in simulation, the big question for us is how long these designs remain stable and performant when faced with real-world wear and tear.
Dev: That’s what I want to know; we need to see if this framework can maintain those tight concentration metrics after several cycles of simulated or actual deployment.
The paper's summary: Rosa: The paper summarizes SurGE by explaining that it creates a differentiable pipeline using a kinodynamic single-rigid-body model and a design-aware control policy to compute surrogate gradients, which are then injected into CMA-ES via mean shift with cosine-annealed step decay.
Dev: Essentially, they're replacing the expensive inner loop optimization—optimizing the control policy for every candidate design—with a single model that predicts that optimal solution beforehand, which is a smart way to save computation time.
Taro: That replacement of the per-design inner optimization with an amortized training model is what I find most interesting from an autonomy perspective; it simplifies the overall process dramatically.
Rosa: It does simplify things by allowing the outer design search to proceed with a policy that’s held fixed, instead of retraining a controller for every new physical configuration we test.
Dev: The methodology also involves aligning states at every time step using a non-differentiable simulator to compute those surrogate gradients, which is necessary because the full dynamics aren't differentiable.
Taro: That state alignment step is the bridge between the smooth surrogate model and the hard reality of the simulator, and it seems crucial for getting meaningful information from those gradients despite the underlying physics being discontinuous.
Rosa: And they use this information to guide CMA-ES through a mean shift process, using a cosine-annealed schedule to manage how aggressively that guidance is applied during different generations of the search.
Dev: That specific mechanism for injecting the gradient into CMA-ES is sophisticated; it’s not just dumping the gradient in and hoping for the best, but scaling it based on both the covariance matrix and that annealing schedule.
Taro: This structured approach to gradient injection suggests a very careful balance between exploration—letting CMA-ES explore widely—and exploitation—using the guidance to focus on promising areas.
Rosa: Overall, they've managed to tackle the non-differentiability of legged robot dynamics by creating a differentiable pipeline for gradient computation and then feeding that information into an evolutionary search algorithm.
Dev: So, they’ve successfully created a hybrid framework where the simulation provides the necessary differentiability for gradient calculation, while CMA-ES handles the broader design search structure.
The paper's improvements: Rosa: The main improvements discussed are that SurGE moves beyond just using standard black-box methods like vanilla CMA-ES by incorporating surrogate gradients to guide the mean shift in the search distribution.
Dev: Specifically, they achieve six times lower cross-seed standard deviation and eighteen percent tighter population concentration compared to vanilla CMA-ES on their four-DOF hopping robot design space.
Taro: The quantitative results on those metrics are what really sell this framework; those numbers show a tangible improvement in the quality of the designs being proposed, not just theoretical correctness.
Rosa: Beyond the simulation, they also demonstrate that starting from a hand-tuned initial design, SurGE reduces the design objective by thirty-seven point six five percent on hardware experiments in a two-dimensional design subspace <ref:2606.21866#pg0,that starting from a hand-tuned initial design, SurGE reduces the design>.
Dev: That hardware result is what really validates the transfer of optimization trends from simulation to physical deployment; it proves the simulation isn't just generating plausible but ultimately useless designs.
Taro: The consistency between the simulated guidance and the physical performance makes this approach much more robust because it overcomes a major hurdle in robotics: ensuring that what works in a simulator actually works on the robot.
Rosa: Furthermore, they incorporate a design-aware control policy trained via Reinforcement Learning and adapted through a teacher-student architecture with regularized online adaptation to help bridge the sim-to-real gap.
Dev: That teacher-student architecture for the policy is clever because it allows us to generate controllers that are inherently optimized for specific physical designs, which should amortize the cost of needing a separate inner loop optimization later.
Taro: Amortizing that per-design optimization cost is a massive win if we consider the entire design space; it means we can explore more configurations without having to solve an inner optimization problem every single time.
Rosa: So, in essence, the improvements are about using gradient information from a differentiable surrogate model to steer an evolutionary search toward better designs that perform better both in simulation and on physical hardware.
Dev: It’s a solid methodology because it combines the gradient guidance with the annealing schedule to manage exploration and exploitation effectively during the search process, which is something vanilla CMA-ES just can't do on its own.
Conclusion: Rosa: To wrap up, SurGE provides a method for co-designing legged robots with parallel elasticity by using a differentiable pipeline to compute surrogate gradients and injecting them into CMA-ES via mean shift with cosine-annealed step decay.
Dev: The key takeaways are that this technique yields six times lower cross-seed standard deviation and eighteen percent tighter population concentration compared to vanilla CMA-ES, and it shows a thirty-seven point six five percent reduction in the design objective on hardware testing.
Taro: The implication is that we have a framework that can more effectively search for complex physical designs by leveraging gradient information derived from differentiable models, which is essential when dealing with the non-smooth nature of real robot dynamics.
Rosa: It really points toward a future where we can autonomously discover and validate optimal physical hardware configurations for legged robots in a computationally efficient way, relying on the SurGE approach to handle the challenges of elasticity.
Dev: I just have to stress that while they show significant simulation-to-real transfer, we still need more data on how long these designs maintain their performance when faced with real-world wear and tear during extended operation.
Taro: That's a critical consideration; if we want this to be truly useful for field deployment, the durability of the designs needs to be rigorously tested over long operational timescales.
Rosa: So, that’s our discussion on SurGE, and it’s been really insightful looking at how they combine surrogate gradients with CMA-ES for robot co-design.
More episodes
- 2610.10846-Cross-Embodiment Robot Foundation World Models with Latent Actions
- 2610.10601-Teaching a Robot Dog New Tricks: Diverse Quadruped Skills via Combined Reinforcement and Imitation Learning with Adversarial Task Selection
- 2610.10637-TacHair: Tactile Contact-Distribution Guided Online Correction for Robotic Hair Stroking and Perception
- 2610.10646-Masked Generative Motion Planning with Geometry-Guided Token Search
- 2610.10812-Skill-SLM: Agent Skill-driven Small Language Models for Reliable Robot Operation
- 2610.10801-Same Action, Different Outcome: Variability in Dynamic Cloth Manipulation
- 2610.10810-Diagnosing and Recovering from Observation-Space Shift at Long-Horizon Skill Seams
- 2610.10748-TAPNAV: Humanoid Navigation through Tactile Active Perception
- 2610.10855-OmniHOI: Dexterous Hand-Object Interaction from Monocular Human Video
- 2610.11003-ActiveReg: Information-Driven Active Regional Probing for Partial-to-Full Bone Registration