Optimize, Learn, Refine: Whole-Body Grasping and Pick-and-Throw with a Spiral Soft Robot

arXiv:2609.38202 · cs.RO · Submitted 2026-09-22 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "Optimize, Learn, Refine".

Dev: Soft continuum robots can exploit distributed compliance for whole-body manipulation, but synthesizing behavior through changing contacts remains difficult.

Rosa: First, who's behind it and why it matters.

Paper summary: Rosa: So to recap where we are, we're discussing how this paper proposes using an outcome-based optimization approach for whole-body grasping and pick-and-throw with a spiral soft robot. The central thesis is that instead of prescribing contact forces or locations, the robot generates its entire motion through optimizing tendon commands based on desired outcomes.

Dev: Right, it’s about synthesizing behavior through changing contacts, which they found difficult before this work was done. They claim that by defining objectives like tip angular sweep and body–object enclosure for grasping and incorporating release direction alignment and minimum release speed for throwing, the robot can achieve these complex behaviors without needing explicit contact force specifications.

Taro: The paper is significant because it demonstrates that this framework can yield whole-body grasping success rates of ninety-eight point four percent in simulation and a perfect one hundred percent success rate on hardware for pick-and-throw tasks.

Rosa: That level of performance across both simulated and physical environments is what makes this work noteworthy, showing the effectiveness of their outcome-based optimization strategy when applied to soft continuum robots like SpiRob.

Dev: The reason it matters is that it shows a systematic way to derive complex behaviors from desired end states, which could change how we approach whole-body manipulation in robotics generally.

Taro: It suggests that for these types of systems, the complexity of the interaction itself can be managed by optimizing the actuation space rather than trying to hardcode every single possible state transition.

Rosa: That makes sense when you consider how much different contact configurations there are; they’re managing that vast search space by focusing only on what matters for reaching a specific goal.

Dev: And the framework moves beyond just solving one problem; it sets up an "optimize–learn–refine" loop, implying that this method is designed to be adaptive across different task conditions.

Taro: It’s not just about solving a single instance; it's about building a system capable of adapting its motion synthesis based on the specific context of the interaction it's currently in.

Rosa: So, we see they are using this method to show that we can get high-quality, complex manipulation from soft robots by letting the system learn how to bridge that gap between desired outcomes and physical execution.

Dev: And that’s the core idea—using outcome costs as the guide for optimization rather than following a predetermined path of control inputs. Now, let's talk about what this means for us in terms of broader applications later on.

Taro: That adaptability is key when we think about real-world scenarios where things go wrong; if the robot can learn to adjust its strategy based on unexpected feedback, that opens up possibilities for much more robust autonomous interaction.

Conclusion: Rosa: Wrapping up our discussion on "Optimize, Learn, Refine: Whole-Body Grasping and Pick-and-Throw with a Spiral Soft Robot," we’ve seen how the authors used an outcome-based optimization loop to synthesize complex whole-body actions. The authors are Marwah Basuhai, Tingcong Liu, Ibrahim Alsarraj, Yuhao Wang, and Ke Wu.

Dev: Indeed, that framework is quite sophisticated because it marries derivative-free optimization with a neural predictor to create a warm start for their CMA–ES optimizer for the new conditions they encounter.

Taro: The implication is that we are moving toward robots where the motion planning isn't just a sequence of pre-defined steps but something that actively evolves based on what it needs to achieve.

Rosa: That’s right, so it’s about giving these soft continuum robots a more intuitive way to handle the messy reality of physical interaction by letting the system optimize its own tendon commands for success.

Dev: It means we can expect systems that are better at handling unstructured environments because they aren't rigidly following a pre-set trajectory, which is something engineers like me think about constantly regarding failure modes.

Taro: If this method scales well, it could mean autonomous agents can interact with objects in ways that are far more nuanced and less prone to simple errors than current methods allow.

Rosa: It’s a powerful demonstration of how outcome-based synthesis works in practice on physical platforms, proving that this approach is viable for achieving high-quality manipulation goals.

Dev: The success rates they reported across both simulation and hardware scenarios give us concrete metrics to judge the practical viability of this method when we bring these soft robots out of the lab and into a more dynamic setting.

Marwah Basuhai, Tingcong Liu, Ibrahim Alsarraj, Yuhao Wang, Ke Wu

cs.RO

Submitted: 2026-09-22

Updated: 2026-10-01

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 83/100

The gist: Soft continuum robots can exploit distributed compliance for whole-body manipulation, but synthesizing behavior through changing contacts remains difficult.

Key concepts

Spiral Soft Robot (SpiRob)
This is the specific soft robot platform used for the experiments. It is a continuous robot that utilizes distributed compliance, allowing it to manipulate objects using tendon commands. The research focuses on using this platform to learn complex whole-body manipulation skills.
Outcome-based Optimize–Learn–Refine Framework
This is a three-step method for learning complex movements. First, CMA-ES optimizes parameters for a task outcome; second, a neural model learns an initial guess based on the task condition; and third, CMA-ES refines that guess using the learned initialization. This cycle synthesizes contact behavior.
Jenc(te) Cost Function
This cost function measures how well the robot achieves a 'compact terminal enclosure' during grasping. It balances two geometric proxies: 'Tip angular sweep,' which tracks movement around the object, and 'body–object proximity,' which measures how close the robot's body is to the object's center.
Actuation Parameterization
The continuous tendon commands are converted into a finite set of parameters (p) that can be searched by optimization algorithms. This involves decomposing the commands into shared components, differential components controlled by scaling factors ($ε$), and smooth modulation functions ($g(t)$). This allows the complex physical problem to be solved as a search over these simpler parameters.

Terminology

Summary

Soft continuum robots can exploit distributed compliance for whole-body manipulation, but synthesizing behavior through changing contacts remains difficult. The gist: An outcome-based optimize–learn–refine framework using CMA–ES and a neural model synthesizes whole-body grasping and pick-and-throw from an ungrasped state by optimizing tendon commands based on task outcomes. This work is significant because it demonstrates the effectiveness of this framework across both simulated and physical whole-body manipulation tasks, achieving high success rates in grasping (98.4% in simulation) and pick-and-throw (100% on hardware).

Problem Formulation and Task Objectives

The paper addresses whole-body grasping and pick-and-throw from an initially ungrasped state using the spiral soft robot SpiRob as the platform. The tasks require a finite-horizon tendon trajectory that achieves specific outcomes. For whole-body grasping, the objective is to form a compact terminal enclosure. This enclosure is quantified by two complementary geometric proxies: Tip angular sweep (capturing progression around the object) and body–object proximity (measuring the mean body-to-object-center distance). The enclosure cost, denoted as Jenc(te), balances these two factors: A larger positive value indicates greater progression of the tip around the object while Smaller values indicate that the robot body remains closer to the object. For pick-and-throw, this enclosure objective is maintained immediately before release, and it is combined with directional requirements.

Actuation Parameterization

The continuous tendon-command trajectory, defined by inputs u(t) = [uL(t), uR(t)], is parameterized into a finite-dimensional search space p. This parameter vector captures the coordinated shared and differential variation of the two tendon commands. The decomposition involves:

  1. Shared components: b(t) and s(t), which change both tendon commands together.

  2. Differential components: controlled by κ (scaling the differential component), ϵ (constant differential bias), kw, and cw (controlling time warping).

  3. Smooth modulation function g(t): parameterized by vectors ρ and a, which shape the differential component through a weighted combination of five smooth phase functions.

The final tendon commands are constructed as uL(t; p) = clip[0,1] (b(t) + κg(t)s(t) + ϵ), and similarly for uR(t; p). This parameterization converts the original functional problem into a finite-dimensional search over p, leading to the objective: u⋆(·) ∈ arg min u(·) J q(·), Po(·) s.t. S q(·), Po(Po, ·), u(·)= 0.

Optimize–Learn–Refine Strategy

The core methodology employs an optimize–learn–refine framework to synthesize contact-rich behavior across varying task conditions.

  1. Optimization: For each task condition ξ, the actuation parameters are optimized using CMA–ES. This involves generating candidate parameter vectors through complete simulation rollouts and retaining the best solution as pCMAi.

  2. Learning (Warm Start): A neural predictor, pb = fθ(ξ), is trained via supervised regression to estimate an actuation-parameter initialization based on the task condition ξ.

  3. Refinement: For a new condition, CMA–ES uses this learned initialization (m0 = pb) and a constant initial state (C0 = I) to refine the initialization according to the corresponding task cost, returning pCMAref. This process is repeated over M task conditions to produce D = ξi, pCMAi M i=1.

Outcome Costs and Performance Metrics

The framework defines distinct outcome costs for each task:

((

Whole-body grasping cost): Evaluated at the end of the horizon: Jgrasp = Jenc(T).

**(Pick-and-throw cost): Evaluated immediately before release, combining enclosure with release dynamics. It is defined as Jthrow = Jenc(t−rel) + wdirAdir + wvelAvel. The directional term (Adir) ensures the velocity aligns with the desired direction dˆdes, and the speed term (Avel) enforces a minimum release speed vmin. Success in pick-and-throw requires a release-direction error below 5◦ and a release speed no lower than the prescribed vmin. Hardware experiments confirmed 100% success for both whole-body grasping and pickand-throw. The learned initialization was shown to increase grasping success from 78.6% to 98.4% while reducing the median rollout count from 1184 to 816 in CMA–ES. This demonstrates that the neural model provides a task-conditioned starting point that places CMA–ES closer to a useful solution region.

Improvements for AI systems

Here are the specific improvements to AI systems based on this scientific paper, and what those improved systems can achieve:


The core improvement lies in developing a robust, outcome-based framework for synthesizing complex, contact-rich behaviors in soft robotics (like whole-body grasping and pick-and-throw) by decoupling the optimization of actuation from the specific task conditions.

Here are the specific improvements:

  1. The development and application of an optimize–learn–refine strategy utilizing a derivative-free global optimizer (CMA–ES) coupled with a neural network for warm-start initialization.

  2. A shared, outcome-based cost function formulation that integrates geometric proxies (tip angular sweep and body-object proximity) into the objective for both grasping and pick-and-throw tasks.

  3. The use of a compact, parameterized representation of tendon commands (decomposed into shared/differential components modulated by smooth phase functions) to convert the continuous control problem into a finite-dimensional search space for CMA–ES.

  4. Training a supervised regression neural model to predict an effective initial actuation parameter vector based on the task condition, allowing subsequent CMA–ES refinement to converge faster and more reliably.

The improved AI system (the Optimize–Learn–Refine Framework) can achieve the following specific capabilities:

  1. A robot can perform complex, dynamic manipulation tasks (like grasping an object from an ungrasped state and then throwing it along a prescribed direction at a minimum speed) using only compliant, soft-body actuation, without needing pre-programmed contact forces or explicit body configurations.

  2. The system can generalize its learned behaviors across different initial object positions and task requirements by leveraging the neural warm-start initialization. This means the robot doesn't need to be retrained from scratch for every new target location or desired throw direction; it only needs a few rapid refinement steps (CMA–ES) based on a task-conditioned starting point.

  3. The AI can achieve near-perfect performance in contact-rich scenarios:

Choose the robot configuration that simultaneously achieves maximum object enclosure (minimizing tip angular sweep and maximizing body proximity) while ensuring the resulting release trajectory aligns with a precise target direction and speed requirement, even when the contact evolution is nonsmooth.

  1. The system can be trained efficiently using simulation data, achieving high success rates (e.g., 98% grasping success in simulation) with significantly reduced computational effort compared to traditional optimization methods (reducing median rollout counts from 1184 to 816).

  2. The framework allows for the direct mapping of abstract task outcomes (grasp enclosure and directional release) into physical actuation commands, effectively synthesizing contact-rich behaviors that emerge naturally from compliant interaction rather than being explicitly prescribed.

Sources

Related papers