Generalizing deep reinforcement learning across cable-driven parallel robot configurations with actuator-level policies

arXiv:2608.07546 · cs.RO, cs.AI · Submitted 2026-07-31 · Read on arXiv

Abir Bouaouda, Mohamed Boutayeb, François Charpillet, Dominique Martinez, Rémi Pannequin

Université de Lorraine · inria · CNRS · International University of Rabat · Aix-Marseille Université

cs.RO, cs.AI

Submitted: 2026-07-31

Updated: 2026-08-11

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Terminology

Summary

Summary

This paper introduces a novel deep reinforcement learning (DRL) approach for controlling cable-driven parallel robots (CDPRs) that is independent of the specific robot configuration. The proposed method, termed actuator-level policy (ALP), trains a policy that controls each motor individually to achieve its target cable length, in contrast to conventional DRL approaches that learn to control the entire robot to reach a desired end-effector position. The authors state: To the best of our knowledge, this is the first work to apply DRL to control CDPRs using an actuator-level policy.

The approach offers two main advantages: (i) a single shared policy can be applied to any CDPR configuration, regardless of actuator count, and (ii) reliance on inverse kinematics, avoiding the more challenging forward kinematics problem. Training is performed in simulation, and the learned policy is successfully transferred to a real CDPR.

The control problem for a CDPR consists of determining the torques to apply to the motors to move the end-effector to a desired 3D position. The inverse kinematics problem (computing cable lengths from end-effector position) is straightforward, while the forward kinematics problem (computing end-effector position from cable lengths) is more complex. The authors assume a black-box model that takes desired motor speeds and current robot state as input and outputs the next state.

In conventional reinforcement learning for CDPRs, the state describes the entire robot: st = (Xt, Xt−1, et, et−1, It) where Xt is the position of the end-effector, et is the error between the current and target positions, and It is the motor current. All observations and actions are normalized to [−1, 1]. Three algorithms are used for comparison: DDPG, PPO, and SAC.

For the ALP approach, the state of each actuator is defined as: stj = (Ljt−3, Ljt−2, Ljt−1, Ljt, Ljt,target, Itj) where Ljt is the length of cable j at step t, Ljt,target is the target length of cable j, and Itj is the current of motor j. Including the cable length at step t−3 provides information about cable dynamics, tension, and speed. Since there are n actuators, at each step t, the agent performs n interactions with the environment but learns a single shared policy. Thus, only one neural network maps actuator state to torque.

The reward function has three components: error-based reward (r1,t = −∥et∥), energy optimization reward (r2,t = −∥It∥2), and current limits reward (r3,t = −∥at − u sat,t∥). The final reward is: rt = α1 r1,t + α2 r2,t + α3 r3,t with coefficients α1 = 1, α2 = 0.001, and α3 = 0.5. The energy term improves efficiency, while the current limits term smooths the control signal.

For ALP, two reward variants are tested: reward 1 (r1) is collaborative, where all actuators receive the same reward based on overall position error; reward 2 (r2) is individual, based on each actuator's own error: rtj = − α1 etj − α2 Itj − α3 atj − u sat,t j where etj is the error between current and target lengths of cable j.

Target trajectories are generated randomly with random speeds and accelerations, respecting workspace boundaries and dynamic limits. When the end-effector reaches workspace limits, the desired speed is randomly reversed to keep it within bounds.

Training was performed only on a 4-cable translational robot with point-mass mobile platform. Results show that PPO does not converge, while DDPG and SAC both converge, with DDPG achieving higher and more stable average reward. Therefore, DDPG was selected for further experiments.

Testing results show that ALP outperforms conventional RL: the error ranges from 0 to 0.5 m for the conventional RL approach and from 0 to 0.25 m for the ALP RL approach. The policy using reward 1 (cooperative) outperforms reward 2 (individual), possibly due to suboptimal tuning of coefficients for the individual reward.

The ALP approach successfully generalizes to different configurations: the ALP reinforcement learning approach successfully tracks the desired trajectory in the new configuration, even though it was trained only on a 4-cable robot with a different setup. It also transfers to an 8-cable robot: the performance of both policies on this configuration is similar to their performance on the initial configuration, demonstrating the transferability of the policy learned on the 4-cable robot to the 8-cable robot.

The policy was tested on a real 8-cable translational robot: the agent is able to track the desired trajectory with high accuracy, and the error between the target position and the end-effector position is less than 0.04 m.

The ALP approach is robust to errors in position estimation because it does not rely on estimating the end-effector position: The control signal for each actuator is generated independently of the lengths of the other cables, allowing the system to remain robust even if there are errors in the measurement of one cable.

The paper includes an appendix with a proof of concept using a dynamic model of a single cable, showing that a control law can be learned by a neural network using reinforcement learning, provided the actuator state includes cable length at previous time steps and motor current. Another appendix describes current constraints, ensuring cables remain in tension and currents stay within safe limits. A third appendix details trajectory generation in bounded workspaces.

Limitations acknowledged by the authors: "The comparison between different reinforcement learning algorithms remains inconclusive, mainly due to the manual tuning of hyperparameters... Similarly, tuning the reward function coefficients is challenging and does not guarantee optimal results. Another limitation is that the CDPR configurations considered in this paper are restricted to translational CDPRs with a point mass mobile platform."

The paper concludes: "In this work, we have proposed a new reinforcement learning-based control approach for cable-driven parallel robots using an actuator-level policy. This method enables systematic generalization, allowing control of any cable robot configuration. Simulation results demonstrate strong performance compared to conventional reinforcement learning approaches, and the proposed method has been validated on a real cable robot, confirming its transferability to different configurations and to real-world hardware."

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to AI systems and what the improved system can do:

1. Actuator-Level Policy Architecture (ALP)

  • Replace the conventional robot-level policy (which maps full robot state → all motor torques) with a shared, single actuator-level policy (which maps individual actuator state → individual motor torque)

  • Use a multi-agent formulation with centralized training and decentralized execution, where each actuator is treated as an independent agent sharing identical network weights

  • Include in the actuator state: cable length at steps t−3, t−2, t−1, t; target cable length; and motor current (all normalized to [−1, 1])

2. Reward Function Design

  • Use a composite reward: r = −α1‖e‖ − α2‖I‖2 − α3‖a − u sat‖ with coefficients α1=1, α2=0.001, α3=0.5

  • For individual rewards (reward 2), compute per-actuator error as the difference between current and target cable length, not end-effector position error

  • Include a current-limit penalty term that penalizes actions outside the saturation bounds to smooth control signals

3. Training Strategy

  • Train exclusively on a 4-cable planar CDPR with point-mass end-effector, then transfer the policy to any other configuration (8-cable, 3D, different workspace sizes) without retraining

  • Generate training trajectories using random acceleration sampled from a Gaussian distribution, with velocity clipping and random velocity reversal at workspace boundaries to ensure full state-space exploration

  • Use DDPG as the preferred algorithm (PPO failed to converge; SAC was less stable)

4. Control Signal Computation

  • Compute desired motor speed as: u = (i sat / K motor) + v m where i sat is the saturated current, K motor is the motor constant, and v m is measured motor speed

  • Apply current saturation to enforce cable tension (minimum current) and prevent motor damage (maximum current)

1. Cross-Configuration Generalization

  • Control any CDPR with any number of actuators (tested from 4 to 8) using a single policy trained on just one configuration

  • Transfer to robots with different actuator positions, workspace sizes, and even different dimensions (2D → 3D) without additional training

  • Handle changes in end-effector mass (0.5–2 kg) without performance degradation

2. Robustness to Sensor Errors

  • Maintain trajectory tracking accuracy even when individual cable length measurements are incorrect or cables sag, because the policy does not rely on forward kinematics or end-effector position estimation

  • Operate without motion capture systems or direct end-effector position sensors

3. Real-World Deployment

  • Achieve tracking error < 0.04 m on a real 8-motor CDPR using a policy trained only in simulation on a 4-motor planar robot

  • Produce smooth, non-oscillatory control signals suitable for real hardware (due to the current-limit penalty term)

4. Sample Efficiency and Training Speed

  • Reduce training time by learning a single shared policy from n interactions per step (one per actuator) rather than one interaction per step for the whole robot

  • Achieve stable convergence with DDPG in approximately 6–10 hours, compared to 50–100 hours for conventional RL approaches

5. Task Flexibility

  • Track arbitrary trajectories (not just point-to-point) at speeds up to 2 m/s with error < 0.25 m

  • Operate in bounded workspaces with dynamic constraints respected (acceleration limits, velocity limits, current limits)

Sources

Related papers