Learning Low-Frequency Motion Control for Robust and Dynamic Robot Locomotion

arXiv:2209.14887 · cs.RO, cs.AI · Submitted 2022-09-29 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "Learning Low-Frequency Motion Control for Robust and Dynamic Robot Locomotion".

Dev: Robotic locomotion can be achieved robustly and dynamically even when using motion controllers operating at very low frequencies.

Rosa: First, who's behind it and why it matters.

Title and authors: Rosa: Let’s talk about the title and the authors of this paper, "Learning Low-Frequency Motion Control for Robust and Dynamic Robot Locomotion." It really highlights their core contribution: showing that you don't need to chase high frequencies to get dynamic movement from a quadruped.

Dev: I agree; it sounds like they’re shifting the focus from raw reactivity to smarter planning capabilities in the control loop.

Taro: The authors seem very focused on proving this claim through empirical evaluations, which is important because many of these claims are just theoretical until you actually see them tested on a physical robot <ref:2209.14887#pg0>.

Rosa: I think the implication here is that we might be over-engineering our control systems by trying to force them to run at frequencies they don't need for stable locomotion, which saves a ton of computational power.

Dev: From an engineering standpoint, if we can get good performance at eight Hz instead of two hundred Hz, the system has more time to settle between commands, which inherently reduces the impact of unavoidable hardware latency <ref:2209.14887#pg1>.

Taro: I think that means autonomy becomes more resilient when things go wrong because the controller isn't constantly fighting against slow response times or noise in the dynamics.

Rosa: It really shifts the paradigm from purely reactive control to a more deliberative approach, which is a big deal for deployment outside of highly controlled labs <ref:2209.14887#pg0>.

The paper's summary: Dev: So, summarizing what the paper actually does, they model the robot as a floating base with four limbs, and their motion controller policy is implemented as a multi-layer perceptron or MLP that maps observations to desired joint states <ref:2209.14887#pg2>.

Rosa: They introduce different types of policies depending on whether they are blind, perceptive, or use joint state history, which changes the input state space significantly <ref:2209.14887#pg3>.

Taro: I see them using history length H to augment the state space dimensionality by H times twenty-four joint states to help the policies better understand what’s happening at a local level <ref:2209.14887#pg3>.

Dev: The key finding they highlight is that low-frequency motion control policies are sufficient for achieving robust and dynamic quadrupedal locomotion, which suggests dynamics randomization or even full actuation modeling might not be necessary for successful sim-to-real transfer <ref:2209.14887#pg0>.

Rosa: That’s the big takeaway: the learned policy itself seems capable of handling the complexities of dynamics and real-world interaction without needing those extra, computationally expensive models <ref:2209.14887#pg0>.

Taro: If that holds true, it simplifies the deployment pipeline immensely; we might skip a lot of heavy offline modeling work to get a working robot controller on the ground <ref:2209.14887#pg1>.

The paper's improvements: Rosa: Now, let’s talk about the suggested improvements derived from this research, which really build on what they found by suggesting how to refine this low-frequency control approach <ref:2209.14887#pg3>.

Dev: One big improvement they suggest is moving away from high-frequency reactive architectures and adopting these low-frequency, planning-based control policies instead Improver one <ref:2209.14887#pg0>.

Taro: That means we stop trying to make the robot react instantly and start treating the policy more like a motion planner that generates targets based on context rather than just chasing the error right now Improver one <ref:2209.14887#pg0>.

Rosa: They also suggest integrating state history for enhanced observability, which allows policies to implicitly encode things like contact detection and actuation dynamics, even in blind settings Improver two.

Dev: That's interesting because incorporating that history helps manage the uncertainties they mentioned earlier, giving the AI a richer picture without needing explicit models Improver two.

Taro: I think this history inclusion is vital for handling unexpected situations, especially when things get messy or when terrain changes rapidly Improver two.

Rosa: And they also propose an adaptive policy training strategy where you test policies across a wide frequency range, rather than assuming one speed is always best for every task Improver three.

Conclusion: Dev: To wrap up the "Learning Low-Frequency Motion Control for Robust and Dynamic Robot Locomotion" paper, the main implication is that we can achieve robust and dynamic locomotion using motion control policies running at as low as eight Hz <ref:2209.14887#pg0>.

Rosa: This means we can expect to see robots perform consistently well on uneven terrain, like achieving a heading velocity of about one point five m/s, without needing to rely on high-frequency control for stability <ref:2209.14887#pg0>.

Taro: I think this means we can focus our research energy on making these low-frequency planners smarter and more adaptive to the unpredictable nature of the real world, rather than obsessing over tracking speed Improver three.

Dev: And from an engineering perspective, it means the system is less sensitive to actuation latencies—up to ninety milliseconds delay—which makes deployment much safer in practical scenarios <ref:2209.14887#pg0>.

Rosa: It really suggests that the ability of the learned policy to operate as a motion planner rather than a high-speed tracker is what unlocks this robustness and allows for successful sim-to-real transfer without needing dynamics randomization <ref:2209.14887#pg0>.

Oxford Robotics Institute

cs.RO, cs.AI

Submitted: 2022-09-29

Updated: 2026-10-02

Comments: 7 pages, 9 figures and 2 tables

Journal ref: IEEE International Conference on Robotics and Automation (ICRA) 2023

Project page: https://ori-drs.github.io/lfmc

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

Importance score: 90/100

The gist: Robotic locomotion can be achieved robustly and dynamically even when using motion controllers operating at very low frequencies.

Key concepts

Low-Frequency Motion Control (LFMC)
This refers to using motion control policies that operate at relatively slow frequencies compared to the robot's dynamics. The paper shows these policies can still achieve robust and dynamic walking, suggesting high-frequency control is not essential for successful locomotion.
Dynamics Randomization (DR) / Actuation Modeling
These are techniques often used in training to make a robot controller robust against uncertainties in the physical system's dynamics or how the motors actually move. The study found that these methods were unnecessary for achieving good results when using low-frequency motion control policies.
State Space Design
This involves defining what information (state) the motion controller uses to make decisions. The state includes robot position, orientation, joint angles, velocities, and sometimes terrain information. The design changes based on whether the policy is 'blind' or 'perceptive'.
Impedance Control
This is a control method where the robot interacts with its environment by controlling its relationship between forces and motion (like stiffness and damping). The paper uses this model to generate joint torques based on desired tracking gains ($k_p, k_d$) to follow the motion commands.

Terminology

Summary

Robotic locomotion can be achieved robustly and dynamically even when using motion controllers operating at very low frequencies. The key finding demonstrated in this work is that low-frequency motion control policies are sufficient for achieving robust and dynamic quadrupedal locomotion, suggesting that dynamics randomization or actuation modeling may not be necessary for successful sim-to-real transfer.

The Gist

Low-frequency motion control policies are sufficient to perform robust and dynamic locomotion, and dynamics randomization or actuation modeling may not even be necessary for successful sim-to-real transfer.

System Model and Control Architecture

The quadrupedal robot is modeled as a floating base B with four attached limbs. The robot state includes the global position of the base, denoted as rB ∈ R3 (position), and its orientation, represented by the rotation matrix RB ∈ SO(3). Each limb has three rotational joints with angular positions qj ∈ R12. The linear and angular base velocities are represented as vB ∈ R3 and ωB ∈ R3, respectively. Joint control torques τ j actuate the system using an impedance control model: τ j = kp∆qj −kdq˙ j, where kp and kd are tracking gains.

The control architecture is composed of a high-level motion controller and a low-level actuation tracker. The motion controller, executed at frequency fm, processes robot state information to generate desired joint states. The actuation tracker operates at frequency fa ≥ fm to track these desired joint states by generating torques using the model described in Eq. 1. The motion controller policy is modeled as a multi-layer perceptron (MLP), πθ: s 7→ a, which maps the input state tuple s to actions a ∈ R12.

Motion Control Policies and State Space Design

The motion control policies are denoted as πftM:H, where ft is the training frequency, M is the mode (b for blind or p for perceptive), and H is the history length of joint states introduced in the state tuple s. The state space design varies by policy type:

  1. For blind policies, πb:0, which has no joint state history (H=0), the state tuple sb:0 ∈ R48 is defined as sb:0:= hRTBez,qj,RTBvB,RTBωB,q˙ j,∆qj, c∗i. The objective of these policies is to track user-generated desired velocity commands.

  2. For perceptive policies, πp:0 (and others), the state space sp:0 augments sb:0 with robo-centric terrain information T ∈ R17×11 observed between [−0.8, 0.8]m along the heading axis and [−0.5, 0.5]m along the lateral axis with a resolution of 0.1 m.

  3. Joint state history augments the state space dimensionality by H × 24, where qtj represents joint positions and velocities recorded at time step tj, allowing for policies like πb:4 with a history length of 4.

Training Methodology

The problem is framed as a sequential Markov decision process (MDP) [38], aiming to maximize the expected cumulative discounted return J (π) = E∼πθ N ∑ t=0 γ t R t. Proximal Policy Optimization (PPO) [39] is used for training. To ensure consistency across different training frequencies ft, the discount factor γ can be computed by γ = exp(log 0.5 / ft × nγ0.5), where nγ0.5 = 3 s is used in this work for an episodic length N = 1 s to maintain a consistent batch size bs = ft × nenv across frequencies. Training times vary between 0.4 s and 1.5 s depending on the training frequency ft, with low-frequency policies converging faster (<10k iterations) than high-frequency ones. Notably, no dynamics randomization (DR) was performed while training the blind policies; an actuator network [26] was used to model real actuation dynamics during training, though this may not be necessary for LFMC.

Evaluation and Key Observations

Empirical evaluations show that LFMC policies are less sensitive to actuation dynamics under the assumption that actuation settling time is less than control step time. Furthermore, LFMC policies do not perform implicit modeling of system dynamics necessary for predictive control at high frequencies but instead can operate as motion planners. Qualitative analysis revealed that high-frequency policies exhibit extremely aggressive actuation tracking resulting in vibrations, whereas lower-frequency policies show reduced vibrations. Robustness to actuation latencies was also demonstrated, with LFMC policies showing higher robustness than high-frequency counterparts (Table I). In dynamic locomotion tests, the robot achieved a heading velocity of approximately 1.

Improvements for AI systems

Here are specific improvements to AI systems based on the findings of this paper, along with what those improved systems can achieve:


The core improvement lies in shifting from high-frequency, reactive control architectures to low-frequency, planning-based control policies. This addresses the limitations imposed by inherent system latencies and model uncertainties.

Here are the specific improvements and capabilities:

  1. Adoption of Low-Frequency Motion Control (LFMC) Policies:

  2. Mechanism: Instead of training policies at high frequencies (e.g., 200 Hz), train them at lower frequencies (e.g., 8 Hz or 15 Hz). The policy's primary role shifts from real-time tracking to acting as a motion planner that generates target joint states based on the current state and historical context.

  3. Improvement in Robustness to Latency/Dynamics:

  4. Mechanism: LFMC policies demonstrate significantly lower sensitivity to actuation latencies (up to 90 ms delay) and variations in system dynamics (e.g., without explicit dynamics randomization or complex actuator modeling). This is because the low control frequency allows the system time to settle into a stable tracking state before the next command is issued, effectively decoupling control decisions from fast, unmodeled disturbances.

  5. Capability: Robust and Dynamic Locomotion in Real-World Environments:

  6. Specific Outcome: The improved AI system (the robot controller) can robustly and repeatably achieve high heading velocities (e.g., 1.5 m/s) while traversing highly uneven, unstructured terrain without requiring expensive hardware for massive parallelization during training and deployment.

  7. Integration of State History for Enhanced Observability:

  8. Mechanism: Incorporate a history of joint states (e.g., the last 4 time steps) into the policy's state space, particularly for perceptive policies, or use it to augment blind policies.

  9. Improvement in Implicit Modeling: The inclusion of joint state history improves domain observability by implicitly encoding actuation dynamics and contact detection information, which is critical for high-frequency control but unnecessary for LFMC.

  10. Capability: Superior Handling of Contact and Stance Phases:

  11. Specific Outcome: The AI system can exhibit more accurate and stable transitions between stance (foot-in-contact) and swing (foot-not-in-contact) phases, leading to better recovery actions in unstable states compared to high-frequency controllers.

  12. Adaptive Policy Training Strategy for Frequency Tuning:

  13. Mechanism: Implement a comparative analysis framework that explicitly tests and compares control policies across a wide range of training frequencies (5 Hz to 200 Hz) and evaluate their resulting performance metrics (success rate, stability).

  14. Improvement in Control Design Philosophy: This allows engineers to move away from the intuitive notion that higher frequency equals better reactivity and instead adopt a design philosophy based on the system's inherent latencies (the motion planning hypothesis).

  15. Capability: Optimized Hardware-Agnostic Deployment: The AI system can be deployed successfully using minimal deployment code in C++ or Python, achieving performance comparable to (or exceeding) policies trained at higher frequencies, provided they are designed as planners rather than high-speed trackers.

  16. Enhanced Terrain Adaptation via Perceptive Policies:

  17. Mechanism: Utilize perceptive control policies that ingest rich, robo-centric terrain information (e.g., height maps) into their state space.

  18. Improvement in Environmental Awareness: The system gains explicit awareness of the environment's geometry and surface characteristics during planning, rather than relying solely on reactive feedback loops to correct errors caused by delayed sensory inputs.

  19. Capability: High-Performance Navigation over Complex Obstacles: The improved AI system can perform dynamic locomotion successfully over complex obstacles like stairs and bricks, achieving high success rates (e.g., 94% success rate on rough terrain) that are difficult to maintain with purely blind or high-frequency controllers when facing unexpected perturbations.

Abstract

Robotic locomotion is often approached with the goal of maximizing robustness and reactivity by increasing motion control frequency. We challenge this intuitive notion by demonstrating robust and dynamic locomotion with a learned motion controller executing at as low as 8 Hz on a real ANYmal C quadruped. The robot is able to robustly and repeatably achieve a high heading velocity of 1.5 m/s, traverse uneven terrain, and resist unexpected external perturbations. We further present a comparative analysis of deep reinforcement learning (RL) based motion control policies trained and executed at frequencies ranging from 5 Hz to 200 Hz. We show that low-frequency policies are less sensitive to actuation latencies and variations in system dynamics. This is to the extent that a successful sim-to-real transfer can be performed even without any dynamics randomization or actuation modeling. We support this claim through a set of rigorous empirical evaluations. Moreover, to assist reproducibility, we provide the training and deployment code along with an extended analysis at https://articulated.robots.ox.ac.uk/lfmc/.

Sources

Related papers