Learning Gait-Aware Quadruped Locomotion with Temporal Logic Specifications
cs.RO, cs.AI
Submitted: 2026-07-01
Updated: 2026-09-17
Project page: https://stl-locomotion.github.io
License: http://creativecommons.org/licenses/by/4.0/
The gist: Reinforcement learning (RL) for quadruped locomotion commonly depends on fixed, hand-crafted, and Markovian reward functions that may limit interpretability of learned policies and may lack explicit
Terminology
Abstract
Reinforcement learning (RL) for quadruped locomotion commonly depends on fixed, hand-crafted, and Markovian reward functions that may limit interpretability of learned policies and may lack explicit control over gait behaviors. We introduce a framework where distinct gaits are specified using parameterized constraints expressed in Signal Temporal Logic (STL). These include safety bounds, gait synchronization constraints, command tracking, and actuation bounds. From these specifications, we develop a reward shaping mechanism that provides learning agents a dense, continuous reward landscape that encodes desired behavior. We define parametric STL templates for three speed regimes (walking-trot, trot, bound), calibrate their parameters from reference rollouts, and compute rewards from using smooth approximations of STL robustness over the rollouts. The generated rewards can be used to provide shaped gradients compatible with Proximal Policy Optimization (PPO). We instantiate the approach on Google's Barkour quadruped robot in MuJoCo XLA (MJX). We use parallelization within the simulator to improve training speeds and use domain randomization to robustify learned policies. Compared with hand-crafted rewards, an expert-switching oracle, and Text2Reward, Human-STL maintains high command-tracking success across the evaluated speed range while exhibiting substantially higher consistency with the intended speed-dependent gait structures. Videos can be found on our project website: https://stl-locomotion.github.io/.
Sources
- Learning Agile Robotic Locomotion Skills by Imitating Animals
- Highly Dynamic Quadruped Locomotion via Whole-Body Impulse Control and Model Predictive Control
- Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning
- Barkour: Benchmarking Animal-level Agility with Quadruped Robots
- Sim-to-Real: Learning Agile Locomotion For Quadruped Robots
- Modular Deep Reinforcement Learning with Temporal Logic Specifications
- Dealing with Sparse Rewards in Reinforcement Learning
- Learning Robust Rewards with Adversarial Inverse Reinforcement Learning
- Model-based Reinforcement Learning from Signal Temporal Logic Specifications
- Learning multiple gaits of quadruped robot using hierarchical reinforcement learning
- Gaitor: Learning a Unified Representation Across Gaits for Real-World Quadruped Locomotion
- MuJoCo Playground
- Proximal Policy Optimization Algorithms
- Brax -- A Differentiable Physics Engine for Large Scale Rigid Body Simulation
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving