CRAX: Fast Safe Reinforcement Learning Benchmarking
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "CRAX: Fast Safe Reinforcement Learning Benchmarking".
Jane: Safety is a core concern for deploying reinforcement learning (RL) agents in real-world domains such as robotics and autonomous driving,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So, we're diving into the paper CRAX: Fast Safe Reinforcement Learning Benchmarking, and it’s a really important piece of work because it tackles that big gap between existing RL benchmarks and what we actually need for real-world robotics.
Jane: Exactly, Tom; this paper is focused on making sure that when we test safe reinforcement learning agents in complex three dee scenarios, we aren't stuck waiting around for slow computer simulations to run.
Lu: I think the title itself tells us a lot about what they're aiming for: speed and safety benchmarking. They are clearly trying to accelerate the process of finding safe AI policies by using powerful hardware acceleration.
Meng: From an engineering standpoint, that speedup is crucial because we can’t afford to spend months on simulations if we want to prototype new physical systems quickly.
Lalam: I see this as a way for the entire culture here at our startup to move faster, giving us concrete data from high-fidelity physics right away instead of relying on slower, less accurate models.
Tom: Right, so they propose a new way to benchmark safe RL that isn't bottlenecked by slow CPU simulations.
Jane: They focus on using the MuJoCo XLA engine with hardware acceleration to achieve speeds up to around one hundred times faster than traditional CPU-based safety benchmarks.
Lu: That speedup is significant because it allows for large-scale experimentation, which was a major bottleneck in previous research areas like this one.
Meng: It’s about scaling up the testing so that we can run many more scenarios to see how these algorithms behave under different safety pressures.
Lalam: And from my perspective as an AI model, having access to this kind of fast data means we can train safer agents more robustly because the feedback loop is much tighter and faster.
Tom: What's really interesting is that they combine this hardware acceleration with realistic three dee physics tasks, which gives them a unique setup compared to what we usually see in safety testing.
Jane: They have designed a benchmark featuring six different environment suites and three distinct agent types to really stress-test the methods.
Lu: It’s interesting because they are not just looking at one problem; they are covering a wide variety of physical challenges, which is essential for building versatile AI.
Meng: The paper mentions these environments span things like navigation and manipulation, so we're looking at a broad application space here.
Lalam: That breadth means the findings won't be tied to just one specific type of robot or task; it’ll give us a more holistic view of what makes an agent truly safe in diverse physical settings.
The paper's summary: Tom: Now, let’s talk about what CRAX actually does according to the paper. Essentially, they are proposing a system that uses the MuJoCo XLA engine to run high-fidelity three dee physics simulations much faster than conventional setups.
Jane: They are presenting this as a solution to the problem where existing safety benchmarks are too slow for large-scale research and prototyping of safe reinforcement learning agents.
Lu: The core idea is that by using vectorized operations and hardware acceleration, CRAX manages to provide a set of simulated tasks, robots, and algorithm baselines for evaluating Safe RL using parallel computing power.
Meng: It’s about moving the evaluation phase from being computationally prohibitive to being something that can be done efficiently.
Lalam: This means we can test more complex safety scenarios in the real world simulation environment much more frequently than before, which is a big step forward for developing reliable AI systems.
Tom: The paper details the benchmark design as well, noting that it involves six environment suites and three specific agent types across different physics tasks.
Jane: Each task within those environments has both a reward signal and a cost signal, which sets up the trade-off we need to manage in constrained Markov Decision Processes.
Lu: They formulate this objective as finding an optimal policy that maximizes expected cumulative reward while keeping the expected cumulative cost below a specific threshold for every state, which is formulated as "max pi in V pi(s) subject to C pi(s) d, s in S (one)".
Meng: That constraint formulation is what makes it truly safe RL; it’s not just about getting a good score, but about guaranteeing safety bounds.
Lalam: It’s fascinating how they make the safety requirements explicit constraints rather than just hoping reward shaping encourages safe behavior.
Tom: The paper also highlights the specific constraints they use across those suites, like hazard proximity costs and velocity thresholds to create different types of safety challenges.
Jane: They also cover various agent morphologies, including things like Point agents, Ant three dee quadruped robots, Humanoids and Spiders, HalfCheetah runners, Walker2d walkers, and Reacher arms.
Lu: The progression of difficulty levels is handled by designing tasks so that lower levels are achievable for existing methods while higher levels introduce more complex hazards or stricter kinematic requirements for those agents.
Meng: This structured difficulty scaling is key because it lets researchers start simple and gradually push the complexity of the safety challenge, which is much better for teaching an algorithm how to handle nuance.
Lalam: Having these distinct agent types means we can test safety across different physical forms, which is vital when we think about deploying AI in such diverse real-world robotic systems.
The paper's improvements: Tom: Moving on to the actual findings and what they actually managed to prove with CRAX, the paper shows some really interesting results regarding how different methods perform.
Jane: They evaluate six popular safe RL methods, including PPO, PPOCost, PPOLag, PPOPID, PPOSaute, P3O, and FOCOPS.
Lu: The key takeaway from their empirical evaluation is that no single approach dominates across all the different tasks they tested; there are clear trade-offs between performance and safety.
Meng: That's interesting because it means we can’t just pick one algorithm and assume it will work everywhere; we have to understand the specific safety requirements for each application.
Lalam: I think this suggests that a generalized, universally safe algorithm might not exist yet; instead, the research needs to be tailored to the specific physics and constraints of the task at hand.
Tom: They also highlight some very important learning strategies they found that can actually improve performance when you train in harder settings.
Jane: Specifically, they found that curriculum learning across difficulty levels and safety transfer can boost performance when training in more challenging environments compared to just training directly in those hard settings.
Lu: That finding is quite insightful because it suggests that a structured learning path is a valid way to improve performance, and it’s not just about brute-force training.
Meng: So, if we want our AI to perform well in high-stakes situations, maybe we should look into structured curricula rather than just throwing more data at the problem.
Lalam: It’s encouraging because it gives us a concrete strategy for improving performance when the constraints get tighter, which is something everyone is looking for in practical deployment.
Tom: Furthermore, they pointed out specific methods that stood out in certain areas of this evaluation.
Jane: They found that P3O and FOCOPS were the strongest baselines overall when looking at their performance on CRAX.
Lu: And they also noted something about PPOLag, which achieved the highest safety percentage, being the only baseline to satisfy all cost bounds on Levels two and three of those environments.
Meng: That’s a strong result for PPOLag; achieving full constraint satisfaction across all levels is a significant milestone in ensuring guaranteed safety.
Lalam: Having a method that reliably meets all constraints in complex scenarios, even the harder ones, gives us confidence when we think about deploying an agent where failure has serious consequences.
Conclusion: Tom: So, to wrap up this discussion on the CRAX paper, the main message is that they've successfully created a fast way to benchmark safe RL using high-fidelity simulation powered by hardware acceleration.
Jane: They demonstrated that evaluating algorithms across six different environments and various agent types reveals clear performance-safety trade-offs rather than finding one single best solution.
Lu: The paper’s implication is that we need to be careful about how we approach safety, realizing that balancing reward maximization and constraint satisfaction requires tailoring the strategy to the specific problem's physical demands.
Meng: From an engineering viewpoint, this means our next steps involve integrating these hardware-accelerated tools into our development pipelines for faster, more rigorous testing of physical AI agents.
Lalam: This work is a vital tool for everyone involved in safe RL research because it provides a concrete platform to develop and analyze future methods, and it sets the stage for much more rigorous testing down the road.
Tom: We’ve discussed how CRAX helps us see where performance bumps up against safety constraints across different tasks and agent types.
Jane: The authors have shown that curriculum learning can be used to improve performance by training in progressively harder settings, which is a useful strategy for getting better results when safety bounds are tight.
Lu: It really shows the interplay between learning structure and the environment design in finding effective solutions for constrained problems.
Meng: So, we’re looking at how this applies to our practical development pipeline moving forward—focusing on making those simulations as efficient and rigorous as possible.
Lalam: I think CRAX gives us a solid foundation for building smarter, safer AI systems because it provides the necessary tools to systematically explore the solution space in a controlled and accelerated manner.
Tristan Tomilin, Mourad Boustani, Mickey Beurskens, Thiago D. Simão
Eindhoven University of Technology
cs.LG, cs.AI
Submitted: 2026-06-18
Updated: 2026-09-27
Code: https://github.com/obertTLange/gymnax
Importance score: 83/100
The gist: Safety is a core concern for deploying reinforcement learning (RL) agents in real-world domains such as robotics and autonomous driving, and this work addresses the gap between existing RL benchmarks
Key concepts
- CRAX Architecture
- CRAX is built on the MuJoCo XLA (MJX) physics engine, which uses hardware acceleration for vectorized operations. This design enables significant speedups compared to traditional CPU-based setups, making it suitable for large-scale parallel experimentation in Safe RL.
- Constrained Markov Decision Process (CMDP)
- CMDP is the mathematical framework used to define the benchmark. The goal is to maximize expected reward while ensuring that a specified cost remains below a fixed threshold across all states. This forces agents to learn policies that balance high performance with strict safety constraints.
- Agent Morphologies
- The benchmark includes diverse robot types, such as Point, Ant (quadruped), Humanoid, and Spider. These agents are tested across different difficulty levels where higher levels introduce more complex hazards or stricter physical requirements to challenge the agent's safety capabilities.
Terminology
Summary
Safety is a core concern for deploying reinforcement learning (RL) agents in real-world domains such as robotics and autonomous driving, and this work addresses the gap between existing RL benchmarks and high-fidelity simulation by proposing CRAX, a hardware-accelerated benchmark that enables large-scale experimentation. The gist is that CRAX provides a hardware-accelerated SafeRL benchmark leveraging MuJoCo to yield up to ∼100x speedups over comparable CPU-based safety benchmarks, revealing trade-offs between performance and safety across six environment suites and three agent types.
CRAX Architecture and Performance
CRAX is built on top of the MuJoCo XLA (MJX) physics engine, which utilizes vectorized operations and hardware acceleration to achieve significant speedups over traditional CPU-based setups. This architectural choice allows CRAX to provide a set of simulated tasks, robots, and algorithm baselines for evaluating SafeRL leveraging parallel computing. The design principles are inspired by BRAX [12] and Safety Gymnasium [16], aiming to facilitate rigorous testing and faster algorithm development for the SafeRL community.
Benchmark Design and Constraints
The benchmark features six environment suites spanning various tasks, such as Safe Navigation,
Safe Velocity,
Safe Pathway,
and Safe Spider.
Each task defines both a reward signal and a cost signal, inducing a trade-off between performance and safety. The objective in the Constrained Markov Decision Process (CMDP) framework is to find an optimal policy that maximizes expected cumulative reward while ensuring the expected cumulative cost remains below a specified threshold, formulated as: max π∈Π Vπ(s) subject to Cπ(s) ≤ d, ∀s ∈ S (1)
. The benchmark employs five key constraint formulations across the suites, including Hazard proximity/contact costs,
Velocity threshold constraints,
and Height constraints.
Agent Morphologies and Difficulty Progression
CRAX includes agents with diverse morphologies, such as Point, Ant (3D quadruped), Humanoid (3D bipedal), Spider (3D hexapod), HalfCheetah (2D planar runner), Walker2d (2D bipedal walker), and Reacher (2-link arm). Each environment suite is designed with three difficulty levels. The lowest level tasks are designed such that most existing methods are capable of learning a reasonable policy and achieving meaningful performance,
while higher levels introduce greater challenges, such as a greater number and variety of hazards
in Safe Goal or requiring the Spider agent to keep one additional leg off the ground
in Safe Spider.
Empirical Evaluation and Findings
The work evaluates six popular SafeRL methods, including PPO [34], PPOCost [41], PPOLag [31], PPOPID [36], PPOSaute [35], P3O [46], and FOCOPS [47]. The evaluation assesses performance-safety trade-offs by varying cost thresholds, curriculum learning, and safety transfer. Key findings include:
Curriculum learning across difficulty levels and safety transfer can improve performance over direct training in harder settings.
The results show that P3O and FOCOPS are the strongest baselines on CRAX,
while PPOLag achieves the highest safety percentage, being the only baseline to satisfy all cost bounds on Levels 2 and 3.
Furthermore, for specific tasks like Safe Reacher, curriculum learning boosts performance for all methods except PPOPID.
Computational Efficiency and Scalability
CRAX demonstrates superior computational efficiency compared to CPU-based benchmarks. A case study comparing CRAX to Safety-Gymnasium shows that CRAX scales far beyond the limitations of CPUbound physics simulation, reaching ∼ 300K steps per second (SPS) at around 8192 parallel environments,
whereas Safety-Gymnasium plateaus early due to CPU and memory bottlenecks.
This throughput advantage is demonstrated by the fact that a full evaluation suite on CRAX completes in two weeks, compared to nearly a year for the equivalent run on Safety-Gymnasium. The benchmark also exposes how varying safety bounds induces safety–performance trade-offs,
with PPOLag reliably adapting to varied bounds across all environments.
Limitations and Future Directions
The current work has limitations, including the fact that MJX does not yet support all features available in the CPU-based MuJoCo, such as certain rigid-body collision types.
Furthermore, the evaluation focuses exclusively on on-policy methods and safety is evaluated only through expected cumulative cost constraints. Future work could investigate off-policy and model-based safe RL approaches,
explore other safety formulations, and extend CRAX to multiagent safety scenarios
using pixel observations. The paper concludes that CRAX serves as a vital tool for developing, analyzing, and benchmarking future safe RL methods.
References
[1] Eitan Altman. Constrained Markov Decision Processes.
Improvements for AI systems
As a fastidious and diligent researcher, I have analyzed the CRAX (Constrained RL Accelerated with JAX) paper. The core innovation lies in providing a high-fidelity, hardware-accelerated benchmark for Safe Reinforcement Learning (SafeRL) using 3D MuJoCo physics with explicit cost constraints.
Here are the specific improvements and capabilities this framework enables for AI systems:
-
The ability to train and evaluate RL agents in complex, physically realistic 3D environments at a speed up to 100x compared to CPU-based benchmarks, thanks to leveraging JAX/MJX hardware acceleration.
-
The capability to rigorously test and compare diverse SafeRL algorithms (like PPO-Lagrange, PPOPID, FOCOPS) across six distinct environment suites (Goal, Button, Circle, Push, Velocity, Height) and various agent morphologies (Point Ant Humanoid Spider).
-
The capacity to explicitly model safety requirements not just as soft rewards but as hard constraints in the form of expected cumulative discounted cost bounds.
Specific improvements enabled by CRAX:
-
A SafeRL agent can be developed that navigates a complex 3D arena (e.g., the Goal suite) while simultaneously being constrained to avoid specific hazard geometries (e.g., cylinders or cubes) with quantifiable proximity costs, ensuring it never enters a
danger zone
defined by a cost threshold. -
A robotic manipulator (Reacher agent) can be trained to reach a target object while adhering to strict collision constraints on its links, minimizing the probability of intersecting predefined hazardous obstacles.
-
Locomotion systems (Ant, HalfCheetah, Hopper) can be optimized for high-speed movement while strictly maintaining velocity limits that prevent mechanical failure or unsafe maneuvers in real-world applications.
-
Humanoid agents can learn to maintain a specific posture (Height constraint) while executing locomotion tasks, ensuring they remain within safe physical operating parameters during dynamic movement.
-
Multi-agent systems (Spider agent) can be trained to execute complex gaits that strictly adhere to kinematic constraints, such as keeping specific legs airborne or restricting foot placement patterns, which is critical for legged robot stability and safety.
-
The framework allows for the systematic study of learning strategies:
narrowing the gap between performance and safety by employing curriculum learning (training on progressively harder tasks) or knowledge transfer (initializing agents with policies trained in related environments), leading to better generalization in high-stakes scenarios.
Sources
- OpenAI Gym
- AI Safety Gridworlds
- Isaac Lab: A GPU-Accelerated Simulation Framework for Multi-Modal Robot Learning
- SafeOR-Gym: A Benchmark Suite for Safe Reinforcement Learning Algorithms on Practical Operations Research Problems
- Benchmarking Batch Deep Reinforcement Learning Algorithms
- Proximal Policy Optimization Algorithms
- DeepMind Control Suite
- MuJoCo Playground
- Penalized Proximal Policy Optimization for Safe Reinforcement Learning
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks