FAITH: Feasibility-Aware Safety-Filtered RL for High-Dimensional Systems

arXiv:2610.12432 · cs.RO, cs.LG, cs.SY, eess.SY · Submitted 2026-10-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.

Dev: Today's paper: "FAITH: Feasibility-Aware Safety-Filtered RL for High-Dimensional Systems".

Rosa: The gist The FAITH framework introduces a feasibility-aware,

Dev: First, who's behind it and why it matters.

Title and authors: Rosa: So we’re looking at this paper today, it’s called FAITH: Feasibility-Aware Safety-Filtered RL for High-Dimensional Systems. It sounds pretty technical, but basically, they are tackling a problem where you want a robot to do a task while also making sure it doesn't crash or hurt someone.

Dev: Right. The title says feasibility aware safety filtered reinforcement learning for high dimensional systems. That points straight at the issue of scaling up to complex robots where safety becomes really hard to track.

Taro: I think what’s interesting is how they handle the inherent conflict between making progress on the task and strictly adhering to a safety rule when both are in one objective function. It suggests separating those two things completely.

Rosa: Exactly, because usually, if you put safety right into the main goal, it starts competing with the actual performance of finishing your job. This paper proposes a way around that competition by splitting them up during execution.

Dev: So they’re not trying to find one policy that does both perfectly at once in the same objective function; they are using a filter mechanism to manage how actions are actually performed.

The paper's summary: Rosa: To break it down simply, FAITH introduces an approach where the task policy is trained only through the task reward, and safety is handled by this learned filter that modifies what actions actually happen. It’s about letting safety change the dynamics the policy experiences without trying to fight against it in its primary goal.

Dev: That means when you’re training, you aren't fighting a penalty term that tries to stop you from moving; instead, you are optimizing for the task reward under these new, safer dynamics defined by the filter.

Taro: I see how that helps with long-term returns because the policy can adapt to safety interventions over time since it’s not fighting a constant safety brake on every single step.

Rosa: Right. The paper suggests they approximate the optimal state-action safety value using a critic and an actor, which then feeds into this filter to figure out the minimal correction needed for an action. This is how they get that feasibility awareness without needing a complex physical model of everything upfront.

Dev: They’re using these learned components to guide the task policy training through filtered rollouts, which means they’re learning how to perform the task based on what the filter allows them to do safely.

The paper's improvements: Rosa: Now, what they propose as improvements is really addressing three specific problems that other safety filtering methods have. They point out that just minimizing the immediate action deviation isn't good enough for long-term task return.

Dev: And they also flag a real issue with this approach: if there’s absolutely no action available that meets the learned safety condition, their original method doesn't tell you what the least harmful thing to do is.

Taro: That sounds like it could leave a robot stuck or in a bad spot if it can't find any safe move at all. So, they suggest making sure that even when constraints are impossible to meet, the system defaults to prioritizing lower predicted harm instead of just stopping or failing.

Rosa: They propose using the learned filter not just for feasibility when actions exist, but also as a way to prioritize the action with minimum predicted peak harm if no safe action is found. It’s about making sure there’s always *some* sensible response even when things look impossible.

Dev: And they address another point by making sure the task policy learns from these filtered rollouts using only the task reward, which keeps safety completely separate from the return optimization objective during policy updates.

Conclusion: Rosa: So to wrap up, FAITH proposes separating safety enforcement from task optimization by training the policy through a learned filter that uses only the task reward, which means safety and return aren't competing terms in the main objective. This allows them to optimize both things in different parts of the system.

Dev: The key idea is approximating that optimal safety value using a critic and actor, then using that information to guide a feedforward filter that handles the action correction dynamically during execution.

Taro: For me, what this means for autonomy is that when the robot encounters a situation where its learned safe set is empty, it doesn't just fail; it uses the filter to pick the action with minimum predicted peak harm instead of needing a whole separate recovery controller.

Rosa: That’s a big deal because it simplifies the hardware side and gives us continuous operation capability even when the constraints are truly infeasible. We’re talking about high-dimensional systems like that twenty-nine-DoF humanoid they tested, showing it works in challenging settings under conditions where feasibility is tough to maintain <ref:2610.12432#pg1>.

Dev: And from an engineering standpoint, the paper shows they can achieve a high safety rate while keeping a lot of the task return—they matched ninety-seven percent of unfiltered return in some tests—which is pretty solid when you consider the complexity involved.

Taro: It proves that we don't need to switch between different modes or have separate reward functions for safety and performance; one unified architecture can handle both through this filtering mechanism, which really scales up the idea of safe interaction in complex physical spaces.

Rosa: So FAITH is a framework that separates these concerns by using learned value functions and feedforward filtering to recover the constrained problem without needing a competing safety term in the task policy objective. That’s what we have today on FAITH.

Songyuan Zhang, Baljeet Singh, Sarthak Ranjeet Kaingade, Chuchu Fan, Bryan Trinh

Massachusetts Institute of Technology · Amazon

cs.RO, cs.LG, cs.SY, eess.SY

Submitted: 2026-10-08

Updated: 2026-10-08

The gist: The gist The FAITH framework introduces a feasibility-aware, model-free reinforcement learning approach that separates safety enforcement from task policy optimization by approximating an optimal

Key concepts

Safety Q-Function
This component uses a neural network (Qh) to estimate the optimal safety value for a given state and action. An action is deemed safe if this estimated value is less than or equal to zero, providing a learned measure of potential harm.
Backup Actor
Trained using the safety critic's Bellman update, this actor learns actions that minimize predicted harm. It helps guide the system toward safer trajectories by minimizing the expected negative safety outcome based on past experiences.
Feedforward Safety Filter
This network approximates a minimal action correction needed to stay within safe action boundaries. It is trained to penalize deviations from the learned set of admissible actions, effectively filtering out unsafe proposals from the task policy.
Separation of Objectives
FAITH trains the task policy using only task reward, while safety is handled by a separate filter. This means safety influences the dynamics experienced by the policy but does not appear as a competing term in its main objective, allowing both goals to be optimized independently.

Terminology

Summary

The gist The FAITH framework introduces a feasibility-aware, model-free reinforcement learning approach that separates safety enforcement from task policy optimization by approximating an optimal state-action safety value and training the task policy through a learned feedforward filter.

How it works

FAITH is designed to address the competing updates between safety and task performance that arise when they are both incorporated into a single policy objective. It achieves this by approximating the optimal state-action safety value using a safety critic and a backup actor, which guides a feedforward filter to approximate the minimal action correction. The task policy is then trained through this filter using only the unmodified task reward, ensuring that Safety therefore changes the dynamics experienced by the policy but does not appear as a competing term in its objective.

Key Components of FAITH

The framework consists of several interconnected modules that operate sequentially during execution and training.

  1. Safety Q-Function: This component approximates the optimal safety critic using a neural network Qh,ϕ where a persistent action admits a safe trajectory if Qh,ϕ(s, a) ≤ 0.

  2. Backup Actor: This actor is trained to approximate the minimum over actions in the Bellman target for the safety critic where samples are drawn from the replay buffer D.

  3. Feedforward Safety Filter: This filter approximates the minimal-intervention correction by using a feedforward residual network to approximate the solution to the projection problem where it is trained with a loss that penalizes deviations from the learned admissible-action set

Training and Optimization Loop

The training process involves three distinct phases that separate safety learning from task-return optimization.

(a) Filtered rollout and deployment:

The task policy proposes a nominal action anom, and the safety filter produces the executed action aexec, which defines the filtered dynamics seen by the task policy.

(b) Off-policy safety learning:

The replay buffer D stores executed transitions, and a backup actor supplies actions to minimize predicted harm using the safety critic's Bellman update.

(c) On-policy task learning:

The task policy πθ is updated using PPO based on filtered-rollout returns, maximizing the return under the dynamics defined by the filter.

Performance and Results

Experiments on a double integrator and a 29-DoF humanoid demonstrate high performance in feasible settings and harm reduction in infeasible settings

(Q1) Feasibility:

Starting from feasible states, FAITH maintains a 100% safety rate while achieving the highest return among all safe methods in the double integrator environment.

(Q2) Harm Minimization:

When harm is unavoidable, FAITH attains the lowest peak harm compared to other methods in the common failure Push-Avoid episodes.

Conclusion

FAITH successfully separates safety enforcement from task policy optimization by training the task policy through a learned filter using only task reward, thereby allowing safety and return to be optimized in different modules. The framework demonstrates that under stated assumptions, filtering preserves the constrained optimum while ensuring that when the learned feasible set is empty, the filter prioritizes lower predicted peak harm. The results are further validated by qualitative demonstrations on a real-world Unitree G1 humanoid. The limitations acknowledged include discounting and finite lambda leaving the learned controller without formal safety or optimality guarantees.

The paper is relevant to high-dimensional systems because it provides a model-free framework that combines learned safety value functions and feedforward action filtering to recover the feasible constrained problem without a competing safety term in the task-policy objective. This matters because it allows for agile behaviors in high-dimensional legged and humanoid robots while ensuring they operate safely around people or equipment. It is significant because it addresses the difficulty of safety depending not only on current geometry but also on whether future actions can prevent a violation eventually in contact-rich, underactuated systems. The paper's findings are important because it shows that safety and task return can be optimized in different modules without competing updates in the policy objective. The paper's contribution is that when no action satisfies the learned safety condition, the same controller approaches the action with minimum predicted peak harm without needing a mode switch or a separate recovery controller. The paper's experiments show that FAITH matches the highest safety rate in Walking-Avoid while retaining 97% of unfiltered return and achieves the lowest harm in Push-Avoid episodes. The paper's ability to scale to a 29-DoF humanoid demonstrates hardware execution of this model<ref:2610.124325VI,We consider a 29-DoF Unitree G1 humanoid in Isaac Lab [41], actuated by joint-position targets>. The paper's overall finding is that FAITH provides a unified architecture allowing protected-region clearance to take precedence over balance without needing a fall-away specific reward or mode switch<ref:2610.

Improvements for AI systems

  1. Bold header: Safety-Filtered Policy Training

The task policy is trained through the filter using only reward, which means safety does not appear as a competing term in its objective. This allows for maximizing return under constraints without needing a separate safety penalty term that might oppose performance.

  1. Bold header: Learned State-Action Safety Value Approximation

FAITH approximates the optimal safety critic, denoted as Q⋆h(s, a), using a neural network approximation Qh,ϕ and samples from a replay buffer D to train the backup actor µψ. This allows the system to estimate the optimal state-action safety value even without an explicit dynamics model.

  1. Bold header: Amortized Minimal-Intervention Filtering

Instead of solving an optimization problem at every step, FAITH uses a feedforward network to approximate the filter, which approximates the minimal action correction. This is achieved by training the filter with a loss function like LF(ω)=Eν∥aexec−anom∥2 + λ[Qh,ϕ(s, aexec)+δ], allowing for fast execution in deployment.

  1. Bold header: Adaptive Harm Minimization on Infeasible Starts

When no action satisfies the learned safety condition, the filter prioritizes the action with minimum predicted peak harm, rather than requiring a mode switch or fallback policy. This ensures continuous operation even when constraints are impossible to meet.

Sources

Related papers