Bridging Reinforcement Learning and Optimal Control via Feasible Action Mapping
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "Bridging Reinforcement Learning and Optimal Control via Feasible Action Mapping".
Dev: Operating constrained dynamical systems requires controllers to efficiently solve complex tasks while enforcing recursive feasibility and safety constraints,
Rosa: First, who's behind it and why it matters.
Paper summary: Dev: So, to wrap up our discussion on "Bridging Reinforcement Learning and Optimal Control via Feasible Action Mapping," the paper presents a framework, FAOC, that uses a specific optimization-based mapping algorithm. The core idea is transforming an RL agent’s action into a state-dependent parameter set for the Optimal Control Problem, which guarantees strict satisfaction of dynamical system constraints.
Rosa: Precisely; we saw how this framework moves beyond just using RL or just using OC by creating this explicit link between them, and the authors demonstrated its topological characterization of feasibility and its invertible mapping algorithm that ensures all feasible actions are included.
Taro: The implication is that for complex, constrained systems, we can leverage the predictive power of Reinforcement Learning while simultaneously enforcing the rigorous safety guarantees derived from Optimal Control theory. This opens up avenues for deploying autonomous agents in environments where strict constraint satisfaction is non-negotiable.
Dev: I think what this means practically is that we can trust the control signals generated by these coupled systems more deeply because they are mathematically guaranteed to respect those boundaries, provided the underlying system dynamics meet certain geometric assumptions.
Rosa: And that’s a big deal because it moves us closer to having truly reliable, high-performance robotic systems operating in real-world scenarios that have inherent physical limits.
Taro: We also see this as a way to push the boundary of autonomy; instead of relying on brittle trial and error, we can design controllers with built-in mechanisms for guaranteed feasibility under dynamic conditions.
Dev: It’s a sophisticated way to handle the complexity, though I do have my engineering reservations about how reliably the mapping performs when those geometric assumptions start to break down in unforeseen ways.
Conclusion: Rosa: So, we're wrapping up our look at "Bridging Reinforcement Learning and Optimal Control via Feasible Action Mapping." This paper essentially shows how to take those two powerful methods, RL and optimal control, and make them talk to each other in a way that guarantees safety constraints are met.
Dev: Right; the core idea is that they've built this mapping algorithm so you don't just get random actions from the RL agent. Instead, it translates those abstract choices into a set of parameters that the optimal control problem actually knows how to solve safely.
Taro: That translation step is what really intrigues me because it means we can use the learning capabilities of AI without worrying that it will immediately fly off into an infeasible state according to physics or system limits.
Rosa: Exactly; I'm wondering if this kind of guaranteed feasibility translates well outside the controlled lab environment, like when a robot has to navigate a messy, unpredictable real-world setting for hours.
Dev: That’s my main concern; the paper focuses on mathematical guarantees based on specific geometric shapes for those parameter sets, but how robust is that mapping if the physical system starts behaving in a way that violates those initial assumptions?
Taro: I think the authors address that by developing methods to handle distortion and even derive target shapes directly from OCP constraints, which suggests a level of adaptability when things go sideways.
Rosa: So we're looking at a framework where an RL agent proposes something, but the control system immediately filters it through this mathematical lens to ensure it stays within the bounds of what the physics allows?
Dev: Precisely; and I’m still focused on how fast that whole mapping process runs. If we need a high loop rate for fast dynamics, can this translation from RL action to feasible parameter happen in time for real-time control?
Taro: That efficiency is crucial because if the computational overhead makes it too slow, the guarantee of feasibility becomes useless in a dynamic situation where you need immediate reaction to unexpected changes.
Rosa: It feels like we’re moving toward systems that are not just smart but also inherently safe and predictable, which could really open up doors for complex autonomous operations.
Dev: I'm ready to hear more about the specific validation results they showed in those robot table tennis experiments before we move on to the details of their mapping algorithm.
Richter Optimization GmbH · Sony AI
eess.SY, cs.RO, cs.SY
Submitted: 2026-07-27
Updated: 2026-10-07
Comments: 26 pages, 6 figures
Code: https://github.com/SonyResearch/feasible_action_for_optimal_control
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 80/100
The gist: Operating constrained dynamical systems requires controllers to efficiently solve complex tasks while enforcing recursive feasibility and safety constraints, which this work addresses by presenting
Key concepts
- Topological Characterization of OCP Feasibility
- This concept proves that for many optimal control problems, the set of parameters that make them solvable (P(x)) has a nice geometric shape: it is compact, solid, and convex. This mathematical guarantee ensures that the RL agent's chosen parameters will always lead to a problem with a solution.
- Invertible Feasible Action Mapping
- This is an efficient algorithm that connects the abstract actions chosen by the RL agent (A¯) to the feasible parameter set P(x). It acts like a bridge, ensuring that every possible feasible action from the RL agent corresponds uniquely to a valid parameter for the optimal control problem, preventing infeasible choices.
- Areamatching Directional Transformation
- This technique is used to fix geometric distortions that occur when mapping sets with different shapes. By using this transformation, the authors prevent 'point accumulation,' which is a problem where many different input actions map to the same output parameter, ensuring better performance and sample efficiency.
- FAOC Framework Operation
- The FAOC process takes an RL agent's output (mean and covariance) and transforms it through three steps: first, mapping the abstract action to a state-dependent feasible parameter set; second, using that parameter to define the OCP; and finally, solving the resulting optimal control problem. This closes the loop between learning and control.
Terminology
Summary
Operating constrained dynamical systems requires controllers to efficiently solve complex tasks while enforcing recursive feasibility and safety constraints, which this work addresses by presenting Feasible Action for Optimal Control (FAOC), a novel control framework integrating Reinforcement Learning (RL) and Optimal Control (OC). The key contribution is a computationally efficient, optimization-based mapping algorithm that transforms the RL agent’s action from a static abstract set into a state-dependent feasible parameter set of the Optimal Control Problem (OCP), guaranteeing strict satisfaction of the dynamical system’s constraints.
Key Contributions
The paper introduces several main contributions to resolve persistent feasibility and exploration challenges in RL-OC coupling:
(i) Topological Characterization of OCP Feasibility: The authors identify mild geometric conditions ensuring the state-dependent parameter set P(x) of a generic OCP is compact, solid, and convex (Lemma 1).
This property is explicitly demonstrated for linear MPC (Corollary 1).
(ii) Invertible Feasible Action Mapping: They propose a computationally efficient radial algorithm (Algorithm 1) bijectively transforming any abstract RL action a¯ ∈ A¯ into a guaranteed-feasible parameter p ∈ P(x) (Proposition 1).
This ensures that all feasible actions are included and infeasible ones are strictly excluded.
(iii) Mitigation of Geometric Distortion: To prevent point accumulation,
they developed an areamatching directional transformation
(Propositions 2 and 3) and a scalable linear transformation utilizing affinely related ellipsoidal surrogates for arbitrary dimensions (Proposition 4).
(iv) Tractability for Implicitly Defined Sets: For real-time scalability without explicit geometric representations of P(x), they derived robust formulations for computing interior points (Propositions 6 and 7) and target shape matrices (Proposition 8) directly from OCP constraints.
(v) Validation in Robot Table Tennis: They apply FAOC to motion planning for an 8-DoF robot arm, showing that it consistently surpasses state-of-the-art RL-MPC controllers in sample efficiency, final performance, and system controllability.
How the Framework Operates
The FAOC framework operates through a four-step process illustrated in Figure 1:
-
An RL agent outputs the parameters (mean µ and covariance Σ) of a multivariate Gaussian distribution, given the system state x and potentially other observations senv. A raw action is sampled from this distribution during training, or taken as the mean during deployment. A nonlinear function b transforms this unbounded raw action into an abstract action a¯ confined within a static, compact abstract set A¯.
-
The abstract action a¯ is mapped via M into a parameter set P(x) that depends on the current state x of the dynamical system. This set encompasses all feasible parameters for the underlying OCP (e.g., terminal state targets).
-
The mapped parameter p fixes the action-encoding vector in the OCP, which is subsequently solved to yield an optimal control input trajectory.
-
This trajectory is applied to the system in open loop, and the resulting subsequent state x+ serves as the initial condition for the next control cycle.
Topological Guarantees for Feasibility
The paper establishes that for a generic OCP (2), under specific assumptions regarding its components—namely that Z is compact, solid, and convex; f is continuous; g is jointly convex and continuous; S(x) has no directions of recession in p; and there exist feasible points z˜ ∈ relint Z and p˜ ∈ Rn—the state-dependent parameter set P(x) is guaranteed to be a compact, solid, and convex subset of Rn
(Lemma 1). This mathematical guarantee ensures that the RL agent only selects parameters p that guarantee problem feasibility.
Furthermore, for linear MPC with a terminal constraint (Corollary 1), feasibility is guaranteed if the terminal constraint matrix C has full row rank.
Invertible Mapping Algorithm
The core mechanism for coupling the layers is Algorithm 1, which performs an invertible mapping between convex sets
X (the abstract action set A¯) and Y (the state-dependent parameter set P(x)). This radial scaling methodology involves:
-
Determining a
point of the base set, applying a bijective directional transformation ϕ(·), and proportionally scaling the transformed ray relative to an interior point of the target set.
-
The proof establishes that this mapping M is bijective, ensuring that any feasible parameter chosen by the optimal controller can be
uniquely projected back into the abstract action space of the RL agent, avoiding action aliasing.
Mitigating Geometric Distortion
To ensure high performance and sample efficiency, the paper addresses point accumulation
which occurs when mapping sets with different aspect ratios. This is mitigated using:
Improvements for AI systems
As a fastidious researcher, I have analyzed this paper, Bridging Reinforcement Learning and Optimal Control via Feasible Action Mapping (FAOC),
and identified several specific, high-impact areas for improvement in current AI systems.
The core innovation is the creation of the FAOC framework: a computationally efficient mapping algorithm that transforms an RL agent's abstract action into a guaranteed feasible parameter set for an Optimal Control Problem (OCP), thereby merging the flexibility of RL with the safety guarantees of OC.
Here are the specific improvements and what they enable:
)1. Enhanced Safety and Constraint Satisfaction in Autonomous Systems
The most critical improvement is moving from RL that might fail
to RL that is guaranteed to be safe.
-
An AI system can now operate in real-world, high-stakes environments (like robotics, autonomous vehicles, or industrial control) where constraint violation is catastrophic. FAOC ensures that the actions suggested by a Reinforcement Learning policy—which are inherently flexible and potentially infeasible—are mathematically mapped to a set of parameters that strictly satisfy physical limits (kinematics, torque bounds, safety envelopes) defined by an underlying Optimal Control Problem (OCP).
-
This enables the deployment of RL agents in safety-critical domains without requiring manual, expert design for every constraint. The RL agent focuses solely on the overarching task reward maximization, while FAOC handles all recursive feasibility and safety constraints.
)2. Superior Sample Efficiency in Complex Manipulation Tasks
The paper demonstrates that FAOC outperforms state-of-the-art RL-MPC baselines in both sample efficiency and closed-loop performance (as shown in Section 5).
-
An AI system can learn complex, high-dimensional motor skills (e.g., delicate manipulation, intricate assembly) significantly faster than traditional methods. Because the mapping algorithm is computationally efficient (using radial scaling or linear transformations), it allows the RL agent to explore its action space more effectively without being immediately penalized by infeasibility.
-
This drastically reduces the number of expensive physical interactions required for training, making complex skills achievable in less time and with less real-world hardware wear.
)3. Robustness Against Geometric Distortion in High-Dimensional Action Spaces
The paper addresses action aliasing
and point accumulation
—problems where different abstract RL actions map to the same or clustered feasible OCP parameters, degrading learning efficiency.
-
The improved system will exhibit higher stability and better generalization across the entire continuous action space. By utilizing advanced mapping techniques (like the area-matching transformation or ellipsoidal surrogates), the system ensures that every meaningful exploration step in the RL policy leads to a distinct, relevant parameter configuration in the OCP, preventing
blind
exploration where many actions yield redundant results. -
This leads to policies that are less prone to getting stuck in local optima caused by geometric mapping artifacts.
)4. Real-Time Scalability for Complex Control Horizons
The framework is designed to handle long-horizon problems and complex, state-dependent constraints without sacrificing real-time performance.
-
An AI system can manage tasks requiring extensive planning (e.g., long-horizon motion planning or multi-stage trajectory generation) that would otherwise cause standard OCP solvers to become computationally intractable. The paper's method for computing interior points and target shape matrices directly from the OCP constraints (Section 4.4) bypasses the need for explicit, high-dimensional geometric representations of the target set during runtime.
-
This allows AI systems to generate complex, multi-step control sequences in real-time, bridging the gap between slow deliberation and fast execution.
)5. Generalizability Across Different System Topologies
The framework is designed to be largely topology-agnostic, relying on the topological properties of the sets (compact, solid, convex).
- This allows AI engineers to apply this solution across a wider variety of physical systems—from simple linear MPC formulations to complex nonlinear dynamics—without needing bespoke mapping algorithms for each new system. The core logic remains invariant; only the specific geometric definitions of the base set and target set need to be supplied.
In summary, the improved AI system is an RL agent that possesses both the intuition of a deep learning model and the rigor of a formal mathematical controller. It can perform complex, real-time physical tasks with guaranteed safety and superior learning speed compared to existing RL-based or OC-based controllers.
Abstract
Operating constrained dynamical systems requires controllers to efficiently solve complex tasks while enforcing recursive feasibility and physical constraints. To address these competing requirements, we present Feasible Action for Optimal Control (FAOC), a novel control framework integrating Reinforcement Learning (RL) and Optimal Control (OC). The core contribution is a computationally efficient, optimization-based mapping algorithm that transforms the RL agent's action from a static abstract set into a state-dependent feasible parameter set of the Optimal Control Problem (OCP), guaranteeing instantaneous parameter feasibility. When paired with invariant terminal sets, FAOC guarantees strict recursive feasibility and safe operation, effectively combining the predictable safety of OC with the behavioral flexibility of RL. Unlike prior work, the abstract action space does not require expert tuning, nor is the OCP formulation compromised by the inability of RL to guarantee feasibility. We evaluate FAOC on real-time motion planning for robot table tennis, where simulated experiments demonstrate superior sample efficiency and closed-loop performance compared to state-of-the-art baselines. We open-source the used implementation of the mapping algorithm and OCP for motion planning https://github.com/SonyResearch/feasible action for optimal control.
Sources
- Synthesis of Model Predictive Control and Reinforcement Learning: Survey and Classification
- Hyperspherical Normalization for Scalable Deep Reinforcement Learning
- 1000 Layer Networks for Self-Supervised RL: Scaling Depth Can Enable New Goal-Reaching Capabilities
Related papers
- One Request, Multiple Experts: LLM Orchestrates Domain Specific Models via Adaptive Task Routing
- A Geometric Decision Procedure for STL Feasibility and Repair
- Submodular Multi-Agent Policy Learning for Online Distributed Task Allocation in Open Multi-Agent Systems
- Policy-Level Recursive Self-Improvement for Embodied AI with a Criticality World Model
- Minimal Experiments for Robust Stabilization: Information, Spectral Geometry, and Duration
- Decentralized Power-Optimal Coordination for Spacecraft Swarms Using Time-Varying Magnetorquer Actuation