FAME: Force-Adaptive RL for Expanding the Manipulation Envelope of a Full-Scale Humanoid
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.
Dev: Today's paper: "FAME: Force-Adaptive RL for Expanding the Manipulation Envelope of a Full-Scale Humanoid".
Rosa: Maintaining balance under external hand forces is critical for humanoid bimanual manipulation, where interaction forces propagate through the kinematic chain and constrain the feasible manipulation envelope.
Dev: First, who's behind it and why it matters.
Title and authors: Rosa: So we're diving into this paper called "FAME: Force-Adaptive RL for Expanding the Manipulation Envelope of a Full-Scale Humanoid." It sounds like they are tackling that really tricky problem of keeping a humanoid balanced when you're also actively manipulating things with both hands, which is something we see constantly in complex robotics.
Dev: Exactly, Rosa. The title suggests they are trying to expand what the robot can actually do and where it can safely operate by making it aware of those external hand forces affecting its balance throughout the entire kinematic chain.
Taro: I'm interested in how this addresses the issue of things going wrong when you're doing something complex; specifically, how does this system handle situations where the world misbehaves while you're actively manipulating objects?
Rosa: Well, basically, they propose a force-adaptive reinforcement learning framework that conditions the policy on a learned context vector that captures both where the upper body joints are and what forces those hands are currently exerting.
Dev: That context vector is key because it lets the base standing policy adjust its lower-body control strategy based on the current loading condition, which means it can adapt in real time.
Taro: That sounds promising for handling unexpected disturbances, but I wonder how robust this adaptation is when those forces are coming from something unpredictable, like a sudden push or an object shifting unexpectedly.
Rosa: They address that uncertainty by training the system with diverse three dee forces applied to each hand in simulation and using an upper-body pose curriculum to gradually increase the difficulty of those scenarios.
Dev: The paper mentions that this training method helps expose the policy to manipulation-induced perturbations, which is necessary because simply training on a fixed set of scenarios wouldn't teach it how to handle a wide range of force interactions.
Taro: So it’s not just about learning a specific balancing maneuver for one pose, but learning an adaptable strategy that works across many different arm configurations and force magnitudes?
Rosa: That’s right; the core idea is that the robot learns a compact representation of those state variations caused by manipulation forces so it can adapt its lower-body balance instantly.
Dev: From an engineering standpoint, I'm looking at the latency here, and they show how this encoder operates during deployment using measured joint torques to estimate these hand forces without needing specialized sensors on the wrists.
Taro: That sensor-free estimation part is interesting; if it can infer those forces just from measuring the robot's dynamics, that opens up a lot of possibilities for practical deployment where adding extra hardware isn't an option.
Rosa: Absolutely, that inference mechanism allows the system to maintain stability even when it’s operating in a real environment without relying on expensive wrist force/torque sensors.
Title and authors: Dev: But we have to consider the loop rate; how fast does this context encoding and subsequent policy conditioning happen when a disturbance is sudden? The paper implies rapid adaptation, but I want to know the practical response time for stabilization.
Taro: If the system can handle asymmetric single-arm loads or symmetric bimanual loads effectively, that moves it closer to real-world scenarios where we expect forces to be highly variable and coupled.
Rosa: That’s what they demonstrate in their experimental validation, showing stability over a larger admissible force region than previous methods, especially in challenging configurations like C1 and C5.
Dev: And those results are encouraging because they show success rates of seventy-three point eight four percent in simulation across five fixed upper-body arm configurations with randomized disturbances.
Taro: I’m curious about the limitations, though; the paper does flag that it relies on a specific formulation of the latent context and the fidelity of that force estimation, which is something we need to watch closely as we deploy this outside of controlled lab settings.
Rosa: That’s a fair point; they are honest about where their method stops working, and knowing those boundaries is crucial for understanding its applicability in a field roboticist's view.
Dev: If it works outside the lab, how long do you think this adaptive balance stays stable before the system might need some form of explicit retraining or recalibration?
Taro: The future work section suggests further exploration into how this context encoding can generalize to completely novel manipulation tasks that weren't explicitly covered in their training curriculum.
Rosa: That points toward a major implication: if this framework proves generalizable beyond the specific tasks it was trained on, it could significantly broaden the range of physically demanding humanoids we can design for.
Dev: I think the impact on control engineering is significant because it moves us toward systems that are inherently aware of their interaction forces rather than just reacting to joint errors alone.
Taro: It suggests a path where autonomous systems don't just react to immediate physical constraints but proactively manage their entire operational envelope based on predicted or sensed interaction loads.
Rosa: So, to wrap up this discussion on "FAME: Force-Adaptive RL for Expanding the Manipulation Envelope of a Full-Scale Humanoid," we see a system that learns to be context-aware about both its body configuration and external forces.
Dev: It shows how incorporating structured learning from coupled state variations can lead to more intelligent stabilization strategies, especially when combined with sensor-free force estimation.
Taro: The potential for generalized force-adaptive control across different manipulation scenarios is what really excites me about this paper’s direction.
Rosa: Indeed, the implication is that we can design humanoids capable of performing complex, force-intensive tasks with much greater stability and safety in unstructured environments.
The paper's summary: Rosa: So, essentially, FAME is about developing a way for humanoid robots to stand stably while they’re simultaneously dealing with external forces from manipulating objects, and it does this by learning how those upper body movements and hand forces are coupled together in a special context representation.
Dev: That coupling aspect is what really caught my attention; if the AI can accurately encode that relationship between joint positions and applied forces, it should be able to anticipate balance issues before they even happen, which is a big step toward reliable control.
Taro: I'm thinking about the implications for autonomy here because this context encoding means the system isn't just reacting to one thing; it’s understanding the whole physical situation at once so it can make smarter decisions when things get messy.
Rosa: Exactly, and they show that this learned context allows for a much wider range of movements and force interactions than older policies could handle safely, expanding what these robots are actually capable of doing in the real world.
Dev: From a control engineering standpoint, the ability to adapt the lower-body strategy based on that force context in real time is impressive because it addresses latency issues inherent in traditional feedback loops when dealing with dynamic loads.
Taro: That real-time adaptation is where I see the most potential for autonomy; imagine a robot suddenly bumped or has an unexpected load shift, and it immediately adjusts its stance without needing a lengthy re-planning process.
Rosa: And that's because they trained the system on diverse force scenarios, which means when it encounters something new in deployment, it has learned enough underlying principles to make a reasonably safe adjustment.
Dev: That brings up the real-world question for me—how long can we trust this adaptation before we need to retrain the policy entirely if the environment changes drastically?
Taro: The paper suggests that by using a curriculum that gradually increases the complexity of those force scenarios, they are building a system that generalizes better, which means it might last longer in varied conditions than policies trained only on simple tasks.
Rosa: That’s what I'm curious about for deployment; we need to know if this robust adaptation holds up under long-term, unpredictable operational stress outside of a perfectly controlled simulation environment.
Dev: If the force estimation method works reliably without those bulky wrist sensors, that’s a huge win for practical hardware design and deployment logistics.
Taro: And when you consider the broader picture, this work suggests we can move toward humanoid robots that aren't just programmed to execute motions but are truly capable of managing complex physical interactions intelligently.
Rosa: It really feels like this is moving us closer to having more sophisticated mobile manipulators that can handle genuinely challenging, dynamic tasks in unpredictable settings.
The paper's improvements: Taro: So, to wrap up on where they suggest taking this work next, the focus is really on making sure this framework can handle a wider variety of real-world physical scenarios, which means focusing heavily on generalizing the learned context representation beyond the specific setups used in their training.
Rosa: I think that’s crucial because if it only works perfectly for C1 through C5 configurations, we need to know how it handles a completely new kind of manipulation or an unexpected external force pattern that wasn't modeled in those initial datasets.
Dev: From a control perspective, the future work seems to emphasize improving the fidelity of that sensor-free force estimation method so that the system can be even more precise when inferring loads during actual operation, rather than just relying on joint torque residuals.
Taro: I agree with Dev; better inference means less reliance on perfect simulation matching, which is a big hurdle for deployment in messy physical environments where friction and dynamics vary constantly.
Rosa: And I'm interested in the idea of online adaptation; if we can develop a mechanism where the policy can adjust its balance strategy instantaneously based on new force contexts as they happen, that would be incredibly useful for unpredictable interactions.
Dev: That instant adjustment capability is what makes it so appealing for high-speed or reactive tasks; if the system has low latency in processing that context vector and outputting a new control action, it can react to disturbances much faster than current methods allow.
Taro: Plus, looking at the broader implications, they are essentially trying to create a robot that is inherently safer because its balance isn't just based on pre-programmed stability margins but on an active understanding of the forces it’s currently experiencing.
Rosa: That brings us to the big picture—if this approach proves robust across asymmetric single-arm loads and symmetric bimanual loads, it could unlock a whole new class of robots capable of handling much more complex, hands-on industrial or assistive tasks.
Dev: The paper also hints at a need for better structural sign herdability in these temporal networks, which suggests that future work might involve designing the underlying control architecture itself to be more inherently stable under varying conditions.
Taro: That points toward a deeper level of theoretical work needed to ensure the system's stability isn't just achieved through clever RL tricks but is grounded in robust system theory.
Rosa: It sounds like they are pushing this framework beyond just balancing and into the realm of truly adaptive, multi-task physical interaction.
Dev: So, we’re looking at a path that combines advanced reinforcement learning with physics-based estimation to create systems that don't just follow instructions but actively manage their physical stability in dynamic situations.
Conclusion: Rosa: So, to wrap up, we’ve seen how FAME tackles the problem of maintaining balance while actively manipulating objects by using a learned context to adapt control in real time.
Dev: It really shows how incorporating structured learning from coupled state variations can lead to more intelligent stabilization strategies when dealing with external forces.
Taro: I think the most significant implication is that we’re moving toward robots that are not just executing pre-programmed motions but are truly capable of managing complex physical interactions intelligently in dynamic environments.
Rosa: That capability, especially with the sensor-free force estimation, opens up so many possibilities for deployment outside of perfectly controlled lab settings.
Dev: I’m still focused on the loop rate and latency; if we can keep this adaptation happening fast enough to handle sudden disturbances, that really changes how responsive these systems are in practice.
Taro: And when you consider the potential for generalized force-adaptive control across different manipulation scenarios, it suggests a path where autonomous systems can handle much more varied physical challenges safely.
Rosa: I’m excited about the potential for this to be applied to everything from delicate assembly to more robust assistive tasks in unstructured settings.
Dev: That robustness is key; if the failure modes are manageable and we understand when the system might need explicit recalibration, that makes it a much more practical piece of control engineering.
Taro: I just think this work sets a really high bar for how we should be thinking about autonomy in humanoids, focusing on understanding the underlying physics of interaction rather than just reacting to errors.
Rosa: That’s the big picture, and it really demonstrates how much progress we're making toward creating more capable physical agents.
Dev: It’s a solid paper that bridges the gap between complex RL and practical control system requirements.
Taro: We definitely need to keep an eye on how this context encoding generalizes to completely novel manipulation tasks, because that’s where the real autonomy is going to be tested.
University of Colorado Boulder
cs.RO
Submitted: 2026-03-09
Updated: 2026-10-02
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 89/100
The gist: Maintaining balance under external hand forces is critical for humanoid bimanual manipulation, where interaction forces propagate through the kinematic chain and constrain the feasible manipulation
Key concepts
- Force Adaptation
- This mechanism uses a learned latent context vector (zt) derived from upper-body joint states and hand forces to condition the base policy. This allows the robot's lower-body control to dynamically adjust its balance strategy in real time, effectively learning a compact representation of how manipulation disturbances affect stability.
- Upper-Body Pose Curriculum
- This training technique gradually exposes the system to more challenging upper-body poses. The difficulty increases by expanding the range of randomized target poses, and this expansion is controlled by a ratio (rhoa) that only increases when the robot achieves better standing quality metrics, ensuring safe and progressive learning.
- Sensor-Free Force Estimation
- FAME estimates interaction forces without needing wrist force/torque sensors. It uses joint torques ($ au$) and gravity compensation torques ($ au_g$), combined with the wrist Jacobian ($J$), to calculate the external forces ($F_{ext}$) acting on the robot, enabling deployment in scenarios where such sensors are unavailable.
Terminology
Summary
Maintaining balance under external hand forces is critical for humanoid bimanual manipulation, where interaction forces propagate through the kinematic chain and constrain the feasible manipulation envelope. FAME proposes a force-adaptive reinforcement learning framework that conditions a standing policy on a learned latent context encoding upper-body joint configuration and bimanual interaction forces, enabling robust balance across diverse arm configurations.
The gist: FAME is a force-adaptive RL framework that conditions a standing policy on a learned latent context encoding upper-body joint configuration and bimanual interaction forces, enabling robust balance through diverse arm configuration scenarios.
How it works
The framework consists of two main components: an upper-body context encoder and a base standing policy. The encoder processes upper-body joint states together with hand interaction forces to produce a latent context vector that conditions the base policy, allowing it to adapt its lower-body control strategy according to the current upper-body loading condition.
During training, the system exposes the policy to diverse manipulation-induced disturbances by:
-
Sampling and applying external forces at each hand.
-
Randomizing upper-body target poses through an
upper-body pose curriculum
that gradually expands the pose range as standing quality improves.
The upper-body context encoder maps an input vector, defined as the concatenated state of torso and arm joint positions and sampled hand forces, to a latent context vector:
xenc = [qub, FL, FR] where qub ∈ R15 is the torso+arm joint positions and FL, FR ∈ R3 denote the left and right wrist forces.
The encoder maps this input to a latent context vector:
zt = µθ(xenc) ∈ R8 where θ denotes encoder parameters for a multi-layer perceptron (MLP).
Key Mechanisms
FAME introduces several key mechanisms to achieve force adaptation and sensor-free deployment:
-
Force Adaptation: The latent context vector, zt, conditions the base policy. This allows the policy to learn a
compact representation of manipulation-induced perturbations
and adapt lower-body balance in real time. -
Upper-Body Pose Curriculum: This curriculum progressively increases the perturbation range of randomized target poses by maintaining a scalar upper-body action ratio ρa ∈ [0, 1], which is increased when the standing quality metric (height-tracking reward) exceeds a threshold.
-
Sensor-Free Force Estimation: At deployment, interaction forces are estimated from robot dynamics without requiring wrist force/torque sensors. This is achieved by measuring joint torques (τ), computing gravity compensation torques (τg), and mapping the residual joint torques into Cartesian space using the wrist Jacobian (J):
Fext = −(J⊤)†(τ − τg) where † denotes the pseudo-inverse.
Training and Policy Formulation
The base policy is formulated as a Partially Observable Markov Decision Process (POMDP), utilizing Proximal Policy Optimization (PPO) to learn a parameterized policy πθ(at ot). The actor network consists of an estimator network E and a follow-up policy network N. The input to the actor includes the one-step proprioceptive observations, the latent context vector zt, and the estimated disturbance I hat t:
oπt = [opropt, oztt] where opropt includes command (scaled), IMU angular velocity (body frame), projected gravity (body frame), joint position error (scaled), and joint velocity (scaled).
The critic network uses a privileged one-step vector that augments the observation with the base linear velocity available in simulation. The policy outputs a 12D action at ∈ R12, corresponding to target joint position offsets for the lower-body joints, which are tracked by a joint-space PD controller.
Experimental Validation
The framework was validated in simulation across five fixed upper-body arm configurations (C1–C5) with randomized 3D hand-force disturbances and commanded base heights. The mean standing success rate achieved by FAME was 73.84%, compared to 51.40% for the Base+Curr baseline and 29.44% for the Base policy (no curriculum or encoder).
In real-world experiments on a Unitree H12, FAME demonstrated robustness under representative load-interaction scenarios:
RE1: Asymmetric Single-Arm Load tests robustness to an asymmetric disturbance where one arm is carrying a load.
RE2: Symmetric Bimanual Load evaluates stability as the carried weight is distributed to both arms.
The results showed that FAME maintained stability over a larger admissible force region compared to the Base+Curr policy, achieving significant success in challenging cases like C1 (Forward Extended) and C5 (Asymmetric Forward Full), where the Base policy collapsed.
Improvements for AI systems
Based on the scientific paper FAME: Force-Adaptive RL for Expanding the Manipulation Envelope of a Full-Scale Humanoid,
here are specific improvements that can be made to AI systems, categorized by capability:
) 1. Robust Bimanual Manipulation Under Uncertainty: The improved system can maintain stable standing while simultaneously performing complex, force-intensive tasks (like carrying or manipulating objects) with external hand forces applied at the wrists.
-
Sensor-Free Force Estimation for Real-Time Adaptation: The improved system can estimate the interaction forces exerted by an object or another agent (e.g., via joint torques and robot dynamics) in real time, eliminating the need for bulky wrist force/torque sensors, which is crucial for mobile and deployment environments.
-
Expanded Feasible Manipulation Envelope: The improved system can operate safely within a significantly larger region of arm configurations and external hand forces—the
manipulation envelope
—than traditional policies, allowing for more complex and physically demanding tasks. -
Online Adaptation to Dynamic Loads: The improved system can adapt its lower-body balance control strategy instantaneously based on the current upper-body configuration and applied force context (latent representation), enabling rapid stabilization against sudden or varying external disturbances without needing to be explicitly retrained for every new force scenario.
-
Structured Learning from Coupled State Variations: The improved system learns a compact, latent representation that explicitly captures the coupling between upper-body joint configurations and bimanual interaction forces, allowing the policy to disambiguate whether a change in balance is due to arm pose or applied force, leading to more intelligent and less conservative stabilization strategies.
-
Generalized Force-Adaptive Control: The system can handle diverse manipulation scenarios, including asymmetric single-arm loads (e.g., pulling) and symmetric bimanual loads (e.g., carrying a balanced load), achieving high success rates across all these challenging scenarios in both simulation and real hardware.
Abstract
Maintaining balance under external hand forces is critical for humanoid bimanual manipulation, where interaction forces propagate through the kinematic chain and constrain the feasible manipulation envelope. We propose FAME, a force-adaptive reinforcement learning framework that conditions a standing policy on a learned latent context encoding upper-body joint configuration and bimanual interaction forces jointly, since the base moment a load induces depends on the arm configuration through which it acts. Training applies isotropically sampled 3D forces at each hand under an upper-body pose curriculum, exposing the policy to manipulation-induced perturbations across continuously varying arm configurations. At deployment the interaction force is not measured but reconstructed online from joint torques and states through rigid-body inverse dynamics, requiring no wrist force/torque sensing. We evaluate over 100 upper-body configurations under swept hand forces, scoring each trial by a task-level criterion that requires the robot both to remain upright and to hold its hands near where the task placed them; all such results run with the estimated force in the loop. At a 150,mm tolerance FAME reaches 38.9% task success, against 16.6% for a policy given the same force without encoding, 4.3% for a pose-conditioned curriculum policy, and 24.7% for an adversarially trained locomotion policy, which stays upright but recovers by stepping and so relocates the hands. We further demonstrate transfer to task-generated interaction forces in a MuJoCo kitchen environment, and to asymmetric and bimanual loading on a full-scale Unitree H1-2. Code and videos are available on the https://correlllab.github.io/fame website.
Sources
- FALCON: Learning Force-Adaptive Humanoid Loco-Manipulation
- RSL-RL: A Learning Library for Robotics Research
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving