CAPABLE: Capability-Aware Policy Adaptation via Behavioral Latent Encoding

arXiv:2610.11971 · cs.RO, cs.LG, cs.SY, eess.SY · Submitted 2026-10-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.

Dev: Today's paper: "CAPABLE: Capability-Aware Policy Adaptation via Behavioral Latent Encoding".

Rosa: The gist The CAPABLE framework introduces a unified capability-aware adaptation framework for frozen Vision-Language-Action (VLA) policies that integrates self-supervised capability inference with residual reinforcement learning to enable fault recovery without…

Dev: First, who's behind it and why it matters.

Title and authors: Rosa: So we're looking at this paper called "CAPABLE: Capability-Aware Policy Adaptation via Behavioral Latent Encoding," and the title itself suggests they're trying to figure out how an AI policy can adapt when a robot joint actually breaks.

Dev: Yeah, it’s about tackling that problem directly. They introduce this framework that lets you recover from actuator failures without needing any pre-labeled fault information or knowing exactly which joint failed.

Taro: That sounds like a big deal for real-world deployment because usually, when something goes wrong on a robot, you need to stop everything and diagnose the exact component affected before you can try to fix it.

Rosa: Exactly, and what makes this CAPABLE approach different is how it infers that capability. Instead of relying on explicit fault labels, they are using command-response history and the robot's current kinematics to figure out what each joint can still actually do.

Dev: And they do this by using a shared temporal encoder across all the joints, which means the system learns common movement patterns instead of having a separate brain for every single joint.

Taro: That shared encoding idea sounds smart because it suggests that if one actuator is missing something, the information about how that failure manifests should be recognizable across other joints as well.

Rosa: Right, and they ground this estimate by using the live Jacobian column for each joint, which ties the abstract history to the robot's physical configuration right at that moment.

Dev: That kinematic grounding is crucial because it ensures the capability inference is based on what’s physically possible *now*, not just some historical averages.

Taro: It seems like they’re trying to build a model of physical potential that changes dynamically as the robot moves or fails, which is pretty advanced for autonomy.

Rosa: And then they use this inferred capability representation to condition a residual policy, which adds small, bounded corrections to the original frozen VLA action without changing the main policy itself.

Dev: So the core idea is that you keep your main AI policy frozen and just let this new module make tiny adjustments based on what you think the robot can still achieve.

Taro: That residual learning approach is interesting because it suggests you don't have to retrain the entire massive VLA system just because one part of the hardware changed.

Rosa: And they tie this whole capability representation to actual physical behavior using self-supervised prediction loss functions, which includes predicting both the joint displacement and the end-effector movement.

Dev: That physics grounding is what makes it more than just a clever statistical guess; they’re forcing the learned capability representation to match how things actually move in space.

Title and authors: Taro: So, if you have a failure, this framework isn't just guessing; it’s predicting the physical outcome of that failure using those prediction losses.

Rosa: Right, and for testing, they used a benchmark like LIBERO with the Franka Panda robot implementation in robosuite and MuJoCo to see how it performs in practice.

Dev: They test it by sampling tasks uniformly, but they specifically sample healthy executions only about ten percent of the time, while usable faults make up the other ninety percent.

Taro: That setup sounds realistic because you don't get enough clean data when things are supposed to be broken all the time in a real-world scenario.

Rosa: The results show that CAPABLE raises success on an actuator excluded from fault training from around twenty-five percent up to fifty-nine point three percent, which is a significant jump.

Dev: They also compared this against a global history baseline and found CAPABLE outperforms it by fourteen point four points in terms of mean success across tasks.

Taro: That transfer advantage they showed isn't just for one specific joint; it works across all six joints when testing on held-out actuators, which is pretty strong generalization.

Rosa: What I find particularly interesting about their ablation studies is that they show the coupled temporal encoding and the kinematic grounding are what actually create that transfer advantage.

Dev: Removing just the history leaves you with a memoryless residual controller using only current Jacobian information, and that performance drops below another baseline, Global-history SAC.

Taro: So it really points to how important it is to have that combination of shared temporal encoding and live kinematic grounding for making this work robustly.

Rosa: Hardware validation on the physical Franka Panda showed CAPABLE reaching eighty-five point zero percent success in two seen-joint conditions, compared to forty-five point zero percent for the global history SAC baseline.

Dev: And even when it comes to unseen joints, CAPABLE hit seventy point zero percent against forty point zero percent for that specific condition.

Taro: That hardware transfer is a key piece of evidence because it means this capability inference mechanism works without needing any extra training just because the robot body itself is different.

Rosa: So, to sum up the main idea of CAPABLE: it represents actuator faults through what each joint can still physically realize, rather than giving you a specific label for the fault.

Dev: It uses that shared temporal encoder and live kinematics to generate bounded residual control that corrects the frozen VLA action.

Title and authors: Taro: This capability formulation lets you recover from failures without needing those explicit fault labels or knowing which joint is faulty.

Rosa: Across twenty-eight LIBERO tasks, they showed this approach outperforms a parameter-matched global-history baseline and performs well on unseen fault families and hardware variations too.

Dev: The main implication here is that we can have frozen VLA policies that are much more resilient when they encounter physical damage in the field.

Taro: It suggests a practical route to fault recovery for frozen VLAs by using factorized command-response representations instead of relying on manual supervision for every possible failure mode.

Rosa: We’ve been talking about how this framework works and why it’s effective, but what does this all mean for the next generation of embodied AI systems?

Dev: It means we can deploy these complex vision-language-action models in environments where hardware wear and tear are expected, without needing a massive retraining pipeline every time a component degrades.

Taro: For autonomy, this capability inference could become a core part of how the robot itself monitors its own health and dynamically adjusts its control strategy on the fly.

Rosa: It suggests that instead of treating failures as catastrophic events requiring full system re-diagnosis, we might be able to treat them as gradual changes in physical constraints that we can adapt to locally.

Dev: The challenge for us engineers is making sure this capability inference happens fast enough, so the residual correction doesn't introduce new latency or instability into the control loop.

Taro: That’s a fair point; we need to ensure that the complexity of inferring capability doesn't cripple the real-time execution speed of the policy.

Rosa: Well, that’s what we need to keep watching as this framework moves from simulation benchmarks into actual field robotics where things get messy and unpredictable.

Dev: We’ll keep tracking those loop rates and latency numbers as they move toward real hardware testing because that's where the rubber meets the road for us.

Taro: And I'm interested to see how they tackle those novel failure modes, like increased friction or damping, which are things you rarely see perfectly in lab setups.

Rosa: We’ll stay tuned to see if this method keeps performing well when we throw it at actuators that were never trained on during the initial setup.

Dev: I'm ready to look at the technical details on how they handle those prediction losses next, because that’s where the physics stuff gets really deep.

Taro: It looks like a lot of interesting work here, focusing on making frozen models more physically aware of their own limitations and adapting intelligently when things go wrong.

The paper's summary: Rosa: So, CAPABLE is basically this system that lets you recover from a robot breaking without ever needing to tell the computer *which* joint broke or give it any labels about the fault.

Dev: Right, so it figures out what each part can still do just by looking at the movement history and where the robot is right now, using this shared encoding idea we talked about.

Taro: And that's pretty powerful because instead of retraining the whole massive vision-language-action policy when a joint fails, you just add these small corrections on top.

Rosa: Exactly, it’s like giving the frozen VLA arm a tiny bit of extra intelligence tailored to what hardware it’s actually got left.

Dev: The core mechanism uses this capability representation to condition a residual policy that adds these bounded corrections to the arm action without changing the main policy itself.

Taro: That residual learning approach is interesting because it suggests you don't have to retrain the entire massive VLA system just because one part of the hardware changed.

Rosa: It’s about inferring what a joint can still realize based on how much of a commanded motion it actually manages to do, tying that directly into the robot's physical configuration through kinematics.

Dev: They use self-supervised physical prediction loss functions, which means they check if their inferred capability matches the actual physics of the robot moving in space.

Taro: That physics grounding is what makes it more than just a clever statistical guess; they’re forcing the learned capability representation to match how things actually move in space when something goes wrong.

Rosa: And for testing, they used benchmarks like LIBERO and Franka Panda to show that this works on real hardware, and the results show success jumping from about twenty-five percent up to fifty-nine point three percent when an actuator is excluded from training.

Dev: They also compared it against a global history baseline and found CAPABLE outperforms it by fourteen point four points in terms of mean success across those tasks.

Taro: What I find particularly interesting about their ablation studies is that they show the coupled temporal encoding and the kinematic grounding are what actually create that transfer advantage.

Rosa: Removing just the history leaves you with a memoryless residual controller using only current Jacobian information, and that performance drops below another baseline, Global-history SAC.

Dev: That really points to how important it is to have that combination of shared temporal encoding and live kinematic grounding for making this work robustly.

Taro: It suggests that instead of treating failures as catastrophic events requiring full system re-diagnosis, we might be able to treat them as gradual changes in physical constraints that we can adapt to locally.

Rosa: So, to sum up the main idea of CAPABLE: it represents actuator faults through what each joint can still physically realize, rather than giving you a specific label for the fault.

Dev: It uses that shared temporal encoder and live kinematics to generate bounded residual control that corrects the frozen VLA action.

Taro: This capability formulation lets you recover from failures without needing those explicit fault labels or knowing which joint is faulty.

Rosa: Across twenty-eight LIBERO tasks, they showed this approach outperforms a parameter-matched global-history baseline and performs well on unseen fault families and hardware variations too.

Dev: The main implication here is that we can have frozen VLA policies that are much more resilient when they encounter physical damage in the field.

Taro: It suggests a practical route to fault recovery for frozen VLAs by using factorized command-response representations instead of relying on manual supervision for every possible failure mode.

Rosa: We’ve been talking about how this framework works and why it’s effective, but what does this all mean for the next generation of embodied AI systems?

Dev: It means we can deploy these complex vision-language-action models in environments where hardware wear and tear are expected, without needing a massive retraining pipeline every time a component degrades.

The paper's improvements: Taro: So we've covered how CAPABLE figures out what a robot can still do based on its physical capability, and now we're looking at how they plan to make this system even better for real use.

Rosa: The authors suggest that the coupling between the shared temporal encoding and the live kinematic grounding is actually what produces that strong transfer advantage we saw in their tests.

Dev: So they aren't just tweaking one part; they’re emphasizing that you need both the history and the current physical state working together for this to be reliable.

Taro: And on top of that, they also show robustness against different kinds of failures, like increased friction or damping, by training on six distinct fault families.

Rosa: That means even if you encounter a failure mode not seen during initial training, the system has a better chance of adapting because it learned the underlying physical principles rather than just memorizing specific scenarios.

Dev: The implication here for an engineer is that they're moving away from brittle models where you need perfect fault labels for every possible thing the robot could break into.

Taro: It’s about creating a system that can handle messy, unpredictable real-world environments where you don't have access to a perfect fault catalog.

Rosa: They also highlight that this approach works on unseen actuators too, which is huge because in the field, you often don't know what hardware is actually installed until it breaks.

Dev: So the goal isn't just to fix known problems; it’s to build a controller that can handle new hardware constraints without needing a complete re-calibration of the entire learning setup.

Taro: That shifts the focus from perfect control in a controlled lab setting to practical resilience in messy, deployed systems.

Rosa: It suggests that for autonomous systems to work reliably in the real world, they need to be inherently capable of understanding and adapting to physical limitations on their own.

Conclusion: Rosa: So we’ve seen how CAPABLE uses behavioral latent encoding to infer physical capability from command responses and kinematics, and it seems like this framework is genuinely pushing frozen VLA policies into a new level of resilience.

Dev: It really does show that you don't need explicit fault labels to recover from actuator issues, just the ability to map those behaviors back to a physical constraint.

Taro: I think what’s most important here is that this capability inference isn't just an academic exercise; it’s a tool for real autonomy where hardware fails unexpectedly in the field.

Rosa: Exactly, because they demonstrated transferability across different joints and even hardware setups without needing any new training data specific to those failures.

Dev: The numbers we saw on the Franka Panda validation—like reaching eighty-five percent success on seen joint conditions—show that this isn't just theoretical; it holds up under physical stress.

Taro: It’s a practical solution because it means a robot can essentially "self-diagnose" its remaining capabilities and adjust its actions accordingly.

Rosa: And the fact that they managed to achieve this with bounded residual corrections, meaning the main policy stays frozen, is really neat for deployment speed.

Dev: From an engineering standpoint, that residual RL setup keeps the loop rate manageable because you’re only adding small adjustments on top of the existing control signal.

Taro: It makes sense that they focus on those prediction losses; they’re grounding the abstract capability in actual physics, which is what separates this from just guessing.

Rosa: So, to wrap up, CAPABLE is a method for recovering from actuator faults by representing a fault as remaining physical capability rather than an explicit joint identifier.

Dev: That shared temporal encoder and kinematic grounding are clearly the core mechanisms that give them that transfer advantage we discussed earlier.

Taro: It’s a big step toward making embodied AI systems more robust in environments where things aren't perfectly controlled or perfectly built.

Rosa: We’ve seen how this works across LIBERO tasks, and it really proves that factorized command-response representations are a practical route for frozen VLAs to handle physical damage.

Dev: I think the next challenge is making sure that capability inference happens fast enough so there isn't any extra latency in the control loop when a fault actually occurs.

Taro: And I’m looking forward to seeing how this extends beyond just joint faults into other types of physical wear or degradation.

Mohammad Khoshnazar, Mohammad Dehghani Tezerjani, Deyuan Qu, Zhiyuan Gao, Yanxiang Zhan, Jeroen Schafer®, Andrew Melnik

University of Bremen · University of North Texas

cs.RO, cs.LG, cs.SY, eess.SY

Submitted: 2026-10-08

Updated: 2026-10-08

The gist: The gist The CAPABLE framework introduces a unified capability-aware adaptation framework for frozen Vision-Language-Action (VLA) policies that integrates self-supervised capability inference with

Key concepts

Capability Inference
This process estimates what each joint can still do based on past command-response data and current physical state. It defines a fault not by labeling a specific joint as broken, but by quantifying the remaining physical potential of that actuator.
Shared Temporal Encoder
A single encoder processes command-response histories from all joints using identical weights. This allows the system to recognize common behavioral patterns across different actuators, creating a unified understanding of how the robot moves rather than learning separate models for each joint.
Self-Supervised Physical Prediction
The framework uses prediction tasks (like predicting joint displacement and end-effector movement) as loss functions. This ties the inferred capability representation to actual physical behavior, ensuring the learned policy generates corrections that are physically plausible.

Terminology

Summary

The gist The CAPABLE framework introduces a unified capability-aware adaptation framework for frozen Vision-Language-Action (VLA) policies that integrates self-supervised capability inference with residual reinforcement learning to enable fault recovery without requiring explicit fault labels or faulty joint identifiers.

How it works

CAPABLE infers the remaining capability of each joint from command–response history and kinematics, which represents a fault by what each actuator can still realize, and uses this representation to condition a residual policy that adds bounded corrections to the frozen VLA arm action without fault labels or faulty joint identifiers<ref:2610.11971#pg6>. This capability inference is achieved through several integrated components A per-joint command–response history is constructed by concatenating the joint position, velocity, and action chunk phase with the live Jacobian column to form a token.

How it works

The core of CAPABLE involves a shared temporal encoder that processes these per-joint histories using identical weights across all joints This shared encoder recognizes common command–response signatures across actuators rather than learning a separate encoder per joint. The resulting representation is then combined with the current kinematic query and projected through a Transformer encoder to form a capability latent vector, z cap t This z cap t conditions a Soft Actor-Critic policy through FiLM modulation to generate bounded corrections.

How it works

The framework incorporates self-supervised physical prediction to tie this representation to physical behavior through several loss functions It utilizes a shared local decoder to predict realized joint displacement and a global decoder to predict realized end-effector displacement. The loss functions include Ljoint, Leef, Lkin, and Lcap Specifically, the kinematic consistency loss enforces physical behavior by requiring the predicted joint displacements mapped through the Jacobian to match observed end-effector displacement.

How it works

The actor and critics utilize a twin-critic Soft Actor-Critic setup where the residual observation o RL t is 169-dimensional This observation includes measured end-effector position and orientation, gripper joints, arm state, action chunk phase, and history elements. The actor and critics receive the normalized observation oe RL t conditioned by FiLM modulation parameters γt and βt derived from z cap t.

How it works

The experimental setup involves training on a benchmark using LIBERO with the Franka Panda implementation in robosuite and MuJoCo The training procedure samples tasks uniformly, with healthy execution sampled with probability 0.10 and usable seen faults sharing the remaining 0.90 in proportion to their screened headroom s H m − s m j. Evaluation compares zero correction with each learned policy on healthy execution, every usable seen fault, and globally unseen j2.

How it works

The results demonstrate that CAPABLE raises success on an actuator excluded from fault training from 24.8% to 59.3%, outperforming a parameter-matched global-history baseline by 17.4 points while preserving healthy performance. Furthermore, the transfer advantage is not specific to one actuator across six joints, as CAPABLE outperforms Globalhistory SAC on all six held-out actuators.

How it works

The mechanism ablation studies support the coupled temporal–kinematic representation shaped by physical prediction rather than any single component. Removing history leaves a memoryless residual with current Jacobian information at 39.4% on unseen j2, below Global-history SAC. The ablation results indicate that the combination of shared temporal encoding and live kinematic grounding is what produces the transfer advantage.

How it works

Hardware validation on a physical Franka Panda shows CAPABLE reaches 85.0% across the two seen-joint conditions against 45.0% for Global-history SAC, and 70.0% against 40.0% on the unseen-j2 condition. This hardware transfer demonstrates that the capability inference is transferable to physical embodiments without hardware training.

How it works

The conclusion is that CAPABLE represents actuator faults through remaining physical capability rather than explicit fault identity. A shared temporal encoder, live kinematics, and bounded residual control enable transfer to actuators excluded from fault training while preserving healthy performance. Across 28 LIBERO tasks, all six leave-one-actuator-out splits, three of five unseen fault families, and hardware, CAPABLE outperforms a parameter-matched global-history baseline. The results show that factorized command–response representations provide a practical route to fault recovery for frozen VLA policies without fault labels or affected-joint supervision.

The gist

CAPABLE represents actuator faults through remaining physical capability rather than explicit fault identity. The paper presents CAPABLE, a residual reinforcement learning controller that keeps the VLA frozen and adds a correction to its arm action using an online estimate of the robot’s remaining actuator capabilities. This method infers capability from command–response history and kinematics using a shared temporal encoder, cross-joint attention, and self-supervised physical prediction to generate bounded residual corrections. CAPABLE outperforms a parameter-matched global-history baseline across 28 LIBERO tasks and demonstrates transfer to unseen actuators and fault families without requiring any fault labels or explicit joint identifiers. This capability is represented by what each actuator can still realize, measured by how much of a commanded motion the joint actually realizes and how that motion affects the end effector in the current configuration. CAPABLE's success on an actuator excluded from fault training is 59.3%, significantly higher than the 24.8% achieved by the base VLA. CAPABLE's transfer advantage over Global-history SAC is +34.5 points on task-balanced mean success across 28 tasks. This capability formulation of actuator faults for frozen VLAs, which represents a fault by what each actuator can still realize, infers it from command–response history and kinematics without fault labels, and uses it to correct a frozen generalist through bounded residual RL.

Improvements for AI systems

  1. Bold header: Capability-Aware Fault Recovery for Frozen VLAs

This system can recover from actuator failures without fault labels or faulty joint identifiers by inferring capability from command–response history and kinematics. It achieves this by using a shared temporal encoder to recognize common command–response signatures across actuators rather than learning a separate encoder per joint.

  1. Bold header: Cross-Actuator Transfer Generalization

The system demonstrates transferability beyond the specific faulty joint, as shown by leave-one-actuator-out experiments across six joints show that this transfer is not specific to one actuator. This capability allows the model to generalize its learned fault recovery mechanism to an actuator whose failure was not observed during training.

  1. Bold header: Physics-Informed Residual Correction

The architecture uses a residual reinforcement learning controller that adds bounded corrections to the VLA arm action without fault labels or faulty joint identifiers. This correction is conditioned on a representation that combines shared per-joint history encoding, current kinematic grounding, and self-supervised physical prediction.

  1. Bold header: Robustness to Unseen Fault Families

The system can generalize to novel failure modes by being trained on persistent locks and then tested against six fault families including increased viscous damping, increased Coulomb friction, retained actuator authority and joint-range fractions in 0.75, 0.50, and 0.25.

Sources

Related papers