SAPS: Shared Autonomy for Policy Steering by Blending Teleoperation with a Pretrained VLA
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "SAPS: Shared Autonomy for Policy Steering by Blending Teleoperation with a Pretrained VLA".
Dev: Recent advancements in Vision-Language-Action (VLA) models demonstrate impressive generalist capabilities in robot manipulation, yet these policies can be brittle under out-of-distribution spatial and semantic perturbations.
Rosa: First, who's behind it and why it matters.
Paper summary: Rosa: So, we're diving into SAPS: Shared Autonomy for Policy Steering by Blending Teleoperation with a Pretrained VLA. The main idea here is that these powerful Vision-Language-Action models are good at manipulation but they get really fragile when the environment throws something unexpected at them spatially or semantically.
Dev: Exactly, and the paper claims they can be brittle under those out-of-distribution perturbations, which is a real problem for deployment.
Taro: And what makes this work interesting is that instead of needing complex auxiliary models or messing with how the policy samples things, SAPS proposes a way to blend human commands with the pre-trained policy actions right at the action level.
Rosa: So, it sounds like this approach aims to give these generalist policies a way to handle unexpected situations without needing massive retraining efforts.
Dev: That's what they claim; they introduce a framework that requires no policy retraining, auxiliary dynamics models, or architectural modifications for steering <ref:2606.15568#pg0>.
Taro: It’s interesting because it moves the steering mechanism post-inference, which is different from methods that try to inject guidance inside the policy's sampler or decoding loop <ref:2606.15568#pg2>.
Rosa: That distinction between where you apply the steering really highlights how SAPS fits into the existing landscape of inference-time policy steering techniques.
Dev: The core mechanism they propose involves computing a final blended action using a formula like a(one:six) blended = alpha a(one:six) VLA + (one-alpha) a(one:six) expert.
Taro: That blending coefficient alpha is what determines the degree of autonomy, and they explore three different arbitration strategies to manage that balance.
Rosa: Three strategies? I'm curious how those specific methods—Fixed Blending, Full Takeover, and the Dynamic Cosine-Similarity Strategy—actually work in practice when things go wrong.
Dev: The Fixed Blending uses a constant coefficient, like alpha = zero point five, during active human intervention and then switches to full autonomy once the expert action's L2 norm drops below a threshold of epsilon=zero point zero zero one <ref:2606.15568#pg1>.
Taro: And then there's the Full Takeover strategy, which lets the human operator completely override the VLA policy whenever they detect active intervention, moving alpha from zero point zero to one point zero.
Rosa: That sounds like a very direct way to handle immediate control situations, but how does that compare when things are just slightly off-distribution?
Dev: The third strategy is the Dynamic Cosine-Similarity Strategy, which scales autonomy based on the geometric agreement between the human and policy actions using a logistic transformation to set alpha = gamma = sigma(k (theta)) where k=six <ref:2606.15568#pg1>.
Taro: That dynamic scaling based on cosine similarity sounds like it could allow for really smooth, continuous transitions between human and autonomous control, which is something prior methods might struggle with <ref:2606.15568#pg2>.
Paper summary: Rosa: Smooth transitions are important if we want the robot to feel responsive rather than abruptly switching control modes when the situation shifts slightly.
Dev: The results they present across various evaluations show that shared autonomy genuinely improves performance over just running the autonomous execution baseline. For example, in the LIBERO perturbation study, when they used Blending and Cosine arbitration, their success rates hit eighty-seven point nine percent and ninety point zero percent, whereas pi zero point five dropped to only twenty-two point nine percent at a perturbation distance of zero point one five m <ref:2606.15568#pg1>.
Taro: That drop for pi zero point five highlights the brittleness of the pure VLA model under spatial and semantic changes, which is exactly what SAPS seems designed to mitigate <ref:2606.15568#pg0>.
Rosa: And looking at the more complex LIBERO-PRO evaluation, Cosine achieved a mean success rate of ninety-seven point four percent across ten out-of-distribution tasks, and it even approached pure Teleoperation at ninety-eight point eight percent, significantly better than pi zero point five 's fifteen point zero percent mean success rate <ref:2606.15568#pg1>.
Dev: Completion times also showed improvement; for instance, in LIBERO-PRO, Cosine achieved a mean completion time of eleven point one s, compared to thirty point seven s for pi zero point five and forty-six point zero s for pure Teleoperation across all tasks <ref:2606.15568#pg2>.
Taro: These metrics suggest that SAPS isn't just surviving the failures; it’s actively helping the robot complete the task much faster when things get tough, which is a key aspect of useful autonomy.
Rosa: It really shows that this framework works well in controlled simulation environments like LIBERO and CALVIN, which are great starting points for testing these ideas.
Dev: The paper then extends these findings to real-world hardware evaluations on the Franka robot, where Cosine achieved an average success rate of ninety-eight point three percent across three tasks: Pick and Place, Close Drawer, and Open Cabinet <ref:2606.15568#pg2>.
Taro: It’s important that they show transferability to physical execution because simulation results don't always translate perfectly to the real world, Rosa needs to know how robust this is outside of a lab setting.
Rosa: And what the authors emphasize in their hardware section is that both shared-autonomy methods substantially reduce human intervention relative to Teleoperation across all three tasks, with Cosine and Blending reducing intervention by thirty to fifty percent <ref:2606.15568#pg2>.
Dev: That reduction in human input while the policy still contributes learned manipulation behavior during physical execution is a strong claim that speaks directly to practical deployment challenges.
Taro: It means operators don't have to micromanage everything; they just provide sparse corrective input, and the policy handles the rest of the learned skill <ref:2606.15568#pg0>.
Rosa: So, when we look at these real-world transfers, it seems like SAPS offers a practical middle ground between total human control and letting the robot figure everything out on its own.
Dev: The authors did note a limitation regarding the reliance on human input; they state that SAPS performance depends heavily on the timing and skill of the human operator <ref:2606.15568#pg2>.
Paper summary: Taro: That's a fair point, it means if the human isn't skilled or if their timing is off, the blending mechanism might not function as intended for optimal recovery <ref:2606.15568#pg2>.
Rosa: So, while the framework is model-agnostic and requires no retraining for deployment, its success still hinges on having a reasonably competent human operator providing those corrective inputs in real-time.
Dev: That dependency on the operator's skill is something engineers have to consider when designing the operational environment around such a system.
Taro: Looking ahead, this framework suggests that we can deploy generalist VLA policies in messy, unpredicted real-world contexts by giving them an intelligent way to incorporate sparse human guidance rather than relying solely on pre-training for every scenario <ref:2606.15568#pg0>.
Rosa: The implication for field robotics is that we could see robots operating reliably in dynamic environments without needing constant, high-level human intervention for every single unexpected event <ref:2606.15568#pg2>.
Dev: Thinking about the broader impact, if we can reduce the human workload by thirty to fifty percent while maintaining high success rates in complex manipulation tasks, that opens up possibilities for deploying robots in settings where continuous human presence is not feasible <ref:2606.15568#pg2>.
Taro: It moves us closer to a scenario where foundation models can be truly useful tools, not just impressive demos in a lab setting <ref:2606.15568#pg0>.
Rosa: I think the title, Shared Autonomy for Policy Steering by Blending Teleoperation with a Pretrained VLA, really captures the essence of what this paper is proposing: combining the learned abilities of the AI with reliable human corrective input at a specific point in time during action execution <ref:2606.15568#pg1>.
Dev: It's a practical engineering solution that tackles the brittleness issue without demanding massive computational overhead or retraining cycles for every new failure mode encountered <ref:2606.15568#pg0>.
Taro: The future work will likely involve exploring how this dynamic cosine strategy performs when the environment introduces even more subtle, continuous perturbations, pushing those limits further than what they tested in LIBERO-PRO <ref:2606.15568#pg2>.
Rosa: It seems like this SAPS framework provides a very solid path forward for making these powerful VLA models reliable enough for actual robotic deployment in the messy physical world, provided we manage that human input aspect effectively <ref:2606.15568#pg2>.
Dev: The stability and latency of the action-level blending are certainly something I'd be watching closely as they move this toward higher frequency real-time control loops, because the success of this approach hinges on that precise timing you mentioned earlier <ref:2606.15568#pg2>.
Taro: And from my perspective, it shows that the combination of pre-trained proficiency and sparse human feedback is a viable route for robust autonomy in manipulation tasks <ref:2606.15568#pg0>.
Conclusion: Dev: I see it as a critical control signal injection point; the fact that it operates post-inference at the action level means we don't have to mess with the policy sampler or introduce massive computational overhead for auxiliary models, which is something I really like. Rosa, how does this shift in focus on steering compare to other methods we've seen?
Rosa: It’s about moving that control decision from a high-level planning stage down into the actual motor commands; I wonder how this affects the responsiveness when we're dealing with those fast, real-time loops Dev is concerned about. Taro, you’ve been looking at the autonomy side; what does this blending strategy actually accomplish when things go sideways?
Taro: The dynamic cosine similarity strategy, in particular, seems powerful because it allows for smooth transitions between human and policy control based on how well the human and AI actions align geometrically <ref:2606.15568#pg1>. That’s how we get that continuous steering when the environment misbehaves or presents an out-of-distribution situation.
Dev: Smoothness is good, but I still want to know about failure modes; Rosa, you asked if it works outside the lab—how long can we trust this blending mechanism to hold up in a truly messy physical environment before those timing dependencies become a problem?
Rosa: That’s my main concern; the authors admit that performance still depends on the human operator's timing and skill, which means we need to figure out how robust this is when the human isn't perfectly synchronized with the robot’s movement. Taro, what about the broader impact if we can deploy these policies reliably in unpredictable settings?
Taro: The implication is that generalist VLA models could become much more useful tools for physical robots instead of just impressive lab demos, because they can incorporate sparse human guidance to handle unexpected situations without requiring constant retraining for every new failure mode.
Dev: So, we're talking about a system that’s lightweight enough not to slow down the loop rate, yet smart enough to leverage human intuition when the AI gets stuck in a tricky spatial or semantic mess? That sounds like exactly what we need for next-generation manipulation systems.
Carnegie Mellon University
cs.RO
Submitted: 2026-06-14
Updated: 2026-10-06
Comments: 9 pages, 8 figures, 4 tables
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 83/100
The gist: Recent advancements in Vision-Language-Action (VLA) models demonstrate impressive generalist capabilities in robot manipulation, yet these policies can be brittle under out-of-distribution spatial
Key concepts
- Shared Autonomy for Policy Steering (SAPS)
- A lightweight framework that mixes a robot's learned actions with direct human commands during operation. It allows a user to steer generalist robot policies in real-world situations without needing to retrain the underlying AI models or add complex auxiliary systems.
- Action-Level Blending
- The core mechanism where the final action is calculated by combining two inputs: the output from the VLA policy and a human expert command. This blending determines how much autonomy (policy control) versus human control is used at every step of movement.
- Dynamic Cosine-Similarity Strategy
- An arbitration method that adjusts the level of autonomy based on how closely the human's action matches the robot's policy action geometrically. It uses a logistic transformation to create a smooth transition between full human control and full policy control, ensuring continuous steering.
- Policy Steering
- The process of guiding a pre-trained robot policy toward better performance in specific scenarios by injecting external control signals. SAPS achieves this by blending expert input with the learned behavior, making the generalist model more reliable under unexpected conditions.
Terminology
Summary
Recent advancements in Vision-Language-Action (VLA) models demonstrate impressive generalist capabilities in robot manipulation, yet these policies can be brittle under out-of-distribution spatial and semantic perturbations. This work introduces Shared Autonomy for Policy Steering (SAPS), a model-agnostic framework that blends real-time human teleoperation commands with pretrained policy actions at the action level to reliably deploy generalist robot policies in real-world contexts.
The gist
SAPS is a lightweight shared-autonomy framework that steers pretrained robot policies and VLAs by blending policy actions with human teleoperation commands, without requiring retraining, auxiliary models, or sampler modification. Across LIBERO, LIBERO-PRO, CALVIN, and real Franka hardware evaluations, SAPS improves robustness over autonomous execution while reducing human intervention relative to pure teleoperation.
Framework Overview
SAPS is designed to combine the manipulation proficiency of foundation models with sparse corrective input from a user through action-level blending. Unlike prior steering methods that require auxiliary dynamics models or modifications to diffusion samplers, SAPS operates post-inference at the action level,
requiring no architectural modifications, auxiliary models, or policy retraining.
The core mechanism involves computing a final blended action:
a(1:6) blended = αa(1:6) VLA + (1−α)a(1:6) expert.
This blending determines the degree of autonomy based on an arbitration coefficient, where α ∈ [0, 1] determines the degree of autonomy.
The gripper state is handled separately to prevent interference between opposing commands.
Arbitration Strategies
The framework proposes three distinct arbitration strategies to balance human and policy control:
-
Fixed Blending: Uses a constant blending coefficient, such as
α = 0.5,
during active human intervention, transitioning to full autonomy when no input is detected (when the L2 norm of the expert action falls below a threshold ε=0.001). -
Full Takeover: Allows the human operator to completely override the VLA policy whenever active intervention is detected, switching control between
α = 0.0
andα = 1.0.
-
Dynamic Cosine-Similarity Strategy: This strategy scales autonomy based on the geometric agreement between human and policy actions. It computes directional agreement via cosine similarity, where the coefficient is determined by a logistic transformation:
γ=σ(kcos(θ))= 1/(1+e−kcos(θ)), k=6.
The arbitration coefficient is set as α = γ,
enabling smooth, continuous transitions between human and autonomous control.
Evaluation and Results
SAPS was evaluated across simulation environments (LIBERO, LIBERO-PRO, CALVIN) and on real-world robot hardware (Franka). The results consistently show that shared autonomy significantly improves performance over the autonomous execution baseline. For instance, in the controlled LIBERO perturbation study, Blending and Cosine arbitration achieved success rates of 87.9% and 90.0%, respectively, compared to π0.5's drop to 22.9% at a perturbation distance of 0.15m (Figure 3). In the more complex LIBERO-PRO evaluation, Cosine reached a mean success rate of 97.4% across the ten out-of-distribution tasks, significantly outperforming π0.5's 15.0% mean success rate and approaching pure Teleoperation at 98.8%. Furthermore, completion times were reduced; for example, in LIBERO-PRO, Cosine achieved a mean completion time of 11.1s compared to 30.7s for π0.5 and 46.0s for Teleoperation across all tasks (Figure 5).
Real-World Transferability
Hardware evaluations confirmed the transferability of SAPS from simulation to physical execution. Across the three real-world tasks (Pick and Place, Close Drawer, Open Cabinet), Cosine achieved an average success rate of 98.3%, while Blending achieved 93.3%. Crucially, both shared-autonomy methods substantially reduce human intervention relative to Teleoperation,
with Cosine and Blending reducing intervention by 30-50% relative to Teleoperation across all three hardware tasks
(Table A5). This demonstrates that SAPS enables the operator to provide partial corrective input while the policy continues to contribute learned manipulation behavior during physical robot execution. Both methods significantly reduce completion time relative to π0.5 across all hardware tasks (p≤0.0212).
Limitations
The paper notes that SAPS performance depends on the human operator's timing and skill, as it relies on "blending with the policy rather than explicit task progress, intent, uncertainty, contact, or safety estimates.
Improvements for AI systems
Based on the provided scientific paper, here are specific improvements that can be made to current Vision-Language-Action (VLA) models and how these improved systems would function:
-
Improve Robustness to Out-of-Distribution (OOD) Perturbations through Action-Level Shared Autonomy (SAPS):
-
Implement a Model-Agnostic Steering Mechanism:
-
Introduce Dynamic, Confidence-Based Control Arbitration:
-
Enable Real-World Deployment with Reduced Human Cognitive Load:
-
Improve Robustness to Out-of-Distribution (OOD) Perturbations through Action-Level Shared Autonomy (SAPS):
The improved system will employ the SAPS framework, which blends the pretrained VLA policy's actions with real-time human teleoperation commands at the action level.
It can now handle spatial and semantic perturbations (e.g., object shifts, incorrect object binding) that cause brittle policies to fail (as seen in LIBERO-PRO) by using sparse human input as an online corrective signal. The system will recover from OOD states without requiring policy retraining or auxiliary dynamics models, significantly increasing task success rates over purely autonomous execution (up to 82% improvement).
- Implement a Model-Agnostic Steering Mechanism:
The system will utilize SAPS, which steers pretrained policies post-inference at the action level. This means it requires no architectural modifications, policy retraining, or auxiliary models (unlike ITPS or DynaGuide methods). It can reliably guide generalist robot policies through OOD states by simply blending their proposed actions with human corrections.
- Introduce Dynamic, Confidence-Based Control Arbitration:
The system will use the proposed Cosine-similarity strategy for arbitration. This involves calculating the geometric agreement between the expert (human) action and the VLA policy action to dynamically scale autonomy based on this agreement using a sigmoid function:
α = σ(k cos(θ)) where θ is the angle between human and policy actions.
This allows for smooth, continuous transitions between full autonomy (when agreement is high) and immediate takeover (when disagreement is high), ensuring corrective intervention while preserving stable behavior.
- Enable Real-World Deployment with Reduced Human Cognitive Load:
The system can be deployed on real-world hardware (e.g., Franka arms) where the VLA policy handles nominal, low-level manipulation, and the human operator provides sparse, high-level corrective input only when necessary. This drastically reduces the required continuous manual control compared to pure teleoperation (reducing intervention by 30–50% across hardware tasks) while achieving faster task completion times than both autonomous execution and pure teleoperation.
In summary, the improved AI system is a more reliable generalist robot policy that maintains its learned manipulation skills while being guided by a human operator only when it encounters novel or out-of-distribution situations, leading to higher success rates and more efficient task completion in real-world assistive applications.
Sources
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- OpenVLA: An Open-Source Vision-Language-Action Model
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation
- A Systematic Study of Data Modalities and Strategies for Co-training Large Behavior Models for Robot Manipulation
- LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization
- Inference-Time Policy Steering through Human Interactions
- VLS: Steering Pretrained Robot Policies via Vision-Language Models
- Do What You Say: Steering Vision-Language-Action Models via Runtime Reasoning-Action Alignment Verification
- Independence in the Home: A Wearable Interface for a Person with Quadriplegia to Teleoperate a Mobile Manipulator
- Bimanual High-Density EMG Control for In-Home Mobile Manipulation by Users with Quadriplegia
- RaC: Robot Learning for Long-Horizon Tasks by Scaling Recovery and Correction
- LAMS: LLM-Driven Automatic Mode Switching for Assistive Teleoperation
- Incremental Learning for Robot Shared Autonomy
- ${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities
- Efficient and Reliable Teleoperation through Real-to-Sim-to-Real Shared Autonomy
- RLinf: Flexible and Efficient Large-scale Reinforcement Learning via Macro-to-Micro Flow Transformation
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving