SpeedTuning: Speeding Up Policy Execution with Lightweight Reinforcement Learning
David D. Yuan, Tony Z. Zhao, Kaylee Burns, Chelsea Finn
Stanford University
cs.RO, cs.AI
Submitted: 2026-08-11
Updated: 2026-08-12
Comments: 10 pages, 12 figures. This arXiv version includes an appendix with qualitative simulation rollouts and additional ablations. Published at ICRA 2025
Journal ref: 2025 IEEE International Conference on Robotics and Automation (ICRA), pp. 1184-1192, 2025
DOI: 10.1109/ICRA55743.2025.11128753
Project page: https://daivdyuan.github.io/speed-tuning
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 95/100
The gist: SPEED TUNING: Speeding Up Policy Execution with Lightweight Reinforcement Learning Abstract Summary: The paper introduces SPEED TUNING, a reinforcement learning framework designed to enhance the
Terminology
Summary
SPEED TUNING: Speeding Up Policy Execution with Lightweight Reinforcement Learning
Abstract Summary:
The paper introduces SPEED TUNING, a reinforcement learning framework designed to enhance the speed of manipulation policies. The authors note that learned robotic policies hold promise for advancing generalizable manipulation, [but] their practical deployment is often hindered by suboptimal execution speed.
They explain that imitation learning policies are inherently limited by hardware constraints and the speed of the operator during data collection,
and there are no established methods for accelerating policies learned via imitation, and the empirical relationship between execution speed and task success remains underexplored.
SPEED TUNING learns to predict the optimal execution speed for actions, thereby complementing a base policy without necessitating additional data collection.
The paper provides "empirical evidence that SPEED TUNING achieves substantial improvements in execution speed, exceeding 2.4x speed-up, while preserving an adequate success rate compared to both the original task policy and straightforward speed-up methods such as linear interpolation at a fixed speed. The approach is evaluated
across a diverse set of dynamic and precise tasks, including pouring, throwing, and picking, demonstrating its effectiveness and robustness in enhancing real-world robotic manipulation."
Introduction Summary:
The paper states that Speed is critical in real-world robotic tasks, as faster policies can increase throughput and enhance the user experience.
However, imitation learning methods largely ignore speed both when training and evaluating a policy.
The authors note that Training robots to perform tasks quickly and successfully through imitation is challenging
because imitation learning is typically concerned with matching the behavior in the demonstration dataset, while it is hard to collect fast and accurate teleoperation data even from experts.
They give examples using ALOHA hardware: cutting a piece of tape and sticking it to a box takes more than 15 seconds, and putting a velcro shoe on a foot takes more than 22 seconds,
while humans can perform either of these tasks in under 5 seconds.
They argue that naive approaches—such as increasing the underlying speed of the robot by a constant multiplier—are insufficient, as they fail to account for the nuanced relationship between speed and task success.
The paper presents SPEED TUNING as a method for accelerating manipulation policies by learning a 'speed-policy' that predicts a speed factor for a chunk of future actions.
The speed policy takes in the observation history and outputs the velocity of the subsequent action
and is trained with reinforcement learning and optimizes a task and speed reward.
The method is designed to be practical and compatible with modern imitation learning methods, most of which predict action chunks at:t+k,
by linearly interpolating predicted chunks by the factor predicted by the speed policy.
The authors emphasize that because the speed policy is trained entirely with reinforcement learning, this acceleration process requires no extra data-collection.
Related Work Summary:
The paper discusses three areas of related work. First, imitation learning, noting that the quality of learned skills remains inherently tied to the quality of demonstrations, which can be limited by unintuitive teleoperation interfaces and demonstrator sub-optimality.
Second, reinforcement learning in robotics, noting that state-of-the art RL algorithms still require extensive real-world interactions to learn complex control policies from scratch,
but that roboticists have successfully fine-tuned policies pre-trained with imitation learning.
Third, fast robot manipulation, which has traditionally centered around optimizing the robot's hardware or integrating speed into the reward function.
The authors position SPEED TUNING as unique because it is compatible with many base policy learning methods, which means that it can be incorporated on top of already performant robotic policies.
Method Summary:
The paper introduces SPEED TUNING as a method for accelerating manipulation policies learned from demonstrations while optimizing for both execution speed and task success.
The approach "decoupl[es] the policy into two components: a task policy trained with imitation learning to mimic expert demonstrations... and a speed policy trained with reinforcement learning to predict optimal action velocities that balance speed and success."
The objective function is defined as: J(ψ) = Eτ∼πψ [Σ(α · rspeed(vt) + rtask(st, at))]
where the reward at each time step t consists of a weighted speed reward α · rspeed(vt) and the task reward rtask(st, at).
The velocities are restricted to a discrete set: V = v1, v2,..., vK to simplify policy optimization.
The paper describes the factorization: We decompose the execution policy into two distinct sub-policies: a task policy, πθ(atst), and a speed policy, πφ(vtst).
The task policy is trained to regress actions with imitation learning.
For the speed policy, the paper explores both constant and non-linear speed policies.
For the speed-adaptive policy, the paper states: "SPEED TUNING seeks to balance swift task execution with maintaining a high success rate. Specifically, we use reinforcement learning with the objective in Equation 3 to train a speed policy πφ(vtst), which predicts an optimal speed vt for a fixed task policy πθ. The optimization uses
Rainbow DQN... because of its strong performance and high sample efficiency."
Key design choices include: limiting the output speed v to a discrete set of possible values,
introducing an additional hyperparameter, β, to further scale the speed reward: rspeed(v) = vβ,
using a frame skip constant kskip ∈ Z+, querying the speed policy only once every kskip steps,
adopting a frame stack of kstack recent observations as input,
and predicting an action chunk at each time step
to make the implementation of the velocity change more straightforward and reduc[e] the effective horizon of the task.
For interpolation over action chunks, the paper defines linear temporal interpolation over a function f: Z 7→ Rn with a step length of v at point t
as: interpf,v(t) = f(⌊vt⌋) + (vt − ⌊vt⌋)/v · (f(⌊vt + 1⌋) − f(⌊vt⌋)).
This effectively accelerates the execution of the function f by a factor of v.
Experiments Summary:
The paper evaluates SPEED TUNING across six fine-grained manipulation tasks, divided equally between simulated and real-world environments.
The tasks include:
-
Simulated Cube Transfer:
requires the robot to pick up and transfer a cube between two grippers
-
Simulated Peg Insertion:
the robot must pick up and accurately insert a peg into a cube with a hole, demanding precise and bimanual manipulation
-
Simulated Tea Bag Transfer:
involves picking up and transferring a tea bag from a desktop into a cup
withthe swinging motion of the tea bag during transfer
presenting challenges -
Tea Bag Disposal (real-world):
the robot must first pick up a mug containing a tea bag and then carefully dispose of the tea bag into a garbage bin
-
Food Preparation (real-world): "involves a sequence of precise maneuvers: picking up a plate with one gripper, picking up and throwing two pieces of chicken from the desktop onto the plate, and finally placing the served food at a designated location on the desktop"
-
Pouring Almond (real-world):
involves opening the lid of a container and pouring almonds from a can into the container
The key results show:
-
SPEED TUNING achieves over a 2.4× speed-up across all six tasks compared to the original ACT task policy.
The authors notea trend where higher speed-ups are attained in tasks requiring less precision.
-
SPEED TUNING effectively retains the success rate at high speed,
andemerges as an outlying point beyond the Pareto curve defined by the baseline
of universal interpolation. -
SPEED TUNING learns the critical parts of the tasks to act accordingly.
For example, in Almond Pouring, "SPEED TUNING learns to maintain a high speed during the lid-opening and can-picking phases, deliberately slows down during the transfer and pouring stages to prevent spilling, and then resumes high speed to place the can back."
Ablations Summary:
The paper conducts ablation studies on Simulated Tea Bag Transfer:
-
Choice of Task Policy:
Replacing the learned ACT task policy in SPEED TUNING with an open-loop, scripted policy eliminates the learned generalization capabilities of the task policy.
The scripted policyresults in lower acceleration and a slightly higher success rate,
underscoringthe importance of generalizing across various speeds for the underlying task policy.
-
Image Observations:
Omitting image observations led to less stable training and lower success rates and acceleration. These results highlight the necessity of incorporating visual information for more efficient task execution.
-
Degree of Speed Rewards: "When β is large, the speed policy prioritizes velocity over task completion, resulting in lower success rates. Conversely, when β = 0, there is no direct incentive for speed, and acceleration relies solely on the standard discount factor γ, leading to unstable training."
-
Frame Skip:
Smaller frame skips with longer horizons lead to suboptimal speed policies, while larger skips sacrifice the granularity needed for achieving higher acceleration.
The selected value of 10balances the two objectives.
Conclusion Summary:
The paper concludes: "We introduced SPEED TUNING, a reinforcement learning framework designed to optimize both execution speed and task success for learned manipulation policies. By training a speed policy to dynamically adjust action velocities based on the current state, SPEED TUNING achieved over 2.4x speed-ups across various dynamic and precise tasks while maintaining adequate success rates compared to baseline methods. The authors state that
Future work should focus on enhancing sample efficiency to make the method more practical for real-world applications."
Improvements for AI systems
Improvements to AI Systems Based on SPEED TUNING:
- Dynamic Speed Adaptation Module for Imitation Learning Policies
-
Add a lightweight, RL-trained
speed policy
layer on top of any existing action-chunking policy (e.g., ACT, Diffusion Policy). This layer predicts a discrete speed multiplier per state, enabling the base policy to execute faster without retraining or new demonstrations. -
Capability: Robots can automatically accelerate slow, human-demonstrated tasks (e.g., assembly, cooking) by 2.4x+ while preserving success rates, adapting speed in real-time to task phases (fast during simple motions, slow during precise or fragile steps).
- Speed-Success Pareto Optimization for Manipulation
-
Integrate a dual-objective reward function (task success + speed) into the RL fine-tuning loop, with a tunable weight (α) and speed reward exponent (β) to control the trade-off. This allows AI systems to explicitly optimize for throughput and reliability simultaneously.
-
Capability: Systems can automatically find the optimal speed profile for new tasks, avoiding naive constant-speed scaling that fails on precision-critical actions (e.g., pouring without spilling, peg insertion).
- State-Conditioned Speed Scheduling via Discrete Action Space
-
Use a discrete set of speed factors (e.g., 0.5x, 1x, 2x, 4x) with frame skipping to simplify RL training and reduce computational overhead. The speed policy queries only every k steps, making it compatible with real-time control loops.
-
Capability: Real-world robots can adjust execution speed on-the-fly based on visual and proprioceptive feedback, mimicking human-like
slow down when uncertain, speed up when confident
behavior.
- Generalizable Acceleration Across Tasks and Policies
-
Design the speed policy as a plug-and-play module that works with any imitation-learned base policy (neural or scripted), as long as the base policy generalizes across speeds. This decoupling enables acceleration without modifying the original policy architecture.
-
Capability: A single speed-tuning framework can be applied to diverse tasks (picking, throwing, pouring, assembly) without task-specific engineering, reducing deployment time for new robotic applications.
- Visual-Input-Driven Speed Control
-
Require the speed policy to take image observations (not just state vectors) as input, enabling it to detect task-relevant visual cues (e.g., object proximity, container edges) to modulate speed.
-
Capability: Robots can slow down when approaching fragile objects or precise insertion points, and speed up during repetitive or low-risk segments, improving both safety and efficiency in unstructured environments.
- Sample-Efficient RL Fine-Tuning for Real-World Deployment
-
Use Rainbow DQN with frame stacking and action chunking to train the speed policy with minimal real-world interactions, reducing the need for extensive RL data collection.
-
Capability: AI systems can be fine-tuned for speed on physical robots within a few hours, making the approach practical for manufacturing, logistics, and home-assistance robots where data collection is expensive.
- Pareto-Optimal Speed Profiling for Task Families
-
Automatically generate a Pareto frontier of speed vs. success rate for a given task, allowing users to select a desired operating point (e.g., max speed with ≥90% success).
-
Capability: System operators can trade off throughput and reliability based on application requirements (e.g., high-speed for low-risk tasks, conservative speed for safety-critical ones).
- Phase-Aware Execution via Learned Speed Policies
-
The speed policy implicitly learns task phases (e.g.,
open lid
vs.pour liquid
) and adjusts speed accordingly, without explicit phase segmentation. -
Capability: Robots exhibit human-like tempo variations—fast during gross motions, slow during fine manipulation—improving both task success and perceived intelligence in human-robot interaction.
Sources
- Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware
- Diffusion Policy: Visuomotor Policy Learning via Action Diffusion
- Auto-Encoding Variational Bayes
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving