TEMPO: Learning Temporal Context for Dynamic Robot Manipulation
cs.RO, cs.CV, cs.LG
Submitted: 2026-09-15
Updated: 2026-09-15
Comments: Accepted at CoRL 2026. Project page: https://tempo-robot.github.io/
Project page: https://tempo-robot.github.io
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Vision-language-action (VLA) models have achieved impressive performance in quasi-static manipulation, but struggle in dynamic manipulation tasks because they operate on a single observation at
Terminology
Abstract
Vision-language-action (VLA) models have achieved impressive performance in quasi-static manipulation, but struggle in dynamic manipulation tasks because they operate on a single observation at inference time. We identify two representational failures that underlie this limitation. The first is motion ambiguity, where a single observation does not include scene dynamics and therefore cannot anticipate the future state of moving objects. The second is state aliasing, where visually similar observations from different points in a task require different actions. We argue that these failures persist regardless of model scale and inference latency, showing that the bottleneck is missing temporal context rather than model capacity. Based on this insight, we propose TEMPO, which augments a pretrained VLA with two temporal inputs: a motion summary extracted from a frozen video foundation model to resolve motion ambiguity and a compact proprioceptive history to resolve state aliasing. TEMPO requires no modification to the backbone and adds minimal compute overhead at training or deployment. Across four dynamic manipulation tasks, it improves Bottle Handover success from 44% to 74% and is the only method that solves state aliasing. Probing and ablation studies confirm that each temporal signal independently addresses its corresponding failure. We further release TEMPO-Bench, a benchmark of over 50k annotated frames for evaluating motion-aware robot perception in both regression and multiple-choice formats. Project Website: https://tempo-robot.github.io/
Sources
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- VLASH: Real-Time VLAs via Future-State-Aware Asynchronous Inference
- SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics
- Leave No Observation Behind: Real-time Correction for VLA Action Chunks
- DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation
- F2F-AP: Flow-to-Future Asynchronous Policy for Real-time Dynamic Manipulation
- Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving