Keep the Future, Drop the Rollout: RIFT for World Action Models
Australian National University · Beijing Normal University
cs.RO, cs.AI
Submitted: 2026-08-12
Updated: 2026-09-25
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 100/100
The gist: World action models (WAMs) condition robot actions on predicted futures, but iterative video rollout increases deployment latency.
Terminology
Summary
World action models (WAMs) condition robot actions on predicted futures, but iterative video rollout increases deployment latency. This paper asks whether action generation requires the evolving rollout trajectory or only its future representation. Across four WAMs on all 40 LIBERO tasks, paired closed-loop interventions show that masking or reassigning future-cache values changes execution and reduces success, indicating sensitivity to future values and their assigned positions. For Joint and Cosmos-2, however, replaying one fixed final-clean key/value (K/V) cache nearly preserves unmodified execution, with 1.7 to 1.9 cm end-effector average displacement error and 97.9% to 98.2% success. This separates cache consumption from production: these models can reuse a fixed cache but still require iterative rollout to construct it. The paper therefore proposes Rift (Rollout-free Imagination via Future Tokens), which uses learned anticipation tokens to construct a complete future K/V cache in one backbone pass while retaining the original future-read interface. On LIBERO, Rift achieves 98.8% success, close to rollout-based Joint, IDM, and LingBot-VA at 98.4% to 98.6%, while reducing action-chunk latency by 68.2% to 89.1%. On RoboTwin 2.0, Rift reaches 92.9/92.6% on clean/randomized scenes, the highest observed among the evaluated methods. These results support rollout-free future conditioning without iterative video generation at deployment.
The paper makes three contributions. First, it introduces a paired closed-loop intervention protocol for future caches. The protocol records the action-independent cache, then either masks the future read or edits its values under the recorded keys before measuring the resulting physical execution with EE-ADE and SR. Second, it identifies two properties of the future cache. Actions are highly sensitive to removing the future read or reassigning values across space and time, while final clean values under the original keys nearly preserve execution. Third, it designs and validates a rollout-free future interface. Rift produces a complete future-position K/V cache in one backbone pass, matches the rollout-based success tier at 1.1× current-only latency, and reaches the highest observed LIBERO and RoboTwin 2.0 success.
The intervention study uses 2,000 paired trials on all 40 LIBERO tasks. Masking the future read yields 18.7 cm EE-ADE and reduces success from 98.4% to 9.7%. Spatial permutation and temporal swapping yield similar EE-ADE (14.3 and 15.6 cm) but sharply different success (65.2% and 0.7%), showing that the action reads future content at its assigned positions. In contrast, replaying final clean values under the original keys yields only 1.9 cm EE-ADE and 97.9% success. Under the original keys, the action therefore depends strongly on future value content and its organization, but little on how the future values evolve across denoising steps.
Rift keeps Fast-WAM-Joint’s architecture and future-read interface. For a video stack with hidden width d and L layers, it replaces rolled-out future tokens with learned anticipation tokens E ∈ Rm×d. Each token inherits its corresponding future spatiotemporal index. If each latent frame contains n tokens and the clip contains Tlat latent frames, full alignment uses m = n(Tlat −1). The paper uses full alignment with m = 196 on LIBERO and m = 240 on RoboTwin 2.0. One video-stack prefill writes the full future-position cache: Cϕ(o, l) = CachePrefillϕ([f0(o); E], o, l), where f0(o) denotes the first-frame tokens extracted from observation o, and ϕ collects the shared video expert’s parameters and the learned tokens E. The action expert reads Cϕ through the rollout model’s per-layer future-position interface: â1:H = ActionDenoise(o, l; Cϕ(o, l)). The action-chunk horizon is H = 32. Cϕ has no action-flow index: the same K/V serve every denoising evaluation, matching the fixed-cache consumption pattern tested in the intervention study.
Training uses two forwards through the same video expert per optimization step. Deployment retains only the second forward’s cache prefill. The first forward applies native video-flow loss Lvid to a clean-first-frame clip with noised future latents and no anticipation tokens or action supervision. The second forward matches deployment through input [f0; E] and its attention mask. Clean rows train the action expert on this cache with Fast-WAM’s inherited action flow-matching loss Lact. After 70% of the configured curriculum horizon, the probability and scale of first-frame latent noise rise linearly from zero to 0.3 and 0.06 times the latent standard deviation. Perturbed rows retain LFM and Lprobe but are masked out before Lact is reduced.
The paper uses conditional flow matching as a distributional auxiliary objective. It samples ϵ ∼ N(0, I) and σ ∈ [0, 1] with the native video-flow schedule, then defines Xσ = (1 − σ)Y + σϵ and vσ⋆ = ϵ − Y. Conditioned on Sϕ, the training-only FM head vψ predicts the velocity from (Xσ, σ). Using the native video-flow timestep weight wvid(σ), its loss is LFM = EY,ϵ,σ[wvid(σ)∥vψ(Xσ, σ; Sϕ) − vσ⋆∥2]. This auxiliary head shapes the cache producer during training but enters neither action nor policy-only deployment. The paper also retains a stopped-gradient linear probe, ŶL2 = gω(RMS(stopgrad(Sϕ))), trained with MSE, Lprobe = mean((ŶL2 − Y)2). Detached input confines this loss to the probe.
Both forwards share the video expert. Lvid updates this expert and its video head; Lact updates the action expert and backpropagates through the cache into the shared video expert and E; LFM updates the shared video expert, E, and the FM head; and detached Lprobe updates only the probe. The total loss is L = Lvid + Lact + λFM LFM + λprobe Lprobe. Both auxiliary weights stay at 1 for the first 70% of the curriculum horizon, then follow a cosine decay to 0.2 over the final 30%. Policy-only deployment retains one [f0; E] prefill, fixed Cϕ, and the standard action flow.
On LIBERO, Rift achieves 98.8% overall success, close to the 98.4% to 98.6% achieved by rollout-based Joint, IDM, and LingBot-VA. Unlike these methods, Rift requires only 247.9 ms per action chunk, reducing latency by 68.2% to 89.1% while remaining close to current-only Fast-WAM at 235.7 ms. Compared with rollout-free Fast-WAM and PFD, Rift improves success by 2.0 and 1.5 percentage points, respectively, at comparable latency. On RoboTwin 2.0, Rift reaches 92.9/92.6 on clean/randomized scenes, the best observed among the evaluated methods, against 92.5/92.1 for PFD, 92.4/91.4 for rollout-based LingBot-VA, 91.9/91.6 for Fast-WAM, and 91.0/91.1 for rollout-based Fast-WAM-Joint.
Ablations show that the base recipe Rift-L2, which regresses future latents with a direct L2 loss, reaches 98.37 ±0.12; the conditional-FM recipe reaches 98.8 ±0.17. Both use the same one-pass graph and 247.9 ms cost, isolating supervision without deployment overhead. The number of anticipation tokens is also swept from m = 2 to full alignment (m = 196); current-only Fast-WAM is the no-cache m = 0 reference (96.75%). Rift-L2 rises from 97.08% to 98.37%; conditional FM exceeds it from m = 4 and peaks at 98.78%. Even small interfaces beat the reference, and full alignment is best for both.
Qualitative results compare Fast-WAM-Joint rollout and Rift one-pass decodes from matched starts on both benchmarks. Both evolve similarly at frames 0, 4, and 8. These visuals diagnose future representations, not cache equivalence. The paper also presents an optional L2–FM uncertainty warning. The stopped-gradient L2 probe and conditional-FM head yield controller-independent future estimates whose normalized disagreement defines a CUSUM warning. Calibrated on 1,967 successful episodes, the mean CUSUM over 33 failed rollouts crosses the alarm threshold η 210 steps before the common t = 420 endpoint.
The paper concludes that world action models combine a representation read by the action expert with the iterative rollout that produces it. Intervening on the future K/V interface shows that actions require values bound to token positions, while one fixed final-clean K/V cache nearly reproduces Original execution within 1.9 cm EE-ADE. Because this cache remains rollout-produced, the intervention establishes consumption-side sufficiency rather than rollout-free production. Rift addresses the remaining production problem with one anticipation-token prefill, preserving the complete future K/V interface while removing video denoising and VAE decoding. It clears the current-only gap on LIBERO at 1.1× baseline latency and reaches the best observed RoboTwin 2.0 success in both evaluation settings.
Improvements for AI systems
Improvements to AI systems:
-
Latency reduction via learned future-cache prefill: Replace iterative video rollout with a single-pass cache construction using learned anticipation tokens, cutting action-chunk latency by 68–89% while maintaining success rates (98.8% on LIBERO, 92.9% on RoboTwin 2.0). The improved system can generate robot actions at near-real-time speeds without sacrificing task performance.
-
Decoupled future representation from rollout dynamics: Use a fixed, final-clean key/value cache for action decoding instead of evolving rollout states. This enables systems to reuse a precomputed future representation across multiple action chunks, reducing computational overhead during deployment.
-
Position-aware future conditioning: Enforce strict spatiotemporal alignment of future tokens (e.g., 196 tokens for full alignment) so actions read values at assigned positions. This improves robustness against spatial permutation errors (success drops from 98.4% to 65.2% if misaligned) and temporal swaps (0.7% success), ensuring reliable action generation under varied future contexts.
-
Conditional flow matching as auxiliary supervision: Train a distributional objective (LFM) alongside the main action loss to shape the cache producer, improving success by 0.43 percentage points over direct L2 regression (98.8% vs 98.37%) without extra deployment cost. The system can better handle stochastic future states.
-
Early failure detection via uncertainty warning: Use disagreement between a stopped-gradient L2 probe and conditional-FM head to generate a CUSUM-based alarm. The improved system can predict task failures 210 steps before the endpoint, enabling preemptive intervention or replanning in long-horizon tasks.
-
Scalable token-efficiency trade-off: With as few as 4 anticipation tokens, the system outperforms current-only baselines (97.08% vs 96.75% success), allowing deployment on memory-constrained hardware while still benefiting from future conditioning.
-
Curriculum-based noise injection for robustness: Gradually increase first-frame latent noise (from 0 to 0.3 probability/0.06 scale) during training, improving generalization to noisy observations in randomized scenes (92.6% success on RoboTwin 2.0 randomized).
Abstract
World action models (WAMs) condition robot actions on predicted futures, but iterative video rollout increases deployment latency. We ask whether action generation requires the evolving rollout trajectory or only its future representation. Across four WAMs on all 40 LIBERO tasks, paired closed-loop interventions show that masking or reassigning future-cache values changes execution and reduces success, indicating sensitivity to future values and their assigned positions. For Joint and Cosmos-2, however, replaying one fixed final-clean key/value (K/V) cache nearly preserves unmodified execution, with 1.7 to 1.9 cm end-effector average displacement error and 97.9% to 98.2% success. This separates cache consumption from production: these models can reuse a fixed cache but still require iterative rollout to construct it. We therefore propose RIFT (Rollout-free Imagination via Future Tokens), which uses learned anticipation tokens to construct a complete future K/V cache in one backbone pass while retaining the original future-read interface. On LIBERO, RIFT achieves 98.8% success, close to rollout-based Joint, IDM, and LingBot-VA at 98.4% to 98.6%, while reducing action-chunk latency by 68.2% to 89.1%. On RoboTwin 2.0, RIFT reaches 92.9/92.6% on clean/randomized scenes, the highest observed among the evaluated methods. These results support rollout-free future conditioning without iterative video generation at deployment.
Sources
- Understanding intermediate layers using linear classifier probes
- Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation
- Motus: A Unified Latent Action World Model
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- RynnVLA-002: A Unified Vision-Language-Action and World Model
- GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
- Privileged Foresight Distillation: Zero-Cost Future Correction for World Action Models
- Vidar: Embodied Video Diffusion Model for Generalist Manipulation
- Gemini Robotics: Bringing AI into the Physical World
- Mastering Diverse Domains through World Models
- DreamGen: Unlocking Generalization in Robot Learning through Video World Models
- Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning
- Causal World Modeling for Robot Control
- Unified Video Action Model
- Video Generators are Robot Policies
- Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation
- RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation
- Being-H0.7: A Latent World-Action Model from Egocentric Videos
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving