Foresight Without Seeing: Latent Futures for World Action Models

arXiv:2608.11605 · cs.AI · Submitted 2026-08-12 · Read on arXiv

Jiakai Huang, Zhongbo Wu, Zheng Zhang, Zihan Wang, Shan You, Tao Huang

Shanghai Jiao Tong University · ACE Robotics · Nanyang Technological University

cs.AI

Submitted: 2026-08-12

Updated: 2026-08-13

Comments: 12 pages, 3 figures

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 95/100

The gist: ForeWAM is a dynamics-conditioned direct-policy World Action Model (WAM) that provides predictive context for action generation without explicitly generating future videos.

Terminology

Summary

ForeWAM is a dynamics-conditioned direct-policy World Action Model (WAM) that provides predictive context for action generation without explicitly generating future videos. It addresses the trade-off between predictive context and inference efficiency in WAMs: explicit-future WAMs (cascaded and joint) expose predicted scene evolution to the action pathway but incur substantial inference costs from iterative video denoising, while direct-policy WAMs (e.g., Fast-WAM) skip future generation for efficiency but lack an explicit inference-time interface for exposing predictive dynamics to the Action DiT.

The core innovation is Future-KV, an implicit interface that transfers predictive context from a Video DiT to an Action DiT. Future-KV preserves the clean visual latent of the current observation, initializes unobserved future slots with noise, and processes them through a single Video DiT prefill. The resulting layer-wise key-value states are cached and reused throughout action denoising, allowing the action pathway to access predictive context over both the current observation and latent future slots without iteratively generating or decoding future video.

To encourage these implicit future states to focus on interaction-induced scene transitions (object motion, contact changes, task progress), the paper introduces dynamics registers supervised by a frozen LaWM latent-action teacher. During training, the teacher extracts compact, non-executable latent-action representations from pairs of real visual observations before and after a transition. These representations supervise the dynamics registers to encode state-transition information. Ground-truth future observations and the teacher are used only during training; deployment requires neither future observations nor the teacher and performs no future video generation.

The model architecture uses four token groups: current-frame tokens C, dynamics registers D, future-slot tokens F, and action tokens A. A structured attention mask routes information such that future-slot tokens can integrate the current frame and dynamics registers, and action tokens can read the complete video sequence together with the registers. The training objective combines three loss terms: a video flow-matching loss on demonstrated future latents, an action flow-matching loss on executable action chunks, and a latent-action distillation loss that encourages the mean-pooled dynamics registers to match the teacher's latent-action target.

Results show that without embodied robot data pretraining, ForeWAM achieves 96.7% average success on the standard LIBERO suites (97.0% Spatial, 99.6% Object, 97.2% Goal, 92.8% Long), and ForeWAM-Flash (an accelerated variant using OneDP distillation to reduce denoising steps from 10 to 2) achieves 96.9% overall. On the observed LIBERO-Plus subset, ForeWAM reaches 61.6% success, surpassing the reported Fast-WAM result of 51.5% by 10.1 percentage points, with the largest gains on camera viewpoint (+46.1 points) and sensor noise (+21.1 points). ForeWAM-Flash reaches 58.2%, 6.7 points above Fast-WAM. ForeWAM reduces mean action-generation latency from 667 ms (Fast-WAM) to 568 ms (a 14.8% reduction), while ForeWAM-Flash lowers it to 220 ms (a 67.0% reduction). ForeWAM uses approximately one-third of the policy parameters of Fast-WAM (2B versus 6B).

Ablation studies on LIBERO-Plus show that combining Future-KV and LA supervision (61.6%) outperforms either alone (Future-KV only: 58.5%; LA supervision only: 58.0%), with a base policy using neither reaching 53.6%. The paper notes limitations: evaluation is limited to LIBERO and LIBERO-Plus, and generalization to different embodiments, task distributions, or real-world deployment remains unclear.

Improvements for AI systems

Improvements to AI Systems:

  1. Add an implicit predictive-context interface (Future-KV) to action-generation models. Instead of generating future frames or relying solely on current observations, the system caches layer-wise key-value states from a single video prefill over latent future slots. This gives the action policy access to predicted scene dynamics at inference time with no iterative video decoding.

  2. Introduce trainable dynamics registers supervised by a frozen latent-action teacher. These registers are mean-pooled and distilled to match compact, non-executable latent-action representations extracted from real pre/post-transition observation pairs. This forces the model to encode interaction-induced state changes (object motion, contact changes, task progress) without requiring future observations or the teacher at deployment.

  3. Use a structured attention mask with four token groups (current-frame, dynamics registers, future-slots, action tokens). This routes information so future slots integrate current frame and dynamics registers, while action tokens read the full sequence plus registers—enabling joint reasoning over observed and predicted states in a single pass.

  4. Combine three loss terms: video flow-matching, action flow-matching, and latent-action distillation. This multi-objective training jointly optimizes future-state prediction, executable action generation, and semantic alignment of dynamics registers, improving generalization under distribution shifts.

  5. Apply OneDP distillation to reduce denoising steps from 10 to 2 for the action pathway. This yields a fast variant (ForeWAM-Flash) that cuts action-generation latency by 67% (from 667 ms to 220 ms) while maintaining high success, enabling real-time or near-real-time control.

What the Improved AI System Can Do:

  • Generate actions for robotic manipulation with predictive awareness of future scene changes (e.g., anticipating object movement, contact events, task progress) without the computational cost of video generation.

  • Achieve high task success on standard benchmarks (e.g., 96.7% average on LIBERO suites) and outperform prior direct-policy models (e.g., +10.1 points on LIBERO-Plus over Fast-WAM), especially under camera viewpoint changes and sensor noise.

  • Run at lower latency (568 ms for standard, 220 ms for flash variant) with one-third the policy parameters of comparable models (2B vs. 6B), making it feasible for edge or embedded deployment.

  • Generalize to unseen task variations by leveraging implicit future-state representations that capture interaction dynamics, without needing future observations or a teacher model at inference time.

  • Support both high-accuracy and high-speed modes via the same architecture, allowing a trade-off between performance and computational budget depending on the application (e.g., simulation vs. real-time robot control).

Sources

Related papers