PACT-WAM: Predicting Actions and Visual Foresight with Compact Temporal Encoding for Robot Manipulation
cs.RO, cs.AI
Submitted: 2026-02-05
Updated: 2026-09-16
License: http://creativecommons.org/licenses/by/4.0/
The gist: Robot manipulation uses temporal context to select actions and visual foresight to assess their consequences, yet dense representations of past and future observations incur substantial processing
Terminology
Abstract
Robot manipulation uses temporal context to select actions and visual foresight to assess their consequences, yet dense representations of past and future observations incur substantial processing costs. We introduce PACT-WAM, a world-action model that jointly generates a 16-step action trajectory and its temporally corresponding visual forecast through conditional flow sampling. Hierarchical history encoding assigns coarse spatial representations to earlier observations and finer representations to recent ones, retaining 16 observations with 256 tokens per view, 75% fewer than dense encoding of the same frames. A shared flow module jointly updates continuous action and visual states through two modality-specific heads under transition-wise causal attention, and a TiTok-VAE decoder reconstructs multi-view future images from the visual latents. Decoded forecasts also support Proposal Review (PR), a vision-language model component for execution-prefix selection and proposal rejection. Without PR, PACT-WAM achieves average success rates of 98.6%, 92.3%, and 78.0% on LIBERO, RoboTwin 2.0, and real-world Piper tasks, respectively. PR provides a test-time enhancement, raising these rates to 99.5%, 93.4%, and 86.7%. Ablations show that hierarchical history allocation and joint action-visual generation improve control success, while analyses of visual capacity and forecast-guided execution characterize the trade-offs between success and proposal-generation cost.
Sources
- Cosmos World Foundation Model Platform for Physical AI
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- WorldVLA: Towards Autoregressive Action World Model
- WorldPrediction: A Benchmark for High-level World Modeling and Long-horizon Procedural Planning
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
- RoboHorizon: An LLM-Assisted Multi-View World Model for Long-Horizon Robotic Manipulation
- Ctrl-World: A Controllable Generative World Model for Robot Manipulation
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- Action-aware Dynamic Pruning for Efficient Vision-Language-Action Manipulation
- FAST: Efficient Action Tokenization for Vision-Language-Action Models
- World4RL: Diffusion World Models for Policy Refinement with Reinforcement Learning for Robotic Manipulation
- Long-Context State-Space Video World Models
- OpenVLA: An Open-Source Vision-Language-Action Model
- Inference-Time Enhancement of Generative Robot Policies via Predictive World Modeling
- Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
- AstraNav-Memory: Contexts Compression for Long Memory
- AVID: Adapting Video Diffusion Models to World Models
- Unified Video Action Model
- Leave No Observation Behind: Real-time Correction for VLA Action Chunks
- DINOv3
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving