What Makes World Action Models Generalize? An Empirical Study of Test-Time Future Modeling
cs.CV, cs.RO
Submitted: 2026-09-28
Updated: 2026-09-29
Project page: https://zrporz.github.io/Simple-WAM-Web
Terminology
Sources
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
- Benchmarking Vision, Language, & Action Models in Procedurally Generated, Open Ended Action Environments
- Efficient-WAM: A 1B-Parameter World-Action Model with Low-Cost Future Imagination
- Causal World Modeling for Robot Control
- Light-WAM: Efficient World Action Models with State-Fusion Action Decoding
- Being-H0.7: A Latent World-Action Model from Egocentric Videos
- mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs
- MemoryVLA++: Temporal Modeling via Memory and Imagination in Vision-Language-Action Models
- Wan: Open and Advanced Large-Scale Video Generative Models
- VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model
- Fast-WAM: Do World Action Models Need Test-time Future Imagination?
- Keep the Future, Drop the Rollout: RIFT for World Action Models
- ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?
- Faster-WAM: Efficient Inference-Time Future Conditioning for Robust World Action Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models