MT-WAM: Reorienting the One-Pass Predictive Representation Toward Action Generation
cs.CV, cs.RO
Submitted: 2026-09-18
Updated: 2026-09-18
Code: https://github.com/Alexi1984/MT-WAM
Terminology
Sources
- Fast-WAM: Do World Action Models Need Test-time Future Imagination?
- Reconstruction or Semantics? What Makes a Latent Space Useful for Robotic World Models
- How Should World Models Be Evaluated for Embodied Decision-Making? A Decision-Making-Centric Position
- ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?
- FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models
- Track4Action: Distilling World-Centric 3D Tracker into Vision-Language-Action Policies
- Making Foresight Actionable: Repurposing Representation Alignment in World Action Models
- GeoSem-WAM: Geometry- and Semantic-Aware World Action Models
- EgoWAM: World Action Models Beyond Pixels with In-the-Wild Egocentric Human Data
- Point Tracking Improves World Action Models
- FlowWAM: Optical Flow as a Unified Action Representation for World Action Models
- Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising
- Learning 4D Geometric Priors for Inference-Efficient World Action Models
- DreamWAM: Beyond RGB Future Prediction for World Action Models
- MV-WAM: Manifold-Aware World Action Model with Value Augmentation
- WorldVLA: Towards Autoregressive Action World Model
- SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space
- ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts
- 4D-WAM: Infusing Spatiotemporal Awareness into World Action Models through Trajectory Fields
- ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models