WorldGuide: Learning Success-Failure Boundaries in Latent World Models for Vision-Language-Action Policies
cs.CV
Submitted: 2026-09-28
Updated: 2026-09-28
Terminology
Sources
- Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- UniVLA: Learning to Act Anywhere with Task-centric Latent Actions
- WorldVLA: Towards Autoregressive Action World Model
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
- Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- OpenVLA: An Open-Source Vision-Language-Action Model
- Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning
- VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators
- Evaluating Real-World Robot Manipulation Policies in Simulation
- World-VLA-Loop: Closed-Loop Learning of Video World Model and VLA Policy
- FAST: Efficient Action Tokenization for Vision-Language-Action Models
- World-Env: Leveraging World Model as a Virtual Environment for VLA Post-Training
- Fast-WAM: Do World Action Models Need Test-time Future Imagination?
- Native Video-Action Pretraining for Generalizable Robot Control
- FLARE: Robot Learning with Implicit World Modeling
- Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models