WorldGuide: Goal-Directed Video World Model for Procedural Task Execution
cs.CV
Submitted: 2026-10-08
Updated: 2026-10-08
Terminology
Sources
- World Simulation with Video Foundation Models for Physical AI
- Qwen2.5-VL Technical Report
- Autoregressive Video Generation without Vector Quantization
- NovaPlan: Zero-Shot Long-Horizon Manipulation via Closed-Loop Video Language Planning
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- LTX-2: Efficient Joint Audio-Visual Foundation Model
- CollabVR: Collaborative Video Reasoning with Vision-Language and Video Generation Models
- OpenVLA: An Open-Source Vision-Language-Action Model
- HunyuanVideo: A Systematic Framework For Large Video Generative Models
- SkyReels-V3 Technique Report
- PhysAgent: Reflective Agentic Physics Control for Physically Plausible Video Generation
- Yume-1.5: A Text-Controlled Interactive World Generation Model
- Yume: An Interactive World Generation Model
- VideoWorld 2: Learning Transferable Knowledge from Real-world Videos
- EgoForge: Goal-Directed Egocentric World Simulator
- WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling
- Bernini: Latent Semantic Planning for Video Diffusion
- MAGI-1: Autoregressive Video Generation at Scale
- ASTRA: Automated Synthesis of agentic Trajectories and Reinforcement Arenas
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models