VIDEAS: Distilling Explicit Action Semantics from Demonstration Videos for World Models via Prior-Guided Simulation
cs.RO
Submitted: 2026-09-27
Updated: 2026-09-27
Terminology
Sources
- Cosmos 3: Omnimodal World Models for Physical AI
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- Qwen2.5-VL Technical Report
- Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation
- AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems
- WorldPrediction: A Benchmark for High-level World Modeling and Long-horizon Procedural Planning
- Planning with Reasoning using Vision Language World Model
- Learning Interactive World Model for Object-Centric Reinforcement Learning
- Learning and Leveraging World Models in Visual Representation Learning
- The Llama 3 Herd of Models
- World Models
- One Demo Is All It Takes: Planning Domain Derivation with LLMs from A Single Demonstration
- From Word to World: Can Large Language Models be Implicit Text-based World Models?
- Qwen3 Technical Report
- LogicEnvGen: Task-Logic Driven Generation of Diverse Simulated Environments for Embodied AI
- Dyn-O: Building Structured World Models with Object-Centric Representations
- UniDomain: Pretraining a Unified PDDL Domain from Real-World Demonstrations for Generalizable Robot Task Planning
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving