CogWAM: Aligning Semantic Cognition with World Action Modeling via Event-Driven Interfaces
cs.RO, cs.CV
Submitted: 2026-09-29
Updated: 2026-09-29
Code: https://github.com/HorizonRobotics/CogWAM
Project page: https://horizonrobotics.github.io/CogWAM
Terminology
Sources
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
- PaLM-E: An Embodied Multimodal Language Model
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- Inner Monologue: Embodied Reasoning through Planning with Language Models
- Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models
- Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- SeqVLA: Sequential Task Execution for Long-Horizon Manipulation with Completion-Aware Vision-Language-Action Model
- DIM-WAM: World-Action Modeling with Diverse Historical Event Memory
- WorldScape Policy 2.0: Empowering Steerable World Action Modeling with Reasoning-Augmented Memory and In-Context Learning
- World Action Models are Zero-shot Policies
- SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space
- Native Video-Action Pretraining for Generalizable Robot Control
- MEM: Multi-Scale Embodied Memory for Vision Language Action Models
- Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- OpenVLA: An Open-Source Vision-Language-Action Model
- Fast-WAM: Do World Action Models Need Test-time Future Imagination?
- WALL-WM: Carving World Action Modeling at the Event Joints
- HarnessWAM: Bridging Prediction and Deliberation in World Action Models
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving