HALO-WA: Hybrid-Attention Latent-Guided Online Reinforcement Learning for World-Action Models
cs.RO, cs.AI
Submitted: 2026-07-05
Updated: 2026-09-21
Code: https://github.com/YeanRoot/HALO-WA
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: World-action (WA) models can generate long-horizon action chunks for general-purpose robotic manipulation, but they remain vulnerable to calibration, perception, and contact-dynamics errors in
Terminology
Abstract
World-action (WA) models can generate long-horizon action chunks for general-purpose robotic manipulation, but they remain vulnerable to calibration, perception, and contact-dynamics errors in real-world precision tasks, often failing in the final few millimeters of alignment or insertion. We propose HALO-WA, a hybrid-attention latent-guided online reinforcement learning (RL) framework for WA models, which leverages latent features and action priors from the WA generation process through a lightweight actor-critic adapter to enable fast online adaptation to real deployment errors. HALO-WA introduces a hybrid-attention structure that preserves the temporal consistency of action chunks while reading task-relevant information from WA latents conditioned on visual context and end-stage correction requirements, thereby producing refined action chunks. We validate HALO-WA on four real-world precision manipulation tasks, where it improves the average success rate from 26.4% for WA-base to 87.1%, outperforming the strongest baseline by 19.2 percentage points while requiring only 45--75 minutes of online training per task. To facilitate reproducibility, we further conduct supplementary simulation experiments in RoboTwin and release the code at https://github.com/YeanRoot/HALO-WA.
Sources
- World Action Models: The Next Frontier in Embodied AI
- World Model for Robot Learning: A Comprehensive Survey
- GigaWorld-0: World Models as Data Engine to Empower Embodied AI
- Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning
- Motus: A Unified Latent Action World Model
- RL-100: Performant Robotic Manipulation with Real-World Reinforcement Learning
- Learning to Manipulate Anywhere: A Visual Generalizable Framework For Reinforcement Learning
- GR-RL: Going Dexterous and Precise for Long-Horizon Robotic Manipulation
- RL Token: Bootstrapping Online RL with Vision-Language-Action Models
- VLA-R1: Enhancing Reasoning in Vision-Language-Action Models
- VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- GigaBrain-0: A World Model-Powered Vision-Language-Action Model
- World Action Models are Zero-shot Policies
- Fast-WAM: Do World Action Models Need Test-time Future Imagination?
- HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models
- OA-WAM: Object-Addressable World Action Model for Robust Robot Manipulation
- Reconstruction or Semantics? What Makes a Latent Space Useful for Robotic World Models
- Neural Machine Translation by Jointly Learning to Align and Translate
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving