M2Tok: Multi-head Multi-codebook Discrete Action Tokenization for Vision-Language-Action Models
cs.RO, cs.AI, cs.CL, cs.CV
Submitted: 2026-09-16
Updated: 2026-09-17
Comments: ECCV 2026
Code: https://github.com/cpaaax/M2Tok
License: http://creativecommons.org/licenses/by/4.0/
The gist: Recent advancements have successfully adapted autoregressive language models to process multimodal signals, such as images and actions.
Terminology
Abstract
Recent advancements have successfully adapted autoregressive language models to process multimodal signals, such as images and actions. Since raw action signals are continuous, effective tokenization is essential to map high-dimensional inputs into compact discrete tokens for autoregressive processing. However, existing discrete action tokenizers often suffer from high reconstruction loss, failing to preserve the fine-grained dynamics required for precise control. This "discretization bottleneck" significantly limits the performance ceiling of downstream Vision-Language-Action (VLA) models. To address this, we propose M squared Tok, a Multi-head Multi-codebook Action Tokenizer designed to minimize reconstruction error and enhance policy performance. Our approach introduces two key structural innovations: (1) we decompose the latent action features into multiple heads, enabling the model to implicitly align specific heads with distinct action dimensions; (2) we assign independent codebooks to each head for quantization. By leveraging the combinatorial nature of multiple codebooks, we significantly expand the representational expressivity of the tokenizer, leading to substantially lower reconstruction loss compared to previous methods. We evaluate the M squared Tok-based VLA on the RoboTwin, Simpler-Env, and 3 zero-shot real-world tasks. Experimental results demonstrate our method not only achieves superior reconstruction fidelity but also significantly boosts the success rate of VLA models. Comprehensive ablation studies further confirm the effectiveness of the multi-head and multi-codebook mechanisms. Code is available at https://github.com/cpaaax/M2Tok.
Sources
- GPT-4 Technical Report
- RT-H: Action Hierarchies Using Language
- RT-1: Robotics Transformer for Real-World Control at Scale
- UniVLA: Learning to Act Anywhere with Task-centric Latent Actions
- GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
- The Llama 3 Herd of Models
- Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- OpenVLA: An Open-Source Vision-Language-Action Model
- Unified Video Action Model
- Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies
- DeepSeek-V3 Technical Report
- HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model
- RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation
- Perceiver-Actor: A Multi-Task Transformer for Robotic Manipulation
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- FAST: Efficient Action Tokenization for Vision-Language-Action Models
- CLIPort: What and Where Pathways for Robotic Manipulation
- Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving