ActionCodec: What Makes for Good Action Tokenizers
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.
Dev: Today's paper: "ActionCodec: What Makes for Good Action Tokenizers".
Rosa: The gist The introduction establishes that action tokenization design remains unanswered, and ActionCodec introduces a high-performance action tokenizer guided by information-theoretic insights to enhance VLA optimization.
Dev: First, who's behind it and why it matters.
Title and authors: Rosa: We’ve talked about the title and authors of ActionCodec: What Makes for Good Action Tokenizers, and it seems they are really digging into the fundamental design choices that have been overlooked in current action tokenization research.
Dev: Right. The core idea is that most existing work has been too focused on how accurately a tokenizer can reconstruct data, but this paper shifts the focus to how those token designs specifically influence the training process for Vision-Language-Action models.
Taro: So, what’s the big picture they are pointing out about why this design question has been left unanswered? What is that fundamental gap?
Rosa: They say that action tokenization has been primarily designed around reconstruction fidelity, and that this focus has missed its direct impact on VLA optimization. The central question they are addressing is what actually makes for a good action tokenizer in the context of learning an action sequence from visual and language inputs.
Dev: They categorize existing tokenizers into heuristic methods, semi-data-driven ones using BPE on frequency signals, and data-driven methods based on Vector Quantization to learn latent discrete representations.
Taro: It seems like they’re arguing that the gap exists because existing Vector Quantization approaches are often treated as black boxes, and there’s a lack of understanding about the specific properties of those VQ tokenizers that either help or hinder VLA training optimization.
Rosa: That's it. They are arguing that we need to look beyond just reconstruction error and start analyzing the complex training dynamics of the VLA backpropagation process when tokens are involved.
Dev: So, they propose a set of four design principles—maximized temporal token overlap, minimized vocabulary redundancy, enhanced multimodal mutual information between tokens and context, and token independence—to guide this new design approach.
Taro: Those principles sound very specific. They’re not just vague goals; they’re measurable things derived from analyzing the expected negative log-likelihood loss decomposition.
Rosa: Exactly. They quantify overlap as a measure of topological stability where small action changes cause jumps in token space, and they bound capacity by an entropy limit to stop encoding unnecessary high-frequency noise.
Dev: And they also have to balance perceptual alignment between visual-language grounding and residual grammar so the model doesn't just over-rely on past actions at the expense of actually understanding where it is in the environment.
Taro: It sounds like a very holistic way to look at it—not just one component in isolation, but how all these factors interact during training.
Rosa: That’s the point. They are establishing a design methodology based on information theory to answer that fundamental question of what makes for good action tokenizers, and then they build ActionCodec around those rules.
The paper's summary: Dev: Now we’re getting into the actual summary of ActionCodec: What Makes for Good Action Tokenizers, which outlines the architecture and how it implements these design principles.
Rosa: They introduce the ActionCodec architecture, which uses a Perceiver-like transformer because it offers inherent flexibility to model diverse token relationships and handle variable-length action sequences effectively.
Taro: So they aren't just sticking to a standard RNN or Transformer structure; they’re using something more flexible to accommodate the complexity of action tokens.
Dev: They also use Vector Quantization, or VQ, for tokenization, but they treat it differently than previous work by focusing on understanding what specific VQ properties actually facilitate or obstruct VLA optimization.
Rosa: They refine the standard VQ approach by incorporating Residual Vector Quantization post-training to improve reconstruction fidelity while still maintaining that topological stability we talked about earlier.
Taro: That sounds like they’re using the architecture not just for what it can do, but for how it can be tuned to meet those specific design requirements.
Dev: They also incorporate embodiment-specific soft prompts into the model, which they suggest are key for facilitating knowledge transfer across different robotic platforms and enabling zero-shot action re-targeting.
Rosa: That’s a big practical step. It means you can adapt the system to novel hardware with minimal fine-tuning just by using those prompts.
Taro: So, the architecture itself is built to be adaptable, which ties back into those principles of independence and multimodal context enhancement they talked about earlier.
Dev: It really is a complete package. They’ve combined flexible architecture, targeted quantization techniques, and specific prompting strategies to create a system that aims to meet all those design requirements simultaneously.
Rosa: So the summary is that ActionCodec isn't just another tokenization scheme; it’s an integrated system built from the ground up to optimize VLA performance by explicitly targeting the training dynamics of the tokens.
The paper's improvements: Taro: So now we’re looking at what they claim are the specific improvements ActionCodec offers over previous tokenizers, beyond just having a new architecture.
Rosa: One major improvement is that they claim it achieves better generalization across diverse platforms. They report that their implementation can achieve a ninety-five point five percent success rate without any robotics pre-training on LIBERO, which is impressive compared to models initialized from general VLMs.
Dev: That number is significant because it demonstrates superior generalization capabilities across different robotic systems, and they link this improvement directly to maximizing the Overlap Rate as a primary design requirement for the action tokenizer.
Taro: So, so improving that overlap rate directly translates into better training efficiency by mitigating overfitting because it keeps things stable in the latent space.
Rosa: That’s right. And they also show that their approach is better for visual-language alignment; they prefer Time Contrastive Learning and CLIP-based objectives over InfoNCE contrastive loss because those yield higher overlap rates and superior training stability.
Dev: They also highlight that their design regarding residual grammar is much more robust than using self-attention or causal architectures because the latter tend to cause reliance on historical tokens leading to temporal hallucinations.
Taro: So, in short, they are claiming tangible performance gains across the board by addressing those specific bottlenecks we discussed.
Rosa: They also mention that ActionCodec can perform zero-shot action re-targeting thanks to those soft prompts, allowing for accelerated fine-tuning on new platforms with minimal effort.
Dev: And finally, they show that the system can handle real-time control with the highest action throughput while still maintaining superior task performance, which is a key thing when dealing with latency constraints.
Conclusion: Taro: So to wrap up this discussion on ActionCodec: What Makes for Good Action Tokenizers, it seems they’ve shown that by integrating these best practices—focusing on temporal overlap, vocabulary size, and multimodal information—they have created a tokenizer that is much better suited for VLA optimization than anything before it.
Rosa: That’s the main message. It provides a clear set of design guidelines for anyone trying to build next-generation physical intelligence by explicitly considering how action tokens affect the training dynamics, not just reconstruction accuracy.
Dev: So, in summary, ActionCodec achieves superior performance on multiple benchmarks and sets a new standard for VLA models that doesn't rely on robotics pre-training to reach high success rates.
Taro: I think the way they’ve framed the integration with Parallel Decoding, Knowledge Isolation, and Block-wise Autoregressive paradigms shows how adaptable this framework is to different ways we structure our models.
Rosa: It’s a solid piece of work that gives us a clear roadmap for developing these more robust action representations for physical intelligence. We’ll keep an eye on how they scale these ideas in the future.
Dev: Yeah, ActionCodec is definitely worth paying attention to because it addresses that long-standing question about what makes for good action tokenizers in this field.
Knowin AI (Work done during an internship) · Tsinghua University · Tianjin University · Fudan University · Shanghai Innovation Institute
cs.RO, cs.AI
Submitted: 2026-02-17
Updated: 2026-10-08
Code: https://github.com/Stanford-ILIAD/openvla-mini
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 83/100
The gist: The gist The introduction establishes that action tokenization design remains unanswered, and ActionCodec introduces a high-performance action tokenizer guided by information-theoretic insights to
Key concepts
- Overlap Rate (OR)
- This measures the topological stability of action tokens. A high OR means that minor changes in the action trigger small, predictable shifts in token space, indicating a well-structured and stable representation of actions.
- Capacity and Vocabulary Size
- These parameters bound how much information a tokenizer can hold. The paper suggests that excessive capacity leads to encoding high-frequency noise instead of useful information, implying there is an optimal balance for effective action representation.
- Perceptual Alignment
- This refers to balancing the model's reliance on temporal patterns against its ability to ground actions in the visual environment. Achieving good alignment prevents the model from relying too much on past actions while ensuring it remains connected to what it sees.
Terminology
Summary
The gist The introduction establishes that action tokenization design remains unanswered, and ActionCodec introduces a high-performance action tokenizer guided by information-theoretic insights to enhance VLA optimization.
Design Principles for Action Tokenizers
The paper identifies four key desiderata for high-performance action tokens: (i) maximized temporal token overlap between adjacent chunks, (ii) minimized vocabulary redundancy, (iii) enhanced mutual information between tokens and multimodal contexts, and (iv) token independence. These principles are derived from analyzing the expected negative log-likelihood loss decomposition which includes Artifact Entropy, Capacity, and Perceptual Alignment components The Overlap Rate is quantified as a measure of topological stability where minor action perturbations trigger stochastic jumps in token space <ref:2602.15397#pg5>. Capacity and Vocabulary Size are bounded by the entropy upper bound n log2 S, where excessive capacity encodes high-frequency noise <ref:2602.15397#pg5>. Perceptual Alignment is decomposed into Visual-Language alignment and Residual Grammar, requiring a balance to prevent over-reliance on temporal priors at the expense of environmental grounding <ref:2602.15397#pg5>.
ActionCodec Architecture and Enhancements
ActionCodec integrates these optimal design choices by leveraging Residual Vector Quantization (RVQ) posttraining to refine reconstruction fidelity while maintaining topological stability <ref:2602.15397#pg7>. It incorporates embodiment-specific soft prompts to facilitate knowledge transfer across diverse robotic platforms, allowing for zeroshot action re-targeting and accelerated fine-tuning on novel robotic platforms <ref:2602.15397#pg9>. The architecture employs a Perceiver-like transformer architecture due to its inherent flexibility, which supports the encoding of variable-length action sequences <ref:2602.15397#pg5>.
Impacts of Design Choices on VLA Performance
The performance on LIBERO-Goal demonstrates that a higher OR consistently enhances training efficiency, asymptotic convergence, and robustness, with the 70% OR tokenizer achieving a 33.4% success rate within only 500 training steps <ref:2602.15397#pg5>. The token budget n exerts a substantially more dominant influence on resilience to overfitting than vocabulary size S, suggesting that while increased capacity preserves discriminability, it leads to excessive dispersion in latent clusters <ref:2602.15397#pg5>. For Vision-Language Alignment, Time Contrastive Learning (TCL) and CLIP-based objectives are preferred over explicit InfoNCE contrastive loss for naturally yielding higher OR values and superior training stability <ref:2602.15397#pg9>. Regarding Residual Grammar, independent tokens are significantly more robust than Self-Attention (SA) or Causal architectures, as the latter foster an over-reliance on historical tokens leading to temporal hallucinations <ref:2602.15397#pg5>.
Experimental Validation and Results
ActionCodec achieves SOTA performance on multiple simulated and real-world benchmarks, notably reaching a 97.4% average success rate on the LIBERO benchmark without any robotics pre-training <ref:2602.15397#pg2>. The comparison with mainstream action tokenizers shows ActionCodec outpaces all baselines in convergence speed, attaining an 89.5% success rate within 5K training steps, significantly outperforming the FAST baseline <ref:2602.15397#pg5>. Furthermore, ActionCodec achieves the highest action throughput while maintaining superior task performance <ref:2602.15397#pg5>. The study also shows that co-training with community data is essential for successfully internalizing long-tail recovery behaviors on low-cost platforms like SO100 <ref:2602.15397#pg9>.
Integration with Prevailing VLA Paradigms
ActionCodec seamlessly integrates with three prevailing VLA paradigms: Parallel Decoding (PD), Knowledge Isolation (KI), and Block-wise Autoregressive (BAR) <ref:2602.15397#pg5>. The PD variant achieves success rates similar to naive autoregression, while the BAR variant establishes a new SOTA for VLA models without robotics pre-training, achieving a 97.4% average success rate <ref:2602.15397#pg5>. The KI paradigm remains superior to its FAST-based counterpart, suggesting it may be better suited for large-scale VLA pre-training <ref:2602.15397#pg9>.
Real-World Application and Transferability
In real-world evaluations, ActionCodec with co-training successfully leverages the broader representational prior to navigate complex distributions of corrective actions during inference, enabling robust closed-loop recovery <ref:2602.15397#pg9>. The model demonstrates effective hardware-agnostic action semantics when tested across different robotic platforms, showing highly consistent action patterns even when decoded into trajectories for LIBERO-Franka, DROID-Franka, and xArm <ref:2602.15397#pg9>. The pre-trained variant of ActionCodec (w/ PT) demonstrates superior optimization dynamics, exhibiting lower L2 reconstruction error and a higher Overlap Rate compared to the w/o PT counterpart <ref:2602.15397#pg5>.
Conclusion
The paper concludes that by integrating identified best practices, ActionCodec achieves superior performance across diverse simulation and real-world benchmarks <ref:2602.15397#pg5>. Future work will focus on scaling these representations to achieve robust in-the-wild transfer across a broader variety of robotic embodiments <ref:2602.15397#pg9>.
Impact Statement
This paper presents work whose goal is to advance the field of Machine Learning <ref:2602.15397#pg5>. The paper's findings provide a clear roadmap for the community to develop next-generation physical intelligence <ref:2602.15397#pg9>.
References
Bai, C., Xu, H., and Li, X. Embodied-ai with large models: research and challenges <ref:2602.15397#pg5>.
Belkhale, S. and Sadigh, D. Minivla: A better vla with a smaller footprint <ref:2602.15397#pg5>.
Bjorck, J., Castaneda, F., Cherniadev, N., Da, X., Ding, R., ˜ Fan, L., Fang, Y., Fox, D., Hu, F., Huang, S., et al. Gr00t n1: An open foundation model for generalist humanoid robots <ref:2602.15397#pg5>.
Black, K., Brown, N., Driess, D., Esmail, A., Equi, M., Finn, C., Fusai, N., Groom, L., Hausman, K., Ichter, B., et al. pi0: A vision-language-action flow model for general robot control <ref:2602.15397#pg5>.
Bu, Q., Yang, Y., Cai, J., Gao, S., Ren, G., Yao, M., Luo, P., and Li, H. Univla: Learning to act anywhere with taskcentric latent actions <ref:2602.15397#pg5>.
Cadene, R., Alibert, S., Soare, A., Gallouedec, Q., Zouitine, A., Palma, S., Kooijmans, P., Aractingi, M., Shukor, M., Aubakirova, D., Russi, M. Capuano F. Lerobot: Stateof-the-art machine learning for real-world robotics in pytorch <ref:2602.15397#pg5>.
Chi, C., Xu, Z., Feng, S., Cousineau, E., Du, Y., Burchfiel, B., Tedrake, R., and Song, S. Diffusion policy: Visuomotor policy learning via action diffusion <ref:2602.15397#pg5>.
Dao, T. Flashattention-2: Faster attention with better parallelism and work partitioning <ref:2602.15397#pg5>.
Dong, Z., Liu, Y., Li, Y., Zhao, H., and HAO, J. Conditioning matters: Training diffusion policies is faster than you think <ref:2602.15397#pg5>.
Driess, D., Xia, F., Sajjadi, M. S., Lynch, C., Chowdhery, A., Ichter, B., Wahid, A., Tompson, J., Vuong, Q., Yu, T., et al. Palm-e: An embodied multimodal language model <ref:2602.15397#pg5>.
Driess, D., Springenberg, J. T., Ichter, B., Yu, L., Li-Bell, A., Pertsch, K., Ren, A. Z., Walke, H., Vuong, Q., Shi, L. X., and Levine, S. Knowledge insulating vision-languageaction models: Train fast <ref:2602.15397#pg5>.
Goyal, A., Hadfield, H.
Improvements for AI systems
-
ActionCodec implementation can achieve
a 95.5% success rate without any robotics pre-training
on LIBERO, demonstrating superior generalization capabilities across diverse robotic platforms compared to existing models initialized from general VLMs. -
The system will exhibit enhanced training efficiency and
mitigates overfitting
by maximizing theOverlap Rate (OR),
which is established as aprimary design requirement for the action tokenizer.
-
ActionCodec can perform zero-shot action re-targeting, leveraging
Embodiment-specific Soft-prompts
to facilitate knowledge transfer across different robotic platforms, enabling models to adapt to novel hardware with minimal fine-tuning. -
The VLA model can be architecturally adapted using three paradigms: Parallel Decoding (PD), Knowledge Isolation (KI), and Block-wise Autoregressive (BAR), where the BAR variant
achieved the highest performance, establishing a new SOTA on the LIBERO benchmark for models without robotics-specific pre-training.
-
The system can handle real-time control with
highest action throughput
while maintaining superior task performance, overcoming the latency issues of naive tokenization schemes like Binning and String.
Abstract
Vision-Language-Action (VLA) models leveraging the native autoregressive paradigm of Vision-Language Models (VLMs) have demonstrated superior instruction-following and training efficiency. Central to this paradigm is action tokenization, yet its design has primarily focused on reconstruction fidelity, failing to address its direct impact on VLA optimization. Consequently, the fundamental question of what makes for good action tokenizers remains unanswered. In this paper, we bridge this gap by establishing design principles specifically from the perspective of VLA optimization. We identify a set of best practices based on information-theoretic insights, including maximized temporal token overlap, minimized vocabulary redundancy, enhanced multimodal mutual information, and token independence. Guided by these principles, we introduce ActionCodec, a high-performance action tokenizer that significantly enhances both training efficiency and VLA performance across diverse simulation and real-world benchmarks. Notably, on LIBERO, a SmolVLM2-2.2B fine-tuned with ActionCodec achieves a 95.5% success rate without any robotics pre-training. With advanced architectural enhancements, this reaches 97.4%, representing a new SOTA for VLA models without robotics pre-training. We believe our established design principles, alongside the released model, will provide a clear roadmap for the community to develop more effective action tokenizers.
Sources
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- UniVLA: Learning to Act Anywhere with Task-centric Latent Actions
- PaLM-E: An Embodied Multimodal Language Model
- Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better
- VLA-0: Building State-of-the-Art VLAs with Zero Modification
- Actions as Language: Fine-Tuning VLMs into VLAs Without Catastrophic Forgetting
- $\pi^{*}_{0.6}$: a VLA That Learns From Experience
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- Perceiver: General Perception with Iterative Attention
- Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
- MolmoAct: Action Reasoning Models that can Reason in Space
- FASTer: Toward Efficient Autoregressive Vision Language Action Modeling via Neural Action Tokenization
- FAST: Efficient Action Tokenization for Vision-Language-Action Models
- SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model
- SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics
- Representation Learning with Contrastive Predictive Coding
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Qwen2.5 Technical Report
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving