eRLT: Efficient VLA Reinforcement Learning via Action-Relevant Token Routing
summary
The gist
Vision-Language-Action (VLA) models provide strong behavioral priors for robotic manipulation, yet efficiently adapting them to downstream tasks remains challenging.
In short
eRLT improves sample efficiency in online reinforcement learning for robotic manipulation by efficiently adapting frozen Vision-Language-Action (VLA) models to specific tasks. It constructs a task-specific state representation by routing action-relevant information across tokens and layers of the VLA, allowing the model to learn how to best aggregate features for better action refinement and value estimation from limited online interactions.
Key concepts
- VLA Models
- Vision-Language-Action (VLA) models are AI systems that combine visual understanding, language comprehension, and action generation capabilities. They provide strong behavioral priors for robots but are challenging to adapt quickly to new, specific tasks without extensive retraining.
- Routing Tokens
- These are learned tokens that dynamically aggregate visual-language features from different depths within the frozen VLA layers. They act as task-specific feature extractors, gathering information relevant to the current action at various points in the model's architecture.
- Layer-wise Routing
- This mechanism involves reading the same positions across multiple selected depths of a VLA and learning weights for each depth. This allows the system to adapt which layers are most important for a specific task by adjusting their relative contributions to form the final action token.
Terminology used across episodes
This episode discusses
- eRLT: Efficient VLA Reinforcement Learning via Action-Relevant Token Routing · Paper Radio
- MMFuser: Multimodal Multi-Layer Feature Fuser for Fine-Grained Vision-Language Understanding
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
- Soft Actor-Critic Algorithms and Applications
- Residual Reinforcement Learning for Robot Control
- Adaptation of Generalist Robot Policies with Minimal Data · Paper Radio
- GR-RL: Going Dexterous and Precise for Long-Horizon Robotic Manipulation
- UniSteer: Unified Noise Steering for Efficient Human-Guided VLA Adaptation
- pi* 0.6: a VLA That Learns From Experience
- Residual Policy Learning
- Improving Robotic Generalist Policies via Flow Reversal Steering
- Steering Your Diffusion Policy with Latent Space Reinforcement Learning
- RL Token: Bootstrapping Online RL with Vision-Language-Action Models
- TORL-VLA: Tactile Guided Online Reinforcement Learning for Contact-Rich Manipulation
- Learning to Act While Waiting: RL Finetuning of Generalist Robot Policies Under Inference Latency · Paper Radio
The paper
eRLT: Efficient VLA Reinforcement Learning via Action-Relevant Token Routing · Read on arXiv
Dehao Huang, Jianbang Liu, Jianpan Gao, Chao Tang, Zilang Cen, Zedong Dan, Jiaheng Wang, Tingguang Li, †Yue Wang
Southern University of Science and Technology, Shenzhen, China. · Beijing Zhongguancun Academy, Beijing, China. · Samsung Robotics eXperience. · Wuhan University, Wuhan, China. · Sun Yat-sen University, Guangzhou, China.
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "eRLT: Efficient VLA Reinforcement Learning via Action-Relevant Token Routing".
Dev: Vision-Language-Action (VLA) models provide strong behavioral priors for robotic manipulation, yet efficiently adapting them to downstream tasks remains challenging.
Rosa: First, who's behind it and why it matters.
Paper summary: Rosa: So, to recap what we've covered so far with eRLT: Efficient VLA Reinforcement Learning via Action-Relevant Token Routing, this paper addresses a real bottleneck in using Vision-Language-Action models for robotics. The authors argue that while these models offer strong behavioral priors for manipulation tasks, efficiently adapting them to new specific tasks through online reinforcement learning is still quite challenging.
Dev: They claim that existing methods fall short because they use either VLA-independent visual encoders or simply compress a fixed layer and token set from the frozen VLA, neither of which explicitly extracts the action-relevant features needed for refining actions and estimating action values efficiently.
Taro: The core thesis is that this task-specific information isn't neatly packaged in a predetermined way; it’s distributed across both the tokens and layers of the frozen VLA, but these useful features change depending on which downstream task you are facing.
Rosa: eRLT claims to solve this by constructing an effective state representation that routes this task-specific action-relevant information across both the tokens and layers of the frozen VLA, which improves sample efficiency in online RL. This is significant because it supports actor-critic learning even when you only have a small number of online interactions.
Dev: The mechanism involves learned routing tokens aggregating features at different depths within selected VLM layers, followed by a lightweight layer router that creates a fixed-dimensional RL token, zt. This zt then feeds into the actor and critic.
Taro: What matters is the two-stage training process: first, teaching the system what internal differences matter for expert actions before online interaction starts, and second, adapting those routing parameters during online RL using critic feedback to estimate action values.
Rosa: The paper shows that this dual routing—token-wise and layer-wise—is key; token-wise allows cues to appear at different positions as the scene or instruction changes, while layer-wise learns how the relative contributions of layer weights adapt across different adaptation tasks.
Dev: Essentially, they are learning *how* to select and combine the most relevant information from the VLA's internal structure on a task-by-task basis rather than relying on a one fixed extraction method.
Taro: And this isn't just theoretical; they tested it across seven simulation tasks and two real-world high-precision manipulation tasks, showing substantial improvements in learning curves and final success rates compared to prior state representations.
Rosa: So the paper boils down to proposing a learned routing mechanism that builds a much more informative state representation by intelligently combining features from different parts of the frozen VLA, thereby making online RL for these complex robotic tasks much more sample-efficient.
Conclusion: Rosa: Considering the title of eRLT: Efficient VLA Reinforcement Learning via Action-Relevant Token Routing, we're looking at a method designed to make adapting powerful Vision-Language-Action models to specific tasks much more efficient using clever routing. The authors are Dehao Huang and his colleagues, and their work is definitely worth paying attention for anyone working on this area.
Dev: I think the real takeaway is that they’ve moved beyond just plugging in standard VLA features; they've figured out how to dynamically select and combine the most useful internal signals from the model based on what action refinement or value estimation actually needs at that moment.
Taro: The implication for autonomy is that we can expect robotic systems to learn new skills faster in real-world scenarios without needing massive amounts of pre-collected interaction data, which could make deploying complex AI agents into physical environments much more practical.
Rosa: Exactly; it suggests a future where robotic agents can handle novel manipulation tasks with better sample efficiency, which is a big step toward making general-purpose physical AI more viable.
Dev: From an engineering standpoint, the implication is that we don't have to be constrained by a fixed architecture for state representation; we can design systems where the representation itself evolves intelligently based on the learning process.
Taro: It means that when a robot encounters something unexpected, its ability to infer what information is actually relevant and how to pull it from its own internal structure becomes much stronger, which is key when dealing with unpredictable environments.
Rosa: So in short, eRLT provides a framework for building smarter state representations by learning the necessary routing logic, and that promises more practical online reinforcement learning for these complex robotic systems.
More episodes
- 2610.11003-ActiveReg: Information-Driven Active Regional Probing for Partial-to-Full Bone Registration
- 2610.10905-Informationally Decoupled Trajectory Design for Sim-to-Real System Identification
- 2610.10934-Higher-Order Morphology Priors for Quadruped Reinforcement Learning Under Actuator Degradation
- 2610.10949-Noise-Induced Navigation in Non-convex Domains and Compact Manifolds
- 2610.10962-iAm.md: Robot Skill Self-Assessment through Agentic Introspection for Unknown Open-Vocabulary Domains
- 2610.11054-A Reconfigurable Fabric Based Pneumatic Actuator with Button Fastened Constraint Modules for Multi Mode Actuation
- 2610.11308-Distributed Relative Localization for Homogeneous Multi-Robot Systems through UWB Ranging and Limited Communications
- 2610.11072-Towards Path-Creative Navigation: Robot Navigation through Embodied Interaction
- 2610.11119-FOCUS: From Privileged States to RGB-D with Controlled Modality Switching and Representation Alignment
- 2610.11141-Distributed Relative Localization Based on Ultra-WideBand and LiDAR for Multi-robot with Limited Communication