eRLT: Efficient VLA Reinforcement Learning via Action-Relevant Token Routing
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "eRLT: Efficient VLA Reinforcement Learning via Action-Relevant Token Routing".
Dev: Vision-Language-Action (VLA) models provide strong behavioral priors for robotic manipulation, yet efficiently adapting them to downstream tasks remains challenging.
Rosa: First, who's behind it and why it matters.
Paper summary: Rosa: So, to recap what we've covered so far with eRLT: Efficient VLA Reinforcement Learning via Action-Relevant Token Routing, this paper addresses a real bottleneck in using Vision-Language-Action models for robotics. The authors argue that while these models offer strong behavioral priors for manipulation tasks, efficiently adapting them to new specific tasks through online reinforcement learning is still quite challenging.
Dev: They claim that existing methods fall short because they use either VLA-independent visual encoders or simply compress a fixed layer and token set from the frozen VLA, neither of which explicitly extracts the action-relevant features needed for refining actions and estimating action values efficiently.
Taro: The core thesis is that this task-specific information isn't neatly packaged in a predetermined way; it’s distributed across both the tokens and layers of the frozen VLA, but these useful features change depending on which downstream task you are facing.
Rosa: eRLT claims to solve this by constructing an effective state representation that routes this task-specific action-relevant information across both the tokens and layers of the frozen VLA, which improves sample efficiency in online RL. This is significant because it supports actor-critic learning even when you only have a small number of online interactions.
Dev: The mechanism involves learned routing tokens aggregating features at different depths within selected VLM layers, followed by a lightweight layer router that creates a fixed-dimensional RL token, zt. This zt then feeds into the actor and critic.
Taro: What matters is the two-stage training process: first, teaching the system what internal differences matter for expert actions before online interaction starts, and second, adapting those routing parameters during online RL using critic feedback to estimate action values.
Rosa: The paper shows that this dual routing—token-wise and layer-wise—is key; token-wise allows cues to appear at different positions as the scene or instruction changes, while layer-wise learns how the relative contributions of layer weights adapt across different adaptation tasks.
Dev: Essentially, they are learning *how* to select and combine the most relevant information from the VLA's internal structure on a task-by-task basis rather than relying on a one fixed extraction method.
Taro: And this isn't just theoretical; they tested it across seven simulation tasks and two real-world high-precision manipulation tasks, showing substantial improvements in learning curves and final success rates compared to prior state representations.
Rosa: So the paper boils down to proposing a learned routing mechanism that builds a much more informative state representation by intelligently combining features from different parts of the frozen VLA, thereby making online RL for these complex robotic tasks much more sample-efficient.
Conclusion: Rosa: Considering the title of eRLT: Efficient VLA Reinforcement Learning via Action-Relevant Token Routing, we're looking at a method designed to make adapting powerful Vision-Language-Action models to specific tasks much more efficient using clever routing. The authors are Dehao Huang and his colleagues, and their work is definitely worth paying attention for anyone working on this area.
Dev: I think the real takeaway is that they’ve moved beyond just plugging in standard VLA features; they've figured out how to dynamically select and combine the most useful internal signals from the model based on what action refinement or value estimation actually needs at that moment.
Taro: The implication for autonomy is that we can expect robotic systems to learn new skills faster in real-world scenarios without needing massive amounts of pre-collected interaction data, which could make deploying complex AI agents into physical environments much more practical.
Rosa: Exactly; it suggests a future where robotic agents can handle novel manipulation tasks with better sample efficiency, which is a big step toward making general-purpose physical AI more viable.
Dev: From an engineering standpoint, the implication is that we don't have to be constrained by a fixed architecture for state representation; we can design systems where the representation itself evolves intelligently based on the learning process.
Taro: It means that when a robot encounters something unexpected, its ability to infer what information is actually relevant and how to pull it from its own internal structure becomes much stronger, which is key when dealing with unpredictable environments.
Rosa: So in short, eRLT provides a framework for building smarter state representations by learning the necessary routing logic, and that promises more practical online reinforcement learning for these complex robotic systems.
Dehao Huang, Jianbang Liu, Jianpan Gao, Chao Tang, Zilang Cen, Zedong Dan, Jiaheng Wang, Tingguang Li, †Yue Wang
Southern University of Science and Technology, Shenzhen, China. · Beijing Zhongguancun Academy, Beijing, China. · Samsung Robotics eXperience. · Wuhan University, Wuhan, China. · Sun Yat-sen University, Guangzhou, China.
cs.RO
Submitted: 2026-10-01
Updated: 2026-10-01
Comments: 21 pages, 9 figures, 8 tables
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 90/100
The gist: Vision-Language-Action (VLA) models provide strong behavioral priors for robotic manipulation, yet efficiently adapting them to downstream tasks remains challenging.
Key concepts
- VLA Models
- Vision-Language-Action (VLA) models are AI systems that combine visual understanding, language comprehension, and action generation capabilities. They provide strong behavioral priors for robots but are challenging to adapt quickly to new, specific tasks without extensive retraining.
- Routing Tokens
- These are learned tokens that dynamically aggregate visual-language features from different depths within the frozen VLA layers. They act as task-specific feature extractors, gathering information relevant to the current action at various points in the model's architecture.
- Layer-wise Routing
- This mechanism involves reading the same positions across multiple selected depths of a VLA and learning weights for each depth. This allows the system to adapt which layers are most important for a specific task by adjusting their relative contributions to form the final action token.
Terminology
Summary
Vision-Language-Action (VLA) models provide strong behavioral priors for robotic manipulation, yet efficiently adapting them to downstream tasks remains challenging. The gist: eRLT constructs an effective state representation by routing task-specific action-relevant information across both tokens and layers of a frozen VLA, improving sample efficiency in online reinforcement learning. This method is significant because it addresses the limitation of existing methods that fail to explicitly extract the most useful task-specific features for action refinement and action-value estimation, thereby supporting sample-efficient actor-critic learning from limited online interactions.
The core problem addressed by eRLT
Existing methods construct state representations either with VLA-independent visual encoders or by compressing representations from a predetermined layer and token set. Neither design explicitly extracts the task-specific action-relevant VLA features most useful for downstream action refinement and action-value estimation, which limits sample efficiency under constrained online interaction budgets. The paper demonstrates that the VLA layers and tokens most useful for online RL vary across downstream tasks, motivating eRLT to learn how they are aggregated.
The eRLT architecture
eRLT constructs an effective state representation by routing task-specific action-relevant information across the tokens and layers of the frozen VLA. Specifically:
-
Learned routing tokens dynamically aggregate visual-language features at multiple depths within selected VLM layers.
-
A lightweight layer router combines these layer-wise summaries into a fixed-dimensional RL token, denoted as zt.
-
The routing module is initialized using expert demonstrations to capture features predictive of expert actions and then refined using critic feedback from online interactions for action-value estimation.
The two stages of training
The routing module undergoes a two-stage training process:
-
Before online interaction, the action-prediction loss updates the routing tokens, layer logits, and temporary action probe to teach both routing stages which internal differences matter for expert actions. This stage differs from reconstruction by focusing on preserving task-specific differences needed for action refinement.
-
During online RL, successful and failed transitions reveal which differences affect task progress and outcomes; the critic objective adapts the routing parameters accordingly for action-value estimation, refining the retained information for action-value estimation. The VLA remains frozen throughout both stages.
The two routing mechanisms
eRLT employs two complementary routing mechanisms to construct zt:
-
Token-wise routing: Task-specific cues appear at different positions as the scene and instruction change, allowing K routing tokens to gather different prefix information at each policy step.
-
Layer-wise routing: The same positions are read at every selected depth, and layer weights are learned across observations within each adaptation task to adapt their relative contributions. The final RL token is constructed by combining these layer summaries: zt = Xm i=1 αiu(i)t, where α is the softmax of task-level logits.
Evaluation and results
eRLT was evaluated across seven LIBERO and RoboTwin tasks, as well as two real-world high-precision manipulation tasks: USB connector insertion and motherboard ribbon-cable insertion. Across the seven simulation tasks, eRLT improves mean normalized learning-curve AUC by up to 23.7% over representative baselines. On the two real-world tasks, eRLT improves learning-curve AUC by 108.9% and 46.7%, respectively, relative to the strongest baseline. For USB connector insertion, eRLT raises the final-window success rate from 50% to 90% under the same interaction budget compared to DSRL∗ and RLT. The results show that eRLT achieves the strongest aggregate performance across both benchmarks and demonstrates improved sample efficiency over existing state representations.
Ablation findings
Ablation studies validate the complementary roles of layer-wise routing, token-wise routing, task-specific action-relevant initialization, and online adaptation. Removing any component reduces AUC; for instance, removing token-wise routing leads to a degradation in AUC compared to uniform token pooling. Offline initialization is also found to be influential as an action-predictive starting point reduces the burden on the critic. Increasing the number of routing tokens does not improve performance monotonically, suggesting that a single routing token is often sufficient.
Real-world performance analysis
In real-world experiments, eRLT improves both the rate of online learning and final reliability on physical interaction tasks. For USB connector insertion, eRLT collects 66 autonomous trajectories with 46 successes compared to 57 for DSRL∗ and 20 for RLT. On ribbon-cable insertion, eRLT collects 70 autonomous trajectories with 63 successes compared to 65 for RLT and 53 for DSRL∗.
Improvements for AI systems
As a fastidious researcher, I have analyzed the eRLT (Efficient VLA Reinforcement Learning via Action-Relevant Token Routing) paper. The core contribution is an efficient method for adapting frozen Vision-Language-Action (VLA) models to downstream tasks using online reinforcement learning (RL).
Here are specific improvements and what the resulting AI system can achieve:
)
- Improved Sample Efficiency in Fine-Tuning VLA Models:
eRLT significantly improves the sample efficiency of lightweight online RL by constructing a task-specific state representation (the RL token) that dynamically aggregates information across both VLA tokens and layers. This allows the actor and critic to focus only on the most relevant features for action refinement and value estimation, rather than relying on fixed compression or independent encoders.
- Enhanced Robustness Across Task Domains:
The system can achieve superior performance across diverse manipulation tasks (e.g., LIBERO, RoboTwin) and high-precision real-world applications (e.g., USB connector insertion, motherboard ribbon-cable insertion). Specifically, the paper shows AUC improvements of up to 23.7% on simulation benchmarks and substantial gains in real-world success rates (108.9% for USB insertion) compared to strong baselines.
- Task-Specific Action Relevance Extraction:
The system moves beyond generic feature extraction by learning which VLA layers and tokens are most useful for a specific task through two stages:
a. An expert action prediction phase initializes the routing module to capture features predictive of expert actions.
b. A critic-driven online adaptation phase refines the routing parameters using critic feedback from successful and failed interactions, ensuring the RL token retains cues necessary for precise action-value estimation under limited interaction budgets.
- High-Precision Physical Manipulation:
The improved system can execute complex, fine-grained physical tasks with high reliability. For example, in USB connector insertion, the eRLT system increases the final-window success rate from 50% to 90% under the same interaction budget by consistently correcting small positional and angular misalignments near contact.
- Efficient Deployment of Large Pretrained Models:
By keeping the massive VLA backbone frozen and only training lightweight routing tokens, the system drastically reduces computational overhead compared to methods that require full fine-tuning or reliance on external, VLA-independent encoders (like DSRL). This makes adaptation feasible in resource-constrained real-robot environments.
Sources
- MMFuser: Multimodal Multi-Layer Feature Fuser for Fine-Grained Vision-Language Understanding
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
- Soft Actor-Critic Algorithms and Applications
- Residual Reinforcement Learning for Robot Control
- Adaptation of Generalist Robot Policies with Minimal Data
- GR-RL: Going Dexterous and Precise for Long-Horizon Robotic Manipulation
- UniSteer: Unified Noise Steering for Efficient Human-Guided VLA Adaptation
- $\pi^{*}_{0.6}$: a VLA That Learns From Experience
- Residual Policy Learning
- Improving Robotic Generalist Policies via Flow Reversal Steering
- Steering Your Diffusion Policy with Latent Space Reinforcement Learning
- RL Token: Bootstrapping Online RL with Vision-Language-Action Models
- TORL-VLA: Tactile Guided Online Reinforcement Learning for Contact-Rich Manipulation
- Learning to Act While Waiting: RL Finetuning of Generalist Robot Policies Under Inference Latency
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving