Attention from Action, for Action: Emergent Visual Bottlenecks for Policy Learning
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.
Dev: Today's paper: "Attention from Action, for Action".
Rosa: The gist: Seeker, an action-supervised module that learns where visual evidence is needed for visuomotor control, turns observation–action data into a progression-aware ROI,
Dev: First, who's behind it and why it matters.
Paper summary: Rosa: To start off, this paper, "Attention from Action, for Action: Emergent Visual Bottlenecks for Policy Learning," argues that visual bottlenecks can improve how robots learn to move because they separate where the robot needs to look from how the robot actually acts.
Dev: They point out that most existing methods for finding these important visual areas rely on external spatial labels, like gaze information or object classes, but they're proposing a label-free alternative.
Taro: The research shows that action-derived crops can be useful spatial priors because they don't need extra labels, but those crops can become misaligned if the task gets more complex or the robot's state changes continuously.
Rosa: That misalignment is a key problem, and Seeker is introduced to solve it by learning the way to map action back to a region of interest directly from observation-action streams.
Dev: Seeker starts with a task- and state-conditioned query over frozen DINOv3 patch features, which they then iteratively refine using visual evidence gathered from image patches during training.
Taro: The architecture involves a multi-head attention head gating linear linear patch feat, where the system produces context along with attention maps and head scores.
Rosa: And what's interesting is that this readout isn't a single lookup; it’s an iterative search process that updates the query based on visual evidence, which means its focus can shift as the task stage changes.
Dev: This emergent ROI extraction is then trained using a diffusion action-prediction loss, which allows the ROI to actually emerge without needing any spatial supervision during training.
Taro: What matters for autonomy is that this system recovers policy-useful ROIs that nearly match the privileged Oracle ROI reference even though it has no spatial labels.
Conclusion: Rosa: So, looking at this paper's title, "Attention from Action, for Action: Emergent Visual Bottlenecks for Policy Learning," it really captures the essence of what they did—they are using action to find where to look.
Dev: The authors, including Zheyu Zhuang and Ruiyu Wang and others, have shown that this emergent visual bottleneck is a way to improve data efficiency in visuomotor learning.
Taro: What this means for us is that we don't necessarily need perfect spatial labels like bounding boxes to guide a robot's vision; action itself can provide enough signal.
Rosa: It suggests that the visual structure the robot recovers from just watching it act is actually very useful for policy learning, and it’s robust enough to handle real-world changes.
Dev: The results show that this approach raises average simulation success from forty-two point six percent up to sixty-two point six percent, and in the real world, it boosts in-domain success from forty-eight point three percent to seventy-six point seven percent over the best baseline they tested against.
Taro: And one of the most practical things they show is that these learned ROIs are reusable; they can be used for mask-guided augmentation and improve robustness under changes in lighting or background, raising shifted-condition success from twenty point zero percent to sixty point zero percent.
Rosa: So, simply put, this paper shows that action supervision can recover a useful spatial bottleneck interface between perception and policy learning without relying on those external annotations we usually have to add.
Department of Robotics, Perception and Learning, KTH Royal Institute of Technology, Sweden · Department of Computer Science, University of Freiburg, Germany · Department of Informatics, Universitat Hamburg, Germany
cs.RO
Submitted: 2026-08-13
Updated: 2026-10-08
Comments: [CoRL 2026] Code: https://github.com/zheyu-zhuang/seeker
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 83/100
The gist: The gist: Seeker, an action-supervised module that learns where visual evidence is needed for visuomotor control, turns observation–action data into a progression-aware ROI, exposing an explicit
Key concepts
- Seeker
- An action-supervised module that learns the necessary visual focus for control. It takes observation-action data and creates a dynamic Region of Interest (ROI) that changes as the task progresses, explicitly showing where visual input is most needed for the robot to act correctly.
- Robot-conditioned Query
- The initial query used by Seeker is conditioned on both the current task embedding and the robot's proprioceptive state. This means the module knows what it's supposed to look for based on what it's doing and where it is in the overall task sequence.
- Action-supervised ROI Interface
- This process uses action supervision to generate an explicit spatial bottleneck (the ROI). This learned attention map is then used downstream for tasks like cropping images or filtering point clouds, providing a concrete visual focus derived only from movement data.
Terminology
Summary
The gist: Seeker, an action-supervised module that learns where visual evidence is needed for visuomotor control, turns observation–action data into a progression-aware ROI, exposing an explicit spatial bottleneck without relying on semantic boxes, gaze, language grounding, or affordance annotations.
How it works
Seeker is proposed as a task- and state-conditioned readout that learns attention from action by starting from frozen DINO features and iteratively updating a query with gathered visual evidence to produce progression-aware ROIs solely from action supervision. This readout is not a static crop predictor, but rather one whose focus can shift as the task stage and robot state change.
The Seeker architecture involves several key components for iterative refinement:
-
Robot-conditioned Query: The initial query Q(0) is task- and state-conditioned, defined as Q(0) = FiLM(Z, s) = (1 + γ(s)) ⊙ Z + β(s), where Z is the task embedding and s is the proprioceptive state.
-
Iterative Gated Readout: Seeker applies gated cross-attention readout for T refinement steps, where the current query Q∗ = ϕq(Q), MHA(Q∗, K, V) determines the attention map and context.
-
Head Gating: To capture diverse spatial cues, Seeker splits DINO patch features into multiple readout heads and uses query-dependent head gating to update the query based on learned head weights ωh = softmaxh Ch, ϕg(Q¯∗).
Action-supervised ROI Interface
The final context Ce is concatenated with proprioception and a task embedding, then passed to a diffusion action head trained with the standard noise-prediction objective. After multi-task Seeker training, this diffusion head is discarded, and the frozen Seeker readout is used for downstream policies. The resulting ROI attention map Ae is converted into a patch-level ROI by selecting the smallest set of tokens whose cumulative mass exceeds a top p threshold. This explicit ROI is then reused in three downstream forms: RGB cropping, mask-guided background augmentation, and point-cloud filtering.
Performance and Robustness
Seeker improves data efficiency and robustness over no-crop, augmentation, and action-derived crop baselines. In simulation, Seeker raises average success from 42.6% to 62.6%, while in the real world, it raises in-domain success from 48.3% to 76.7% over the best baseline. Furthermore, Seeker ROIs support mask-guided augmentation and improve real-world robustness under both lighting and background shifts, raising average shifted-condition success from 20.0% to 60.0%.
Evaluation Across Modalities
The evaluation tested three claims about action-supervised visual bottlenecks:
-
Action supervision can recover policy-useful ROIs, as Seeker nearly matches the privileged Oracle ROI reference for policy learning without spatial labels.
-
The resulting bottleneck improves policy learning, raising average simulation success from 42.6% to 62.6% and real-world indomain success from 48.3% to 76.7% over the best baseline.
-
The learned ROI is reusable and supports robustness, improving real-world robustness under both lighting and background shifts from 20.0% to 60.0%.
Ablations and Limitations
Ablation studies show that the main benefit comes from exposing the progression-aware ROI as an explicit crop, not merely telling the network where the box is, as low-res cropping nearly matches Seeker on average. Seeker Localizes Rather Than Controls, as direct rollout from its context achieves near-zero average success across tasks. Limitations include that demonstration coverage matters more than raw count, and the training overhead is roughly 1.5–1.8× over training the downstream policy alone.
Conclusion
Seeker turns observation–action data into a progression-aware ROI, exposing an explicit spatial bottleneck without relying on semantic boxes, gaze, language grounding, or affordance annotations. These results point to action-grounded visual bottlenecks as a practical interface between perception and policy learning. This highlights that the visual structure recovered from action supervision depends on the coverage and variation present in the demonstrations.
References
[1] C. Wang, L. Fan, J. Sun, R. Zhang, L. Fei-Fei, D. Xu, Y. Zhu, and A. Anandkumar Mimicplay: Long-horizon imitation learning by watching human play In Conference on Robot Learning (CoRL), 2023
[2] R. Takizawa, I. Karino, K. Nakagawa, Y. Ohmura, and Y. Kuniyoshi Enhancing reusability of learned skills for robot manipulation via gaze information and motion bottlenecks IEEE Robotics and Automation Letters (RA-L), 2025
[3] S. James, K. Wada, T. Laidlow, and A. J Davison Coarse-to-fine q-attention: Efficient learning for visual robotic manipulation via discretisation In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022
[4] R. Wang, Z. Zhuang, D. Kragic, and F. T Pokorny PALM: Enhanced generalizability for local visuomotor policies via perception alignment IEEE Robotics and Automation Letters (RA-L), 2026
[5] S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, Q Jiang, C Li, J Yang, H Su, J Zhu., and L Zhang Grounding DINO: marrying DINO with grounded pre-training for open-set object detection In European Conference on Computer Vision (ECCV), 2024
[6] A. Stone, T. Xiao, Y. Lu, K Gopalakrishnan, K.-H Lee, Q Vuong, P Wohlhart, S Kirmani, B Zitkovich, F Xia., C Finn., and K Hausman Open-world object manipulation using pretrained vision-language models In Conference on Robot Learning (CoRL), 2023
[7] T. Kwon, N Di Palo, and E Johns Language models as zero-shot trajectory generators IEEE Robotics and Automation Letters (RA-L), 2024
[8] J Hu, L Wang, S Li, Y Jiang, X Li, P Weng., and Y Ban Generalizable coarse-to-fine robot manipulation via language-aligned 3d keypoints In International Conference on Learning Representations (ICLR), 2026
[9] V Myers, C Zheng, O Mees, K Fang, and S Levine Policy adaptation via language optimization: Decomposing tasks for few-shot imitation In Conference on Robot Learning (CoRL), 2024
[10] M Shridhar, L Manuelli, and D Fox Perceiver-actor: A multi-task transformer for robotic manipulation In Conference on Robot Learning (CoRL), 2023
[11] A Goyal, V Blukis, J Xu, Y Guo, Y.-W Chao., and D Fox Rvt2: Learning precise manipulation from few demonstrations In Proceedings of Robotics: Science and Systems (RSS), 2024
[12] O Sim’eoni, H V Vo, M Seitzer, F Baldassarre, M Oquab, C Jose, V Khalidov, M Szafraniec., S Yi., M Ramamonjisoa, F Massa., D Haziza., L Wehrstedt., J Wang., T Darcet.
Improvements for AI systems
-
Bold ROI interface for policy learning: Seeker turns
action supervision into an explicit ROI interface without spatial labels,
which allows policies to learn from aprogression-aware ROI
that shifts with task stage and robot state, improving data efficiency over baselines by raisingaverage in-domain success from the best baseline’s 48.3 to 76.7%.
-
Robust background generalization: The learned ROI supports mask-guided augmentation where
Seeker-Guided augmentation preserves predicted control-relevant regions while perturbing task-irrelevant appearance, reducing the burden on the policy to learn under visually distracting textures.
-
Transferable spatial filtering: Seeker's image-plane ROI can serve as a
plug-and-play spatial filter for point-cloud policies,
improving point cloud success rates by30.3 absolute points on average
over manual workspace cropping, even when trained only from RGB action supervision.
Sources
- Generalizable Learning for Frequency-Domain Channel Extrapolation under Distribution Shift
- Revisiting Feature Prediction for Learning Visual Representations from Video
- The Temporal Trap: Entanglement in Pre-Trained Visual Representations for Visuomotor Policy Learning
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving