ECHO: Event-Augmented Context with Hindsight and Outlook for Wrist-Only Manipulation
cs.CV, cs.RO
Submitted: 2026-09-28
Updated: 2026-09-28
Project page: https://echo-wam.github.io
Terminology
Sources
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- OpenVLA: An Open-Source Vision-Language-Action Model
- CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes
- Event-VLA: Action-Conditioned Event Fusion for Robust Vision-Language-Action Model
- WorldVLA: Towards Autoregressive Action World Model
- Ctrl-World: A Controllable Generative World Model for Robot Manipulation
- Latent Action Pretraining from Videos
- CLAM: Continuous Latent Action Models for Robot Learning from Unlabeled Demonstrations
- Motus: A Unified Latent Action World Model
- CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning
- FLARE: Robot Learning with Implicit World Modeling
- $\mu$VLA: On Recurrent Memory for Partially Observable Manipulation in VLA Models
- AtlasVLA: Persistent World-Ego State Modeling for Vision-Language-Action Models
- Mem-World: Memory-Augmented Action-Conditioned World Models for Persistent Robot Manipulation
- SEBVS: Synthetic Event-based Visual Servoing for Robot Navigation and Manipulation
- Event-based Vision for Early Prediction of Manipulation Actions
- REASSEMBLE: A Multimodal Dataset for Contact-rich Robotic Assembly and Disassembly
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models