SimpleMemVLA: A Simple but Effective Native-Video Memory for Vision-Language-Action Models
cs.CV, cs.LG, cs.RO
Submitted: 2026-09-02
Updated: 2026-09-25
Comments: 29 pages, 12 figures
Code: https://github.com/wadeKeith/SimpleMemVLA
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Qwen3-VL Technical Report
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- WorldVLA: Towards Autoregressive Action World Model
- GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
- RMBench: Memory-Dependent Robotic Manipulation Benchmark with Insights into Policy Design
- Memory, Benchmark & Robots: A Benchmark for Solving Complex Tasks with Reinforcement Learning
- RoboMME: Benchmarking and Understanding Memory for Robotic Generalist Policies
- LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models
- Gated Memory Policy: In-Context Memorization and Adaptation
- Ctrl-World: A Controllable Generative World Model for Robot Manipulation
- ECHO: Continuous Hierarchical Memory for Vision-Language-Action Models
- ChainVLA: Chaining Vision-Language-Action Queries through a Unified Execution State for Long-Horizon Manipulation
- NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- ContextVLA: Vision-Language-Action Model with Amortized Multi-Frame Context
- HAMLET: Switch your Vision-Language-Action Model into a History-Aware Policy
- RoboMemArena: A Comprehensive and Challenging Robotic Memory Benchmark
- RoboMemory: A Brain-inspired Multi-memory Agentic Framework for Interactive Environmental Learning in Physical Embodied Systems
- MAP-VLA: Memory-Augmented Prompting for Vision-Language-Action Model in Robotic Manipulation
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models