RoboMME-Interference: Benchmarking Robot Memory Under Interference
cs.RO, cs.AI, cs.LG
Submitted: 2026-06-21
Updated: 2026-08-26
Comments: 9 pages, 5 figures. Updated results; added subgoal-memory systems
Project page: https://robotmemorybench.com
License: http://creativecommons.org/licenses/by/4.0/
The gist: Robots deployed in realistic settings will accumulate experience across many sessions and tasks over their deployment.
Terminology
Abstract
Robots deployed in realistic settings will accumulate experience across many sessions and tasks over their deployment. The robot's tasks may often require it to remember information from multiple sessions ago, making long-context robot memory important for real-world deployments. However, most robot-memory benchmarks today are based on single episodes or a short context. To measure how current robot memory systems perform on longer sessions with more distractions, we introduce RoboMME-Interference, a cross-session benchmark built on RoboMME (Dai et al., 2026). For each query episode, we construct a session history using the query's relevant prior demonstration followed by a controlled number of unrelated sessions, which we provide to the VLA as memory and measure accuracy. Running RoboMME's released memory-augmented π 0.5 variants unmodified through this benchmark, we find that while perceptual memory variants improve success when given the history without any distractors, they decay strongly and steadily as unrelated sessions accumulate. The subgoal variants, which read the history with a vision-language model and pass written subgoals to the policy, improve less at their best but hold more of that improvement as distractors accumulate. Adding a retrieval step to the strongest perceptual variant, which selects the section of history most visually similar to the robot's current view and passes only that section to the policy, restores its no-distractor success rate at every interference level. With this release, we emphasize the importance of long-context memory and robustness to interference and show that current systems largely fail on such capabilities. The project page, videos, code, and data are at https://robotmemorybench.com.
Sources
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- RMBench: Memory-Dependent Robotic Manipulation Benchmark with Insights into Policy Design
- RoboMME: Benchmarking and Understanding Memory for Robotic Generalist Policies
- Chameleon: Control-Indexed Prospective Memory for Visuomotor Manipulation
- OpenVLA: An Open-Source Vision-Language-Action Model
- RoboMemArena: A Comprehensive and Challenging Robotic Memory Benchmark
- MAP-VLA: Memory-Augmented Prompting for Vision-Language-Action Model in Robotic Manipulation
- Evaluating Very Long-Term Conversational Memory of LLM Agents
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
- MemER: Scaling Up Memory for Robot Control via Experience Retrieval
- Learning to (Learn at Test Time): RNNs with Expressive Hidden States
- Beyond a Million Tokens: Benchmarking and Enhancing Long-Term Memory in LLMs
- MEM: Multi-Scale Embodied Memory for Vision Language Action Models
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving