MIKASA-Robo-VLA: Benchmarking Memory in VLA Models for Long-Horizon Manipulation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "MIKASA-Robo-VLA: Benchmarking Memory in VLA Models for Long-Horizon Manipulation".
Jane: As a fastidious and diligent researcher, I have meticulously analyzed the provided excerpts from the paper "MIKASA-Robo-VLA:
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So, wrapping up our talk on "MIKASA-Robo-VLA: Benchmarking Memory in VLA Models for Long-Horizon Manipulation," what does this benchmark really mean for us moving forward? It seems like the authors have given us a very clear picture of the current limitations regarding how much memory these models can sustain during complex, lengthy operations.
Jane: I think it shows that if we want AI to handle tasks that require sustained interaction and planning over extended periods, we need to focus our research on improving the underlying memory mechanisms themselves rather than just increasing raw processing power for vision or language inputs.
Lu: The implication is that future work should heavily focus on designing better explicit memory systems or more sophisticated recurrent state management within the AI architecture to handle those long retention intervals effectively <ref:2610.00604#pg2>.
Meng: For implementation, this paper gives us a yardstick. Instead of just saying "our model is good," we can now say, "our model performs better on the temporal memory tasks defined in MIKASA-Robo-VLA." It provides a measurable way to improve our systems in a way that matters for real deployment <ref:2610.00604#pg0>.
Lalam: Culturally, this work pushes the entire field toward recognizing that persistence and memory are not just technical hurdles; they are fundamental requirements for true generalist AI capabilities interacting with the physical world <ref:2610.00604#pg1>.
Tom: Exactly. The title itself, "MIKASA-Robo-VLA: Benchmarking Memory in VLA Models for Long-Horizon Manipulation," perfectly captures the essence of this study—it’s a detailed look at testing memory limits in VLA systems during long tasks.
Jane: It’s a lot to digest because it forces us to think about the architecture of intelligence, not just the input processing layers, when we design these advanced AI systems. We're moving into an era where sustained context is key for real-world utility.
Conclusion: Tom: So we’ve spent some time looking at how MIKASA-Robo-VLA sets up this benchmark for testing memory in VLA models, and now we need to wrap up with some big takeaways.
Jane: Yeah, it really boils down to this paper's title, "MIKASA-Robo-VLA: Benchmarking Memory in VLA Models for Long-Horizon Manipulation," which is just super descriptive of what they’ve done.
Lu: It shows that they didn't just look at short tasks; they deliberately designed things that force the models to remember stuff over very long stretches, up to two thousand one hundred sixty steps.
Meng: From my side, the authors are testing how far we can push these systems before the memory starts failing in a real-world scenario.
Lalam: I think what this benchmark really suggests is that for AI to handle complex manipulation tasks over time, it needs a way to keep track of context much more robustly than current methods allow.
Tom: Exactly, and when you consider the authors of this work, they’ve put together a very structured set of tests that gives us a real measure of performance.
Jane: It’s important for everyone listening to grasp that MIKASA-Robo-VLA isn't just another test; it’s a standardized way to see how different VLA models handle the actual difficulty of remembering things during long operations.
Lu: The implication here is that we need to think about building memory architectures into the very core of these models so they can sustain that context reliably.
Meng: Practically, this means we can start setting clearer targets for how much temporal or spatial information an AI needs to retain before it starts making mistakes in a long sequence.
Lalam: For our culture and future applications, this pushes us toward designing AI systems that have more persistent working memory, which could lead to much more intuitive and capable interactions with the physical world.
Tom: It's clear this research is laying down the groundwork for what’s next in making AI truly capable of handling long-term physical tasks.
Egor Cherepanov, Nikita Kachaev, Aleksandr I. Panov, Alexey K. Kovalev
HSE University
cs.LG, cs.AI, cs.RO
Submitted: 2026-09-30
Updated: 2026-09-30
Comments: 57 pages, 39 figures, 38 tables
Code: https://github.com/CognitiveAISystems/MIKASA-Robo
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 89/100
The gist: As a fastidious and diligent researcher, I have meticulously analyzed the provided excerpts from the paper "MIKASA-Robo-VLA: Benchmarking Memory in VLA Models for Long-Horizon Manipulation." My
Key concepts
- Memory Types
- These are the different kinds of information a model must remember during a task. Examples include remembering the identity of an object, its exact location in space, or maintaining a sequence of actions. The benchmark tests if models can successfully use these specific memory types to solve complex problems.
- Long-Horizon Episodes
- These are extended sequences of actions that can last up to 2,160 steps. They are divided into short, medium, and long categories. This structure allows researchers to see how well models maintain context and plan effectively over very long time scales.
- Information Gap Metric
- This metric measures the time delay between when a necessary piece of information is last seen and when the model needs to use it. A large gap indicates a challenging situation where the model must rely heavily on its stored memory to proceed successfully.
Terminology
Summary
As a fastidious and diligent researcher, I have meticulously analyzed the provided excerpts from the paper MIKASA-Robo-VLA: Benchmarking Memory in VLA Models for Long-Horizon Manipulation.
My objective is to synthesize this information into a comprehensive, detailed summary suitable for rigorous academic review.
Here is the detailed synthesis:
The paper introduces MIKASA-Robo-VLA, a sophisticated and memory-intensive benchmark designed specifically to evaluate the capabilities of vision-language-action (VLA) models in long-horizon manipulation tasks. The core contribution of this work is not a single performance metric, but rather the creation of a standardized suite—comprising the benchmark itself, the comprehensive dataset, and a diagnostic apparatus—to rigorously test how effectively these models can utilize various forms of memory over extended temporal horizons.
A. Task Complexity and Diversity:
MIKASA-Robo-VLA comprises 90 language-conditioned manipulation tasks distributed across 26 distinct task families. These tasks are meticulously designed to probe different aspects of required memory, varying in their content (e.g., object attributes, sequences, counts) and structural complexity.
B. Memory Requirements:
A critical feature of the benchmark is that every memory-dependent task requires retention by construction. The suite explicitly covers ten distinct memory types:
-
Object: Retaining identity or location of an object.
-
Spatial: Recalling a specific physical location within the environment.
-
Capacity: Maintaining set membership (e.g., remembering which items are in a collection).
-
Temporal: Retaining counts or durations over time.
-
Negative: Defining a goal through exclusion (what not to do or what is absent).
-
Sequential: Remembering the order of items within a sequence.
-
Procedural: Recalling a prior trajectory or motion plan.
-
Prospective and Tracking: These types further test future planning and the maintenance of moving referents, respectively.
C. Episode Structure and Horizon Regimes:
The benchmark spans diverse temporal scales, with episode horizons ranging from 25 to 2,160 steps. These episodes are categorized into three distinct splits: Short (25–200 steps), Medium (250–600 steps), and Long (700–2,160 steps).
D. Data Scale and Format:
The benchmark releases a substantial dataset comprising 22,500 oracle demonstration episodes, totaling approximately 6.39 million timesteps. This data is provided in two formats:
-
RLDS: Storing dense per-step rewards suitable for online Reinforcement Learning (RL) fine-tuning.
-
LeRobotDataset v3: Storing the raw state and action arrays used for training and evaluation.
Crucially, every released episode is guaranteed to satisfy an oracle success rate of 1.000, meaning the demonstrations are perfectly executed by oracle policies or scripted planners, ensuring a ground truth for evaluation.
The environment interface is standardized to ensure fair comparison across models:
-
Visual Observations: The robot receives two visual inputs: a static top-down RGB image and a wrist-camera RGB image, both 128 times 128 times 3.
-
Proprioception: A 7-dimensional vector is provided, detailing the end-effector position, roll–pitch–yaw orientation, and gripper state.
-
Actions: The robot executes 7 dimensions of action, controlling delta end-effector pose and gripper control at a frequency of 10 Hz.
-
Language Input: Language instructions explicitly define the goal and specify the required memory operation for the task.
The benchmark goes beyond mere success rates by providing robust diagnostic tools:
-
Information Gap Metric: This metric characterizes environment-level information gaps by measuring the interval between when a task-relevant cue is last available and the start of the phase demanding a cue-dependent action. For 70 tasks, these timings are explicitly specified; for 28 tasks, this gap exceeds 16 steps, establishing this as the widest fixed-context VLA window surveyed.
-
Chance Floor: A baseline success rate derived from running a uniform-random policy provides a lower bound against which learned policies can be measured.
A reference policy baseline, pi 0.
Improvements for AI systems
Here are the specific improvements to AI systems that can be derived from MIKASA-Robo-VLA, and what these improved systems can achieve:
The core improvement lies in moving VLA policies beyond short-term visual attention toward robust, long-horizon reasoning by explicitly modeling and managing memory constraints.
-
The ability to characterize and manage different types of memory demands (Object, Spatial, Capacity, Temporal, Negative, Sequential/Procedural) allows for the development of a
Memory-Aware Policy Architecture.
-
The introduction of an explicit information gap metric (the interval between cue disappearance and action requirement) enables policies to be trained or evaluated based on their ability to predict or compensate for expected memory depletion.
Improvement Specifics:
-
A VLA policy can be explicitly designed with a dedicated, parameterized memory module that mimics the different types explored in MIKASA-Robo (e.g., a
Sequential Memory Buffer
for tracking orders, or aSpatial Map
for recalling occluded locations). -
Policies can be trained using memory-specific loss functions derived from the task family they are solving (e.g., if the task is
RememberColor,
the loss function prioritizes color identity retention during occlusion). -
The system can incorporate a dynamic
Memory Load
estimator that tracks which memory types are currently being utilized, allowing for resource allocation in complex multi-stage tasks.
What the Improved AI System Can Do:
-
A robot can successfully complete long-horizon manipulation tasks (up to 2,160 steps) where critical information is temporarily occluded or requires recall of a sequence/order of events.
-
The system will demonstrate superior performance in tasks requiring complex memory operations, such as:
Small-Scale Memory Retrieval: Recalling the identity (color or shape) of an object after it has been hidden for extended periods (e.g., Observe which color is under each cup, track them as they shuffle, then touch the cup matching the lamp color
).
Long-Sequence Tracking: Maintaining and reproducing a specific order of elements through a series of actions, even when intermediate states are obscured (e.g., Observe colors in sequence, then touch them in that same order
).
Temporal Reasoning: Accurately counting or tracking time-dependent events (e.g., Count how many times the blue lamp blinks and press the red button exactly that many times
).
State Prediction Under Uncertainty: In scenarios where cues are absent for long durations, the system can leverage learned memory structures to predict where an object is located or what action is required next, even if it cannot perfectly retrieve a static fact.
Adaptive Task Execution: The policy can dynamically switch its internal memory strategy
based on the current task demands—shifting from pure visual tracking (Spatial/Tracking) to explicit sequence recall (Sequential) when the instruction changes.
Abstract
Vision-language-action policies often see only one or a few recent frames, which makes it difficult to evaluate how they use information that disappears during a task. We introduce MIKASA-Robo-VLA, a benchmark of 90 language-conditioned manipulation tasks. All but 10 hide the cue an action depends on. Those 10 are reactive controls. MIKASA-Robo, the suite it rebuilds, has 32 tasks and uses language only in a representative VLA subset. Here every task provides an instruction, while memory-dependent tasks hide a task-relevant cue and reactive controls keep it available. For 70 tasks, environment phase timings specify an information gap, and for 28 of them the gap exceeds the 16-frame window of the widest fixed-context VLA we survey. The gap counts only the interval the cue is provably absent, not the full duration a policy must retain it, so every memory-dependent task still requires memory by construction, including the ones whose measured gap is short. We release 22,500 oracle trajectories across 10 memory types in RLDS and LeRobotDataset v3. A reference π 0.5 baseline with current images and proprioception, but no observation history or explicit memory module, is fine-tuned on 14 tasks and achieves 0.211 plus or minus 0.044 mean task success. Its lower success on the evaluated Long-split tasks is confounded by open-loop chunking and the memory types represented in that subset. Project page: https://mikasarobo.github.io/
Sources
- MemER: Scaling Up Memory for Robot Control via Experience Retrieval
- MEM: Multi-Scale Embodied Memory for Vision Language Action Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks