Divide-and-Remember: Recursive Action-Relevant Memory for Long-Horizon VLA Policies
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.
Dev: Today's paper: "Divide-and-Remember: Recursive Action-Relevant Memory for Long-Horizon VLA Policies".
Rosa: Vision–language–action (VLA) models struggle on history-dependent manipulation tasks where current observations alone do not determine action, necessitating a memory mechanism to retain relevant past information.
Dev: First, who's behind it and why it matters.
Title and authors: Rosa: Moving on to the title and authors of this paper, "Divide-and-Remember: Recursive Action-Relevant Memory for Long-Horizon VLA Policies." It sounds like they are focusing intensely on solving the problem of history dependence in Vision-Language-Action models when those tasks span many steps.
Dev: I see they’re tackling that core difficulty—where just looking at the current scene doesn't tell the AI what to do next because it needs context from earlier events. The authors are also listing a team of researchers including Xuehui Yu, Eason Yu, Meiyi Wang, and Haozhe Du.
Taro: I’m interested in what this specific focus on "Recursive Action-Relevant Memory" implies for autonomy research; does this mean they are targeting the kind of memory needed for long-term planning or just short-term context maintenance?
Rosa: It seems they are aiming for more than just short-term context; the paper points to history-dependent tasks like picking up an object and then moving it to a target in a specific manner, which requires recalling actions from earlier in the sequence.
Dev: That kind of task demands more than just tracking where things are now; it needs to recall *how* things were done previously, which is exactly where standard memory methods often fail because they pick what seems visually salient rather than action-relevant.
Taro: So, their core idea is that the optimal memory isn't just a snapshot of the environment; it has to be a distillation of the history that directly informs the next action. That sounds like a much more sophisticated form of contextual awareness for an autonomous system.
Rosa: Precisely; they view memory as an optimization problem where we select from past steps what carries the key fact that the current observation lacks, which is then used by the policy to make a correct move.
Dev: That distinction between general history and action-relevant memory is key because it helps filter out irrelevant historical data, which should be beneficial for keeping computational loads down in our control loops.
Taro: If this approach works well for remembering specific sequences or procedures, what kind of complex behavioral patterns are they imagining the AI being able to handle autonomously?
Rosa: They're looking at things like accumulating counts over repeated events or reproducing a demonstrated sequence, which points toward tasks that require procedural memory rather than just raw state tracking.
Dev: That sounds challenging for a VLA model because it means the memory needs to encode not just spatial facts but also temporal dependencies and sequential logic, which puts a heavy demand on what that selected subset of tokens can capture.
Taro: So, their ambition is for an AI that can exhibit more complex behaviors that rely on procedural knowledge, moving beyond simple reactive responses based only on the immediate input.
Rosa: That’s the direction they are pushing; they want an agent capable of performing tasks that require understanding and repeating complex maneuvers without needing to re-experience every single step.
Dev: It's a significant leap from what we see in memory-augmented VLAs today, which often struggle with losing fine spatial details while trying to compress history into something compact.
The paper's summary: Rosa: To summarize the paper, "Divide-and-Remember: Recursive Action-Relevant Memory for Long-Horizon VLA Policies," they are proposing a solution to the struggle of Vision-Language-Action models on tasks where current observations aren't sufficient.
Dev: Essentially, they show that the optimal way to handle this is by defining memory as an optimization problem where it must maximize the conditional mutual information between the action and that memory given what we see now.
Taro: That means they are not just picking what looks visually important; they are mathematically optimizing for preserving information that directly dictates future actions, which is a very different approach to memory design.
Rosa: Exactly; this optimal memory function acts as a policy-sufficient statistic, meaning the AI only needs that distilled piece of history to make the correct decision, simplifying its dependence on the entire history.
Dev: That simplification is what makes it viable for long sequences; instead of processing every single frame ever seen, the system focuses on a concise representation that is actually useful for control.
Taro: The mechanism they use to achieve this practical implementation is Divide-and-Remember, which recursively divides the history into smaller subproblems and selects relevant tokens at each level of recursion.
Rosa: So, D andR is their proposed architecture that allows them to map a function from the full history down to a subset of at most K tokens using this recursive selection technique.
Dev: And they train the selector function end-to-end by optimizing a lower bound of that mutual information objective, which links directly back to maximizing action utility in the policy's training.
Taro: This end-to-end learning of the selector is what makes it powerful because it learns *what* to remember automatically based on what actually helps the policy succeed, rather than relying on human intuition about pixel changes.
Rosa: It means they are teaching the AI how to distill its own experience into a highly compressed, action-relevant summary tailored for control, which is a very different way of learning memory.
Dev: That compression mechanism sounds like it addresses the tension between retaining necessary spatial details and keeping the representation computationally light enough for real-time operation.
The paper's improvements: Rosa: Now let’s discuss the specific architectural improvements they suggest in "Divide-and-Remember: Recursive Action-Relevant Memory for Long-Horizon VLA Policies." The main improvement is moving away from heuristic memory methods to this information-theoretic approach.
Dev: They propose implementing this recursive memory structure where the full history is divided into subproblems of top- K selection over 2K tokens, using a single, lightweight selector to manage an unbounded history efficiently.
Taro: That recursive division sounds like a clever way to manage complexity; having that single selector capture the common rule across all blocks should simplify the training and make it more scalable for very long contexts.
Rosa: And they train this memory function using the variational lower bound of the mutual information objective, which gives them a tractable way to optimize this complex objective through reinforcement learning.
Dev: That training signal derivation is really solid because it directly connects maximizing that theoretical measure to minimizing the actual action loss of the policy when it uses those selected tokens.
Taro: The core improvement is that they replace fixed selection rules, like using frames with largest pixel changes or frames a VLM labels as events, with a dynamic rule learned end-to-end based on what actually helps control.
Rosa: This means the AI learns to prioritize facts that are action-relevant over just visually salient ones, which should lead to better generalization across different manipulation tasks.
Dev: The paper also points out that D andR can achieve state-of-the-art performance metrics on RoboMME, specifically noting gains in frame-like facts, showing its success across various data types.
Taro: And I'm glad they highlighted that it performs well on real robots under noisy conditions and human intervention, which validates the method beyond just simulation results.
Conclusion: Rosa: So wrapping up the paper "Divide-and-Remember: Recursive Action-Relevant Memory for Long-Horizon VLA Policies," they’ve proposed a memory system that mathematically optimizes for conditional mutual information between action and memory given observation.
Dev: Essentially, they showed that this leads to a policy-sufficient statistic and introduces the Divide-and-Remember architecture to implement it recursively through top- K selection over 2K tokens.
Taro: The implication is that we can design memory not based on pre-defined rules but based on what the action actually needs, which should lead to more robust autonomy when things go wrong in complex scenarios.
Rosa: It suggests a significant shift toward learning memory as an optimization problem that maximizes control utility directly within the VLA model's training process.
Dev: It’s a method that promises better performance on long-horizon tasks, especially where we need to remember specific sequences or procedures rather than just static state tracking.
Taro: If this method proves reliable outside of the lab and handles noisy real-world data effectively, it could enable more capable physical agents in challenging physical situations.
Rosa: We've really seen how this paper introduces Divide-and-Remember as a structured way to learn memory that is directly tied to maximizing control performance.
Dev: It’s a method that offers better performance metrics compared to existing memory methods, especially when the context length grows quite large without sacrificing computational efficiency.
Taro: We should watch closely how this framework evolves in future work, especially concerning its ability to handle truly open-ended and unpredictable world interactions.
Xuehui Yu, Eason Yu, Meiyi Wang, Haozhe Du, Stefano V. Albrecht, Harold Soh
National University of Singapore · Nanyang Technological University, Singapore
cs.RO, cs.AI
Submitted: 2026-10-01
Updated: 2026-10-01
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 92/100
The gist: Vision–language–action (VLA) models struggle on history-dependent manipulation tasks where current observations alone do not determine action, necessitating a memory mechanism to retain relevant
Key concepts
- Optimal Memory Function
- The ideal memory for visuomotor policies is a function that maximizes the conditional mutual information between the action and the memory, given what is currently observed. This means the memory should store only the most relevant historical data needed to predict future actions.
- Policy-Sufficient Statistic
- Optimal memory is characterized as a policy-sufficient statistic, meaning that once you have this specific memory alongside the current observation, the policy's decision about what to do becomes independent of the entire history leading up to that point.
- Divide-and-Remember (D&R)
- D&R is an architecture designed to implement this optimal memory. It recursively divides a long history into manageable subproblems, using lightweight selectors trained end-to-end to keep only the most important tokens, ensuring scalability for very long contexts.
Terminology
Summary
Vision–language–action (VLA) models struggle on history-dependent manipulation tasks where current observations alone do not determine action, necessitating a memory mechanism to retain relevant past information. The core finding of this work is that the optimal memory function maximizes the conditional mutual information between the action and the memory given the current observation, leading to a proposed recursive method called Divide-and-Remember (D&R) that learns this memory within the pretrained token space of VLA models.
The gist
The optimal memory for visuomotor policies is defined as a policy-sufficient statistic that maximizes the conditional mutual information I(at; mt ot) between the action and the memory given the current observation.
Theoretical Foundation of Optimal Memory
The paper casts the memory problem within a Partially Observable Markov Decision Process (POMDP) formulation. The objective is to find a memory function, denoted as mt = M(ht), that maximizes I(at; mt ot). By applying the chain rule, this optimization is shown to be equivalent to minimizing the discarded information:
The optimal memory for visuomotor policies is M⋆ = arg max M I(at; M(ht) ot).
This optimal memory, mt = M⋆(ht), is characterized as a policy-sufficient statistic, meaning that the policy conditioned on the observation and this memory is independent of the full history:
The memory mt = M⋆(ht) is a policy-sufficient statistic, i.e., π⋆(at ot, ht) = π⋆(at ot, mt) for all (ot, ht).
Divide-and-Remember (D&R) Architecture
To implement this optimal memory in practice while staying compute-light and scalable to long contexts, the authors propose Divide-and-Remember (D&R). This method addresses the challenge of mapping a function M from the full history Ht to a subset mt with at most K tokens. D&R achieves this by dividing the full history recursively into subproblems:
-
The selection over the full history is divided recursively into subproblems of top-K selection over 2K tokens, where a fixed-size, lightweight selector learns end-to-end to support an unbounded history.
-
All recursion blocks share one selector, which captures the common selection rule and maintains efficiency.
-
The kept tokens from two neighboring slices are merged into a new 2K-token slice and re-selected level by level, with the final selections passed to the policy via an attention mask.
Learning the Selector Function
The memory function Mψ is parameterized by a selector fψ, which maps 2K candidate history tokens to their top-K selection. This selector is trained end-to-end using the variational lower bound of the mutual information objective:
Optimising a mutual-information objective through such a tractable bound has proven effective for generalisable control in reinforcement learning (Yu et al., 2024).
The selection score si for each token xi is calculated using an encoder (SigLIP) and a transformer structure, with the memory tokens mˆ being the top-K tokens selected based on scores gi. The training signal for fψ is derived from the action loss of the policy that reads these selected tokens, which corresponds to maximizing Equation 8:
max M I(at; mt ot) ≥ max M, θ ED[log πθ(at ot, mt)] + H(at ot).
Performance and Ablation Results
D&R was evaluated on RoboMME, a benchmark of 16 long-horizon manipulation tasks. Under a memory budget of only 64 tokens, D&R achieved a state-of-the-art average success rate of 38.6%, outperforming latent memory (32.2%) and the memoryless policy (17.9%). The method demonstrated superior gains across all four suites, including frame-like facts and state-like facts. Ablation studies showed that D&R gains the most on frame-like facts: 47.5% on Motion-Centric against 31.0% for FrameSamp and 24.8% for HAMLET.
Furthermore, real-robot experiments confirmed these gains under conditions of perception noise and human intervention. The method's efficiency is highlighted by its ability to maintain high success rates even as the candidate pool size grows, showing that it keeps its cost independent of the growing history length.
Real-World Validation
The effectiveness of D&R was validated on a real robot across four tasks: PutBottles (Counting), TrackCube (Permanence), RepickCube (Reference), and DrawPattern (Imitation).
Improvements for AI systems
Based on the provided research paper, here are specific improvements for AI systems and what those improved systems can achieve:
The core improvement lies in replacing existing, heuristic-based memory methods with a theoretically grounded, information-theoretic approach called Divide-and-Remember (D&R).
- Improve the VLA policy's memory mechanism by implementing the D&R method, which optimizes for maximizing the conditional mutual information:
I(at; mt ot) between action and memory given current observation.
-
Implement a recursive memory structure where the full history is divided into subproblems of top-K selection over 2K tokens, using a single, lightweight selector to manage an unbounded history efficiently. This ensures the selected memory remains compute-light while supporting long contexts (up to 64 tokens in the benchmark).
-
Enable end-to-end training where the memory function (the selector) is trained directly against a tractable lower bound of the mutual information objective:
max M I(at; mt ot) ≥ max M, θ ED[log πθ(at ot, mt)] + H(at ot).
-
The resulting improved AI system can perform complex, long-horizon manipulation tasks that require remembering specific historical facts (like
in the same manner as before,
ormove the cube to the target
) without running out of context length. -
The D&R system can achieve state-of-the-art performance (e.g., 38.6% on RoboMME) across diverse task suites, specifically excelling at remembering:
at a) Frame-like facts (e.g., object locations and spatial relations).
b) State-like facts (e.g., tracking a running state like elapsed time or a count of events).
-
The system will demonstrate superior generalization compared to existing methods because the D&R memory is optimized in the pretrained token space of the VLA, allowing it to retain high fidelity spatial details while remaining optimized for control, rather than relying on loose perceptual heuristics (like pixel change).
-
In real-world robot experiments, this improved system will show increased reliability under challenging conditions (perception noise and mid-task human intervention), as evidenced by achieving the highest success rates in tasks like
PutBottles
andTrackCube.
Sources
- Recursive Belief Vision Language Action Models
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- EventVLA: Event-Driven Visual Evidence Memory for Long-Horizon Vision-Language-Action Policies
- KEMO: Event-Driven Keyframe Memory for Long-Horizon Robot Manipulation with VLA Policies
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving