Divide-and-Remember: Recursive Action-Relevant Memory for Long-Horizon VLA Policies

summary

Video file (mp4)

The gist

Vision–language–action (VLA) models struggle on history-dependent manipulation tasks where current observations alone do not determine action, necessitating a memory mechanism to retain relevant

In short

Vision-language-action (VLA) models struggle with long tasks because they cannot remember past information effectively. This work proposes a memory function that maximizes mutual information between actions and memory given observations. The Divide-and-Remember (D&R) method implements this optimal memory recursively, allowing VLA models to retain relevant past context efficiently.

Key concepts

Optimal Memory Function
The ideal memory for visuomotor policies is a function that maximizes the conditional mutual information between the action and the memory, given what is currently observed. This means the memory should store only the most relevant historical data needed to predict future actions.
Policy-Sufficient Statistic
Optimal memory is characterized as a policy-sufficient statistic, meaning that once you have this specific memory alongside the current observation, the policy's decision about what to do becomes independent of the entire history leading up to that point.
Divide-and-Remember (D&R)
D&R is an architecture designed to implement this optimal memory. It recursively divides a long history into manageable subproblems, using lightweight selectors trained end-to-end to keep only the most important tokens, ensuring scalability for very long contexts.

Terminology used across episodes

This episode discusses

The paper

Divide-and-Remember: Recursive Action-Relevant Memory for Long-Horizon VLA Policies · Read on arXiv

Xuehui Yu, Eason Yu, Meiyi Wang, Haozhe Du, Stefano V. Albrecht, Harold Soh

National University of Singapore · Nanyang Technological University, Singapore

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.

Dev: Today's paper: "Divide-and-Remember: Recursive Action-Relevant Memory for Long-Horizon VLA Policies".

Rosa: Vision–language–action (VLA) models struggle on history-dependent manipulation tasks where current observations alone do not determine action, necessitating a memory mechanism to retain relevant past information.

Dev: First, who's behind it and why it matters.

Title and authors: Rosa: Moving on to the title and authors of this paper, "Divide-and-Remember: Recursive Action-Relevant Memory for Long-Horizon VLA Policies." It sounds like they are focusing intensely on solving the problem of history dependence in Vision-Language-Action models when those tasks span many steps.

Dev: I see they’re tackling that core difficulty—where just looking at the current scene doesn't tell the AI what to do next because it needs context from earlier events. The authors are also listing a team of researchers including Xuehui Yu, Eason Yu, Meiyi Wang, and Haozhe Du.

Taro: I’m interested in what this specific focus on "Recursive Action-Relevant Memory" implies for autonomy research; does this mean they are targeting the kind of memory needed for long-term planning or just short-term context maintenance?

Rosa: It seems they are aiming for more than just short-term context; the paper points to history-dependent tasks like picking up an object and then moving it to a target in a specific manner, which requires recalling actions from earlier in the sequence.

Dev: That kind of task demands more than just tracking where things are now; it needs to recall *how* things were done previously, which is exactly where standard memory methods often fail because they pick what seems visually salient rather than action-relevant.

Taro: So, their core idea is that the optimal memory isn't just a snapshot of the environment; it has to be a distillation of the history that directly informs the next action. That sounds like a much more sophisticated form of contextual awareness for an autonomous system.

Rosa: Precisely; they view memory as an optimization problem where we select from past steps what carries the key fact that the current observation lacks, which is then used by the policy to make a correct move.

Dev: That distinction between general history and action-relevant memory is key because it helps filter out irrelevant historical data, which should be beneficial for keeping computational loads down in our control loops.

Taro: If this approach works well for remembering specific sequences or procedures, what kind of complex behavioral patterns are they imagining the AI being able to handle autonomously?

Rosa: They're looking at things like accumulating counts over repeated events or reproducing a demonstrated sequence, which points toward tasks that require procedural memory rather than just raw state tracking.

Dev: That sounds challenging for a VLA model because it means the memory needs to encode not just spatial facts but also temporal dependencies and sequential logic, which puts a heavy demand on what that selected subset of tokens can capture.

Taro: So, their ambition is for an AI that can exhibit more complex behaviors that rely on procedural knowledge, moving beyond simple reactive responses based only on the immediate input.

Rosa: That’s the direction they are pushing; they want an agent capable of performing tasks that require understanding and repeating complex maneuvers without needing to re-experience every single step.

Dev: It's a significant leap from what we see in memory-augmented VLAs today, which often struggle with losing fine spatial details while trying to compress history into something compact.

The paper's summary: Rosa: To summarize the paper, "Divide-and-Remember: Recursive Action-Relevant Memory for Long-Horizon VLA Policies," they are proposing a solution to the struggle of Vision-Language-Action models on tasks where current observations aren't sufficient.

Dev: Essentially, they show that the optimal way to handle this is by defining memory as an optimization problem where it must maximize the conditional mutual information between the action and that memory given what we see now.

Taro: That means they are not just picking what looks visually important; they are mathematically optimizing for preserving information that directly dictates future actions, which is a very different approach to memory design.

Rosa: Exactly; this optimal memory function acts as a policy-sufficient statistic, meaning the AI only needs that distilled piece of history to make the correct decision, simplifying its dependence on the entire history.

Dev: That simplification is what makes it viable for long sequences; instead of processing every single frame ever seen, the system focuses on a concise representation that is actually useful for control.

Taro: The mechanism they use to achieve this practical implementation is Divide-and-Remember, which recursively divides the history into smaller subproblems and selects relevant tokens at each level of recursion.

Rosa: So, D andR is their proposed architecture that allows them to map a function from the full history down to a subset of at most K tokens using this recursive selection technique.

Dev: And they train the selector function end-to-end by optimizing a lower bound of that mutual information objective, which links directly back to maximizing action utility in the policy's training.

Taro: This end-to-end learning of the selector is what makes it powerful because it learns *what* to remember automatically based on what actually helps the policy succeed, rather than relying on human intuition about pixel changes.

Rosa: It means they are teaching the AI how to distill its own experience into a highly compressed, action-relevant summary tailored for control, which is a very different way of learning memory.

Dev: That compression mechanism sounds like it addresses the tension between retaining necessary spatial details and keeping the representation computationally light enough for real-time operation.

The paper's improvements: Rosa: Now let’s discuss the specific architectural improvements they suggest in "Divide-and-Remember: Recursive Action-Relevant Memory for Long-Horizon VLA Policies." The main improvement is moving away from heuristic memory methods to this information-theoretic approach.

Dev: They propose implementing this recursive memory structure where the full history is divided into subproblems of top- K selection over 2K tokens, using a single, lightweight selector to manage an unbounded history efficiently.

Taro: That recursive division sounds like a clever way to manage complexity; having that single selector capture the common rule across all blocks should simplify the training and make it more scalable for very long contexts.

Rosa: And they train this memory function using the variational lower bound of the mutual information objective, which gives them a tractable way to optimize this complex objective through reinforcement learning.

Dev: That training signal derivation is really solid because it directly connects maximizing that theoretical measure to minimizing the actual action loss of the policy when it uses those selected tokens.

Taro: The core improvement is that they replace fixed selection rules, like using frames with largest pixel changes or frames a VLM labels as events, with a dynamic rule learned end-to-end based on what actually helps control.

Rosa: This means the AI learns to prioritize facts that are action-relevant over just visually salient ones, which should lead to better generalization across different manipulation tasks.

Dev: The paper also points out that D andR can achieve state-of-the-art performance metrics on RoboMME, specifically noting gains in frame-like facts, showing its success across various data types.

Taro: And I'm glad they highlighted that it performs well on real robots under noisy conditions and human intervention, which validates the method beyond just simulation results.

Conclusion: Rosa: So wrapping up the paper "Divide-and-Remember: Recursive Action-Relevant Memory for Long-Horizon VLA Policies," they’ve proposed a memory system that mathematically optimizes for conditional mutual information between action and memory given observation.

Dev: Essentially, they showed that this leads to a policy-sufficient statistic and introduces the Divide-and-Remember architecture to implement it recursively through top- K selection over 2K tokens.

Taro: The implication is that we can design memory not based on pre-defined rules but based on what the action actually needs, which should lead to more robust autonomy when things go wrong in complex scenarios.

Rosa: It suggests a significant shift toward learning memory as an optimization problem that maximizes control utility directly within the VLA model's training process.

Dev: It’s a method that promises better performance on long-horizon tasks, especially where we need to remember specific sequences or procedures rather than just static state tracking.

Taro: If this method proves reliable outside of the lab and handles noisy real-world data effectively, it could enable more capable physical agents in challenging physical situations.

Rosa: We've really seen how this paper introduces Divide-and-Remember as a structured way to learn memory that is directly tied to maximizing control performance.

Dev: It’s a method that offers better performance metrics compared to existing memory methods, especially when the context length grows quite large without sacrificing computational efficiency.

Taro: We should watch closely how this framework evolves in future work, especially concerning its ability to handle truly open-ended and unpredictable world interactions.

More episodes

← Home