ElasticMem: Latent Memory as a Learnable Resource for LLM Agents
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "ElasticMem: Latent Memory as a Learnable Resource for LLM Agents".
Tom: Based on the provided text snippets, I have meticulously synthesized a detailed description of the paper's core concepts, methodology, and empirical results.
Jane: First, who's behind it and why it matters.
Title and authors: Tom: So we're looking at ElasticMem today. The title is "ElasticMem: Latent Memory as a Learnable Resource for LLM Agents." It sounds like they are moving away from just dumping all the memory into the context window, which is what some models do.
Jane: That’s right. They are proposing that memory shouldn't be a fixed thing you have to manage; it should be something the AI learns how to use flexibly. Think of it less like a filing cabinet and more like a dynamic resource pool that changes based on what the AI needs right now.
Lu: It suggests they’re tackling the problem where current memory methods are either too expensive in terms of tokens or they force you into using rigid retrieval systems, which can be kind of wasteful if you don't know which memory is actually useful for a specific task.
Meng: So instead of just grabbing the top few similar things, this framework tries to figure out exactly how much "space" each piece of memory gets based on the current question. That sounds like it could save a lot on computation time if it’s smart about what it keeps active in its head.
Lalam: I think they are treating memory as an elastic latent resource, which means the model learns to prioritize useful evidence by assigning each retrieved memory a variable latent budget through a learned policy. That sounds like a big step toward making agents truly adaptive instead of just being static retrievers.
The paper's summary: Tom: Okay, so what does ElasticMem actually do? They build an offline latent memory bank with retrieval keys and content caches first, and then when the agent needs information, it queries the reasoner’s hidden state to retrieve memories adaptively.
Jane: It’s not just about searching; it’s about using a learned policy to decide which memories are relevant for the current context before they even get injected into the model. It's that adaptive retrieval part you mentioned earlier.
Lu: The core idea is assigning each retrieved memory a variable latent budget governed by a learned policy that looks at two things: how relevant the query is to that specific memory, and what utility you expect from using it downstream.
Meng: So it’s trying to solve the problem of mismatch between query-dependent memory utility and fixed allocation. They are saying you don't need a uniform capacity for every piece of retrieved data because some memories might be way more important than others for the next step.
Lalam: And they inject these selected latent states as soft memory tokens, scaled by their allocated budget, which lets the model selectively incorporate relevant context without getting bogged down by irrelevant or redundant data. That’s a neat way to control the flow of information.
The paper's improvements: Tom: The authors point out that they are optimizing the whole pipeline end-to-end using group-relative policy optimization, which means they're not just focusing on getting good retrieval or budget allocation in isolation.
Jane: They are supervising everything from the retrieval control token to the memory budget decisions and even the latent projector itself through that optimization process, making it a holistic system.
Lu: The biggest theoretical point they make is that ElasticMem learns to distinguish useful evidence from memories that look superficially similar, which is better than just using simple similarity metrics like cosine similarity alone.
Meng: They also show it aligns memory budgets with transferable plan structures instead of arbitrary task labels, which means the system understands the *structure* of a plan rather than just a label on a piece of text.
Lalam: So they are trainable components include the LoRA-adapted reasoner backbone, that budget policy network, and that latent projector. That shows they are designing memory use as an actual learnable resource allocation problem for the AI.
Conclusion: Tom: So to wrap up, ElasticMem is about treating memory as a dynamic pool where the system learns how to prioritize and allocate capacity based on what it needs next, which is really important for long-horizon tasks.
Jane: It moves beyond just finding similar chunks; it teaches the AI *how much* weight to give each chunk when it’s using that information. This selective use is what allows for better performance without just piling on more data or making the memory bigger in a fixed way.
Lu: The result is that they achieve high accuracy on benchmarks like ALFWorld, improving success rates significantly, while also showing the lowest token cost among all compared methods at that level of performance.
Meng: For practical application, this suggests we can build agents that use their memory much more efficiently in real-world scenarios because they aren't wasting tokens on things that aren't actually helpful for the immediate goal.
Lalam: I think ElasticMem is a promising direction for building more efficient, adaptive, and long-horizon LLM agents by transforming latent memory into a dynamically managed, learnable resource.
Tom: That’s what we have on ElasticMem today. It really shows how learning to manage resources can unlock better performance for these complex AI systems.
Tao Feng, Chongrui Ye, Tianyang Luo, Jingjun Xu, Xueqiang Xu, Haozhen Zhang, Ge Liu, Jiaxuan You
University of Illinois Urbana-Champaign
cs.CL
Submitted: 2026-05-29
Updated: 2026-10-04
Code: https://github.com/ulab-uiuc/ElasticMem
Project page: https://langchain-ai.github.io/langmem
Importance score: 91/100
The gist: Based on the provided text snippets, I have meticulously synthesized a detailed description of the paper's core concepts, methodology, and empirical results.
Key concepts
- Elastic Latent Resource
- Memory is treated as a dynamic resource that the model learns to use adaptively. Instead of having a static pool, the system learns how much capacity each piece of retrieved memory should occupy based on current needs and expected utility.
- Variable Latent Budget Allocation
- This core mechanism dynamically assigns a specific amount of latent space to each retrieved memory. The model learns this allocation by considering both how relevant the query is to that specific memory and how useful using that information will be for the final task.
- Soft Memory Injection
- Instead of rigidly inserting full memories, ElasticMem scales the selected latent states according to their allocated budget and injects them as 'soft tokens.' This allows the model to selectively incorporate relevant context without being cluttered by irrelevant data.
- Group-Relative Policy Optimization
- The entire memory pipeline—from retrieving memories to allocating budgets and generating output—is trained together. This joint optimization ensures that the system learns the best sequence of memory interactions tailored to achieve specific task rewards.
Terminology
Summary
Based on the provided text snippets, I have meticulously synthesized a detailed description of the paper's core concepts, methodology, and empirical results.
Here is the comprehensive summary:
Detailed Summary of ElasticMem: Latent Memory as a Learnable Resource for LLM Agents
The paper introduces ElasticMem, a novel memory-augmented Large Language Model (LLM) framework designed to treat memory not as a static, fixed resource, but as an elastic latent resource that the model learns to utilize adaptively. The central thesis of ElasticMem is that effective long-horizon reasoning and agentic decision-making require not just having access to memory, but intelligently managing which memories are retrieved and how much latent capacity each retrieved piece of information should occupy.
Core Methodology and Architecture:
ElasticMem operates through a sophisticated, multi-stage process that jointly optimizes several interconnected components:
-
Offline Memory Bank Construction: The framework first builds an offline latent memory bank. This bank is structured to facilitate retrieval, incorporating both retrieval keys and content caches.
-
Adaptive Retrieval: During inference or reasoning, ElasticMem retrieves memories adaptively by querying the reasoner's hidden state. Crucially, this retrieval process is not rigid; it is guided by a learned policy that determines which memories are relevant for the current context.
-
Variable Latent Budget Allocation (The Elastic Core): This is the defining feature of ElasticMem. After retrieving a set of candidate memories, the framework assigns each retrieved memory a variable latent budget. This allocation is governed by a learned policy that dynamically adjusts the capacity allocated to each memory based on two key factors:
-
The relevance of the query to that specific memory.
-
The downstream utility expected from using that specific piece of information.
-
Soft Memory Injection: The selected latent states, scaled according to their assigned budget, are then injected into the generation process as soft memory tokens. This mechanism allows the model to selectively incorporate relevant contextual information without being overwhelmed by irrelevant or redundant data.
-
Group-Relative Policy Optimization: The entire memory-use pipeline—encompassing retrieval control, budget allocation, latent projection, and final generation—is jointly optimized end-to-end using group-relative policy optimization. This means the system is trained not just to retrieve memories well or allocate budgets efficiently in isolation, but to optimize the entire sequence of memory interaction based on downstream task rewards.
Key Contributions and Theoretical Insights:
ElasticMem moves beyond traditional memory approaches by framing memory use as a learnable resource allocation problem. The trainable components of the framework are specifically:
-
The LoRA-adapted reasoner backbone.
-
The budget policy network responsible for variable allocation.
-
The latent projector that handles the injection of soft tokens.
A significant theoretical insight is that ElasticMem learns to distinguish useful evidence from superficially similar memories, overcoming limitations inherent in methods relying solely on rigid similarity metrics like cosine similarity. Furthermore, the framework demonstrates a superior alignment between memory usage and task structure: it aligns memory budgets with transferable plan structures rather than arbitrary task labels.
Empirical Performance and Superiority:
The empirical evaluation across two demanding benchmarks—MemorySuite (for memory-intensive QA) and ALFWorld (for embodied agentic decision-making)—demonstrates ElasticMem's exceptional performance:
-
MemorySuite-QA: ElasticMem achieves the highest weighted average accuracy, showing substantial improvements over strong baselines. For Qwen2.5-3B-Instruct, accuracy rose from 0.588 to 0.742, and for Qwen2.5-7B-Instruct, it improved from 0.668 to 0.832.
-
ALFWorld: The framework achieves the highest average success rate, improving significantly for both model sizes: from 0.270 to 0.449 for Qwen2.5-3B-Instruct, and from 0.416 to 0.529 for Qwen2.5-7B-Instruct.
The paper explicitly states that ElasticMem provides a strong accuracy-efficiency trade-off, outperforming both text-space and fixed-capacity latent methods by demonstrating gains derived from selective memory use rather than simply increasing interaction length or memory capacity.
**In conclusion, ElasticMem is presented as a promising direction for building more efficient, adaptive, and long-horizon LLM agents by transforming latent memory into a dynamically managed, learnable resource.
Improvements for AI systems
-
Elastic latent memory provides an
elastic latent resource
that learns to prioritize useful evidence, allowing it toassign each retrieved memory a variable latent budget through a learned policy.
This enables the system to distinguish betweenuseful or evidence-bearing memories
and those that areredundant or misleading,
which improves accuracy by learning not only what to answer but alsowhich memories to retrieve and how much representational capacity each memory deserves.
-
Query-conditioned retrieval adapts memory access based on internal reasoning, as the system derives the query from
the reasoner’s own hidden state rather than a standalone retrieval encoder,
leading to more effective selection compared to rigid methods. This adaptive retrieval is achieved by sampling aretrieval-control token
and using the resulting state, such thatretrieval is conditioned on how the model interprets the current input.
-
Group-relative policy optimization jointly optimizes all memory components, as it supervises
the retrieval-control decision, the memory budget allocations, the latent projector, and the LoRA-adapted reasoner through group-relative policy optimization,
ensuring thatthe reward signal supervises not only final generation but also the retrieval-control token and the memory-budget decisions.
Sources
- Multi-Agent Evolve: LLM Self-Improve through Co-evolution
- Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
- SePer: Measure Retrieval Utility Through The Lens Of Semantic Perplexity Reduction
- Memory-T1: Reinforcement Learning for Temporal Reasoning in Multi-session Agents
- LightMem: Lightweight and Efficient Memory-Augmented Generation
- Memp: Exploring Agent Procedural Memory
- Similarity is Not All You Need: Endowing Retrieval Augmented Generation with Multi Layered Thoughts
- SmartRAG: Jointly Learn RAG-Related Tasks From the Environment Feedback
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- PersonaMem-v2: Towards Personalized Intelligence via Learning Implicit User Personas and Agentic Memory
- Adaptation of Agentic AI: A Survey of Post-Training, Memory, and Skills
- Training Language Models to Self-Correct via Reinforcement Learning
- DeepSeek-V3 Technical Report
- Understanding R1-Zero-Like Training: A Critical Perspective
- Decoupled Weight Decay Regularization
- Skill-Pro: Learning Reusable Skills from Experience via Non-Parametric PPO for LLM Agents
- On Memory Construction and Retrieval for Personalized Conversational Agents
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering