Memory as a Controlled Process: Learned Adaptive Memory Management for LLM Agents
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Memory as a Controlled Process".
Tom: Large Language Model (LLM) agents increasingly rely on external memory systems to accumulate experience across tasks, yet nearly all existing approaches access memory through fixed, hand-designed heuristics.
Jane: First, who's behind it and why it matters.
Paper summary: Tom: So we're looking at this paper today, "Memory as a Controlled Process: Learned Adaptive Memory Management for LLM Agents." Basically, they're arguing that the way AI agents handle memory right now is too rigid.
Jane: They claim that existing methods use fixed rules for accessing memory, and that this static view is actually holding back how well these agents can learn across different tasks.
Lu: The core idea they push is that optimal memory behavior really depends on the context of the task at any given moment, so you need a way to adapt your retrieval strategy.
Meng: I mean, if we treat memory access like a set of fixed steps instead of something flexible, it seems like we're missing out on a lot of potential performance gains.
Lalam: From my side, this means the agent doesn't just blindly pull whatever is nearest; it learns when to pull and how much to pull based on what it needs for the current goal.
Tom: Exactly, and this paper introduces MEMCON as a framework that models these memory operations as a Markov Decision Process, or MDP.
Jane: So they're not just looking at memory as data storage anymore; they are treating every memory move—retrieving something, planning an injection—as an action in some kind of decision problem.
Lu: They define the action space pretty broadly, including things like retrieve, plan inject, re-retrieve, consolidate, forget or even do nothing.
Meng: That’s a lot of choices to make for the AI to decide on every single step it takes during a complex task.
Lalam: And they learn a policy online using something called a tabular contextual bandit with UCB exploration to figure out the best action for that specific situation.
Tom: So, what's the big picture here? The abstract says they show that early stages need less retrieval because memory is sparse, and recurring goals benefit from plan reuse instead of just looking up things one by one.
Jane: They also mention that when agents get stuck, they should use re-retrieval with different queries to try and find an alternative path forward.
Paper summary: Lu: And for long tasks, the authors argue the memory store itself needs to be consolidated and pruned so it stays useful over time.
Meng: That brings us to the practical side: how does this translate into real-world savings? The results they show are quite compelling when you look at token consumption across six different benchmarks and three agent frameworks.
Lalam: They report that MEMCON consistently beats existing memory baselines by up to fifteen point two points in task success, while also cutting down on token usage by about five to twenty percent across those tests.
Tom: That's a solid number for efficiency, and it shows that this adaptive control layer is actually helping the agents be smarter without necessarily making them use more massive language models for every single memory operation.
Jane: The theoretical underpinning of this work is quite interesting, as they show that the Memory MDP can be broken down into a family of per-state stochastic bandits.
Lu: That decomposition leads to some strong mathematical guarantees, specifically an O(log n) per-state regret and an O(ΦA log T) global regret.
Meng: So, theoretically, even though the AI is learning dynamically during deployment using that UCB rule, there are predictable bounds on how much it will be wrong over a long run.
Lalam: That means we can have a pretty good idea of how much effort it's going to take for the agent to get better as it interacts with the environment.
Tom: And what does this mean for those of us who just watch the AI from afar? It suggests that instead of designing one monolithic memory system, we should design a system that allows for this adaptive control layer.
Jane: If you only listen to the show, you can think about how an agent might decide to stop looking up old information and start using a pre-made plan instead.
Lu: The paper also shows they built this as a thin wrapper, making it backend agnostic so it can sit on top of any existing memory system.
Meng: That backend agnosticism is huge for us because we don't have to rewrite the entire data storage layer just to get this adaptive decision-making on top of it.
Paper summary: Lalam: And they introduced two special operations, plan injection and goal decomposition, which are tailored specifically for long-horizon agentic tasks.
Tom: Those augmentations prepending a general success plan when one is available, and handling composite tasks by injecting templates for sequential steps.
Jane: So the paper isn't just proposing a new way to store data; it’s proposing a smarter way to manage the flow of information within that storage.
Lu: This work builds on several areas, including multi-agent reinforcement learning and agentic memory research, positioning this adaptive control layer within that broader context.
Meng: From an engineering standpoint, I'm focused on how lightweight this controller is compared to using another full LLM call for every single memory query.
Lalam: They show that MEMCON achieves its adaptive behavior without needing those extra LLM calls, which is a significant practical win for running these agents efficiently.
Tom: So when we look at the conclusion of "Memory as a Controlled Process: Learned Adaptive Memory Management for LLM Agents," it really boils down to this adaptive approach.
Jane: The authors are essentially arguing that static memory systems are a bottleneck because they don't account for the changing needs of an agent during its life cycle.
Lu: They call their framework MEMCON, which models the operations as a Markov Decision Process, and they claim it learns an online policy to decide when to retrieve or consolidate information.
Meng: It’s about moving from fixed heuristics to a system where the memory management itself becomes part of the learning process.
Lalam: This means agents can learn how to manage their own experience in a way that is tailored specifically for the task they are currently performing, rather than relying on one-size-fits-all settings.
Tom: It really shifts the focus from just building bigger memory stores to building smarter systems that actively control how that memory is used.
Jane: The implication for us, as listeners, is seeing agents that can handle complex tasks more reliably because they are making intelligent choices about their own experience base.
Conclusion: Tom: So we've been talking about how agents need to manage their memory better than just keeping everything in one big dump of data.
Jane: Yeah, this paper, "Memory as a Controlled Process: Learned Adaptive Memory Management for LLM Agents," it basically says that the way an agent decides what to remember or forget should be a learning process itself.
Lu: They model memory operations like a decision-making problem, an MDP, where the AI learns an online policy to pick actions like retrieve or consolidate.
Meng: It’s about giving the agent control over its own experience—not just letting it blindly pull what it finds in its long-term storage.
Lalam: I think this means agents get smarter at deciding which memories are worth keeping and which ones to discard based on how they're actually doing the task.
Tom: So, the authors aren't just proposing a new way to store data; they’re building a system where memory management becomes part of the agent’s learning process.
Jane: That shift from fixed rules to an adaptive policy is really what makes this work different from previous methods we've seen.
Lu: They showed that by treating memory access as a series of choices, you can learn how to handle things like plan injection or re-retrieval dynamically.
Meng: From an engineering standpoint, it’s interesting because they wrap this up so that it works on top of whatever existing memory backend you already have installed.
Lalam: And the results are pretty solid; they show real improvements in task success while cutting down on the amount of text the agent has to process overall.
Tom: It suggests that agents can get better at handling complex, long-term goals by making intelligent choices about their own memory base throughout the entire run.
Jane: And this implies we’re moving toward agents that aren't just reactive tools, but systems that actively manage their own knowledge base for better performance.
Lu: It opens up a lot of possibilities for how we design agent architectures going forward if we want them to handle really deep, multi-step tasks reliably.
Meng: I’m curious about the practical side; it shows that this adaptive control layer is relatively light on resources compared to running another full language model every time the agent needs to decide what memory operation to do.
Lalam: And if we can make agents manage their own experience more effectively, it could lead to cultures where AI systems are much more self-correcting and less reliant on massive upfront training data for every single new situation.
Eric Hanchen Jiang, Zhi Zhang, Yuchen Wu, Levina Li, Dong Liu, Xiao Liang, Rui Sun, Yubei Li, Edward Sun
University of California Los Angeles
cs.CL, cs.AI
Submitted: 2026-07-15
Updated: 2026-10-05
Code: https://github.com/ericjiang18/MemCon
Importance score: 91/100
The gist: Large Language Model (LLM) agents increasingly rely on external memory systems to accumulate experience across tasks, yet nearly all existing approaches access memory through fixed, hand-designed
Key concepts
- Memory MDP
- This framework treats every memory action—like retrieving data or consolidating it—as a decision within a Markov Decision Process. The system learns the best sequence of these decisions to maximize long-term task success, much like an agent navigating a complex environment.
- Online Policy Learning
- MEMCON uses a lightweight contextual bandit with UCB exploration to learn the optimal memory strategy while the agent is actively working. This means it continuously adapts its memory retrieval and management choices based on real-time task progress and feedback, rather than following fixed rules.
- Augmented Operations
- The framework introduces specialized actions like PLANINJECT, which injects a generalized success plan when available. This allows the agent to proactively use high-level strategies for long tasks, and goal decomposition helps break down complex goals into manageable sequential steps.
Terminology
Summary
Large Language Model (LLM) agents increasingly rely on external memory systems to accumulate experience across tasks, yet nearly all existing approaches access memory through fixed, hand-designed heuristics. This paper introduces MEMCON, a framework that models memory operations as a Markov Decision Process and learns an online policy to adaptively decide when, what, and how much to retrieve or consolidate.
The gist
MEMCON is a framework that models memory operations as a Markov Decision Process and learns an online policy that adaptively decides when, what, and how much to retrieve, when to inject a distilled plan, and when to consolidate or forget.
How it works
MEMCON reformulates memory access as a sequential decision problem by casting the choice of memory operation (RETRIEVE, PLANINJECT, RE-RETRIEVE, CONSOLIDATE, FORGET, NOOP) together with its parameters (top k, insight k, graph hop) as actions in a Memory MDP. The state captures both task progress (goal type, step phase, stuck indicator, locations visited) and memory status (size, plan availability, learning phase). The policy is learned online during deployment using a lightweight tabular contextual bandit with UCB exploration.
Decision Making and Learning
The action space A is defined by an operation (op) and its parameters (θ), where op ∈ 9 actions, including RETRIEVE, PLANINJECT, RE-RETRIEVE, CONSOLIDATE, FORGET, NOOP. The policy selects an action via the Upper Confidence Bound rule: a = arg max a∈A
Q(ϕ(st), a) + c ln N(ϕ(st)) / N! ϕ s". The state S is discretized into a compact hashable key ϕ(s) for tabular learning. Credit assignment is performed via reverse-discounted Monte-Carlo return, updating Q-values using the rule: Q(ϕj, aj) ← Q(ϕj, aj) + α h γ ep−j−1 · r i.
Augmented Operations and Backend Agnosticism
MEMCON is a thin wrapper that intercepts the abstract retrieve and store entry points of any existing memory backend, making it backend-agnostic
by wrapping any existing memory implementation. It introduces two augmented operations tailored for long-horizon agentic tasks: generalized plan injection (PLANINJECT) and goal decomposition. PLANINJECT prepends a generalized success plan when one is available, and goal decomposition handles composite tasks by injecting templates for sequential steps. These augmentations consume only the action transcript and concatenate text to the retrieved context, making them independent of the backend.
Performance and Theoretical Guarantees
MEMCON consistently outperforms multiple memory baselines by up to 15.2 points in task success while reducing token consumption by 5–20% across 6 benchmarks, 3 agent frameworks, and 3 LLM backbones. Theoretically, the Memory MDP decomposes into a family of per-state stochastic bandits. The UCB rule admits a sub-Gaussian concentration inequality yielding an O(log n) per-state regret and an O(ΦA log T) global regret. This demonstrates that the average per-decision regret E[RT]/T → 0 as T → ∞ at rate O(log T /T).
REFERENCES
[1] Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. ReAct: Synergizing reasoning and acting in language models. In International Conference on Learning Representations (ICLR), 2023.
[2] Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan. Tree of thoughts: Deliberate problem solving with large language models. In Advances in Neural Information Processing Systems (NeurIPS), 2023.
[3] Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, Wayne Xin Zhao, Zhewei Wei and Jirong Wen. A survey on large language model based autonomous agents. Frontiers of Computer Science. 2024.
[4] Zhiheng Xi, Wenxiang Chen, Xin Guo, Wei He, Yiwen Ding, Boyang Hong, Ming Zhang, Junzhe Wang, Senjie Jin, Enyu Zhou et al. The rise and potential of large language model based agents: A survey. arXiv preprint arXiv:2309.07864. 2023.
[5] Joon Sung Park, Joseph C. O’Brien, Carrie J Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. Generative agents: Interactive simulacra of human behavior. In ACM Symposium on User Interface Software and Technology (UIST), 2023.
[6] Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. Reflexion: Language agents with verbal reinforcement learning. In Advances in Neural Information Processing Systems (NeurIPS), 2023.
[7] Andrew Zhao, Daniel Huang, Quentin Xu, Matthieu Lin, Yong-Jin Liu, and Gao Huang. ExpeL: LLM agents are experiential learners. 2024.
[8] Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar. Voyager: An open-ended embodied agent with large language models. 2024.
[9] Guibin Zhang, Muxin Yue, Xiangguo Li, Jiaxuan Ran, Rui Song, Ran Cheng, Zheng Wang, and Shirui Pan. G-Memory: Tracing hierarchical memory for multi-agent systems. arXiv preprint arXiv:2506.07398. 2025.
[10] Haoran Ou, Jianyu Li, Weiran Chen, Yuxin Liu, Ting Sun, and Dian Yu. Latent memory: Distilling cross-task experience into learnable tokens for language agents. arXiv preprint arXiv:2509.18432. 2025.
[11] Charles Packer, Sarah Wooders, Kevin Lin, Vivian Fang, Shishir G Patil, Ion Stoica, and Joseph E Gonzalez. MemGPT: Towards llms as operating systems. arXiv preprint arXiv:2310.08560. 2023.
[12] Prateek Chhikara, Deshraj Khant, Saket Aryan, Taranjeet Singh, and Deepak Yadav. Mem0: Building production-ready ai agents with scalable long-term memory. arXiv preprint arXiv:2504.19413. 2025.
[13] Bingbing Wu, Xian Yang, Xinyan Chen, Jiayin Liu, Haowen Li, Yu Su, and Yu Zhang. MemP: Procedural memory from trajectories for language agents. arXiv preprint arXiv:2508.06433. 2025.
[14] Peter Auer, Nicolo Cesa-Bianchi, and Paul Fischer. Finite-time analysis of the multiarmed bandit problem. Machine Learning, 47:235–256, 2002.<ref:2607.
Improvements for AI systems
- Bold header: Learned Memory Control via Memory MDP
MEMCON learns an online policy that adaptively decides when, what, and how much to retrieve, when to inject a distilled plan, and when to consolidate or forget,
modeling memory operations as a sequential decision problem captured by an MDP.
This allows agents to select actions like RETRIEVE (varying depth), PLANINJECT, RE-RETRIEVE
based on the current task state and memory status.
- Bold header: Zero-Call Adaptive Controller
MEMCON is a lightweight controller that is adaptive without any extra LLM calls,
which is achieved by learning a policy using a lightweight tabular contextual bandit with UCB exploration.
This contrasts with systems like MemGPT, which incur an additional LLM call per memory operation.
- Bold header: Task-Progress State Fusion
The Memory MDP state captures both task progress (goal type, step phase, stuck indicator, locations visited)
and memory status (size, plan availability, learning phase).
This fusion enables adaptive behavior based on the specific regime: Early tasks should retrieve less,
and Stuck agents that repeat actions should trigger re-retrieval with an alternative query.
- Bold header: Plan Injection for Efficiency
The system implements a learned action called PLANINJECT
which prepends a generalized success plan when one is available, allowing the agent to use a distilled, object-generalized template of a prior success
instead of generic nearest-neighbor lookup.
Sources
- The Rise and Potential of Large Language Model Based Agents: A Survey
- G-Memory: Tracing Hierarchical Memory for Multi-Agent Systems
- Generation of pure, spin polarized, and unpolarized charge currents at the few cycle limit of circularly polarized light
- MemGPT: Towards LLMs as Operating Systems
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- Memp: Exploring Agent Procedural Memory
- WebWalker: Benchmarking LLMs in Web Traversal
- GAIA: a benchmark for General AI Assistants
- Signal-First Architectures: Rethinking Front-End Reactivity
- Is Independent Learning All You Need in the StarCraft Multi-Agent Challenge?
- Hierarchical Auto-Organizing System for Open-Ended Multi-Agent Navigation
- AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
- Mindstorms in Natural Language-Based Societies of Mind
- Multi-Agent Collaboration: Harnessing the Power of Intelligent LLM Agents
- MemLLM: Finetuning LLMs to Use An Explicit Read-Write Memory
- Think-in-Memory: Recalling and Post-thinking Enable LLMs with Long-Term Memory
- Agent Workflow Memory
- Skill-Pro: Learning Reusable Skills from Experience via Non-Parametric PPO for LLM Agents
- MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents
- HiAgent: Hierarchical Working Memory Management for Solving Long-Horizon Agent Tasks with Large Language Model
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering