Shared Selective Persistent Memory for Agentic LLM Systems
summary
The gist
This paper introduces a "shared selective persistent memory" architecture designed to solve the statelessness of agentic LLM systems.
In short
Apple researchers developed an architecture for agentic LLM systems using shared selective memory. By organizing information into task specifications, data schemas, tool configurations, and output constraints rather than full histories, the system achieves a 96% task completion rate while significantly reducing token costs and task time.
Key concepts
- Agentic LLM Systems
- These are AI systems that go beyond simple chatting to perform active tasks like writing code or managing files. They function more like professional teammates capable of working within shared digital spaces rather than just conversational assistants.
- Selective Memory
- Instead of saving entire conversation histories, which can confuse the AI, this method organizes memory into four buckets: task specifications, data schemas, tool configurations, and output constraints. This helps the AI focus on essential rules while discarding distracting past mistakes.
- Zero-token Data Refresh
- This technique allows an AI to update its information without needing new prompts. By writing programs that look for data at specific points, the system can run with updated numbers automatically, reducing task completion time by 14 times and lowering token costs significantly.
Terminology used across episodes
This episode discusses
- Shared Selective Persistent Memory for Agentic LLM Systems · Paper Radio
- Extending Context Window of Large Language Models via Positional Interpolation
- Active Learning Framework for Cost-Effective TCR-Epitope Binding Affinity Prediction
- Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models
- MemGPT: Towards LLMs as Operating Systems
- Gorilla: Large Language Model Connected with Massive APIs
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
The paper
Shared Selective Persistent Memory for Agentic LLM Systems · Read on arXiv
Apple Inc.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Shared Selective Persistent Memory for Agentic LLM Systems".
Jane: The paper was written by Sanjana Pedada, Aditya Dhavala and Neelraj Patil from Apple Inc..
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: We’re looking at a fascinating new paper today called "Shared Selective Persistent Memory for Agentic LLM Systems."
Jane: That title sounds a bit like a mouthful, Tom, but I think I can help unpack it for everyone.
Tom: It definitely does, Jane, so how would you explain what an "agentic" system actually is?
Jane: Think of it as an AI that doesn't just chat with you, but actually goes out and does tasks like writing code or managing files.
Lu: And the "shared" part of that title is what really gets my imagination racing.
Tom: You mean because the memory isn't just stuck in one person's head, Lu?
Lu: Exactly, because this architecture allows different people to work in the same digital space and share what the AI has learned.
Meng: That sounds great for collaboration, but I'm curious about the authors behind this.
Jane: It's a team from Apple Inc., including Sanjana Pedada, Aditya Dhavala, and Neelraj Patil.
Tom: Apple is clearly looking at how these agents can move from being simple assistants to becoming part of a professional team.
Meng: I wonder if they're thinking about how this actually scales in a real company.
Jane: They certainly are, since they're focusing on how these agents can remember the rules of a job without having to be retaught every single morning.
Lu: It's like giving an AI a permanent professional identity instead of a goldfish's memory.
Lalam: This shift toward shared intelligence could actually change how we build community in digital workspaces.
Tom: That's a big thought, Lalam, so let's look at what this memory actually looks like in practice.
Summary: Tom: We've established that these agents are getting a much better memory, but the paper says they have to be "selective" about it.
Jane: That's the clever part, because they aren't just saving every single word the AI says.
Tom: Right, because if you save everything, doesn't the AI just get confused by its own old mistakes?
Jane: It does, and the paper explains that saving the whole history can actually make the AI perform worse.
Lu: It's like if you tried to learn a new recipe but kept reading all your old, failed attempts at cooking along with the new instructions.
Meng: So instead of a giant pile of old chats, they've broken memory down into four specific buckets.
Jane: They've got task specifications, data schemas, tool configurations, and output constraints.
Tom: That sounds a lot more organized than just a long transcript, Meng.
Meng: It is, and it means the AI knows the rules and the data types without needing to see the messy reasoning steps from yesterday.
Lu: They are essentially teaching the AI to remember the "what" while forgetting the "how" of its previous struggles.
Jane: That "selective forgetting" is such a human way to handle information, isn't it?
Lalam: By discarding the noise, we allow the AI to focus on the essence of the task, which helps preserve the clarity of human intent.
Tom: It sounds like they're building a much cleaner brain for these agents.
Improvements: Tom: Now we have to talk about the actual results, because the numbers in this paper are pretty wild.
Jane: They found that using this selective memory led to a ninety-six percent task completion rate.
Tom: And that's a huge jump compared to the seventy-one percent you get if you just give the AI its entire old history.
Meng: Wait, so giving the AI more information actually made it perform worse?
Jane: Yes, because the old reasoning traces act like distractions that bias the agent toward wrong paths.
Lu: It's a total paradigm shift in how we think about "more context" being "better context."
Tom: And there's this "zero-token data refresh" thing that sounds like magic.
Meng: I was going to ask about that, because how can you update data without talking to the AI again?
Jane: They've designed the AI to write programs that look for data at a specific point, so when the data changes, the program just runs again with the new numbers.
Tom: They're claiming that can reduce the time it takes to finish a task by fourteen times!
Lu: Imagine a dashboard that updates itself instantly without a single new prompt being sent.
Meng: That would save a massive amount of computing power and money in a real production environment.
Jane: They even saw a ninety-seven times reduction in token costs compared to just throwing raw data at the model.
Lalam: This efficiency isn't just about saving resources; it's about making AI interactions feel seamless and invisible in our daily lives.
Tom: It really makes you realize that managing context is just as important as the model itself.
Conclusion: Tom: We've covered a lot of ground today, from the Apple team's new architecture to those incredible efficiency gains.
Jane: It really feels like we're moving away from "chatting with a bot" and toward "working with a teammate."
Tom: The paper "Shared Selective Persistent Memory for Agentic LLM Systems" really sets a new standard for how these systems should behave.
Lu: I can see a future where these workspaces are the primary way humans and AI co-create everything.
Meng: From my side, I'm just excited to see how this makes deploying agents in actual enterprise software much more stable.
Lalam: This technology will help bridge the gap between raw computation and meaningful, shared human culture.
Tom: Well, that's all the time we have for this one, thanks for joining us.
Jane: See you next time!
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language