Shared Selective Persistent Memory for Agentic LLM Systems
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Shared Selective Persistent Memory for Agentic LLM Systems".
Jane: The paper was written by Sanjana Pedada, Aditya Dhavala and Neelraj Patil from Apple Inc..
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: We’re looking at a fascinating new paper today called "Shared Selective Persistent Memory for Agentic LLM Systems."
Jane: That title sounds a bit like a mouthful, Tom, but I think I can help unpack it for everyone.
Tom: It definitely does, Jane, so how would you explain what an "agentic" system actually is?
Jane: Think of it as an AI that doesn't just chat with you, but actually goes out and does tasks like writing code or managing files.
Lu: And the "shared" part of that title is what really gets my imagination racing.
Tom: You mean because the memory isn't just stuck in one person's head, Lu?
Lu: Exactly, because this architecture allows different people to work in the same digital space and share what the AI has learned.
Meng: That sounds great for collaboration, but I'm curious about the authors behind this.
Jane: It's a team from Apple Inc., including Sanjana Pedada, Aditya Dhavala, and Neelraj Patil.
Tom: Apple is clearly looking at how these agents can move from being simple assistants to becoming part of a professional team.
Meng: I wonder if they're thinking about how this actually scales in a real company.
Jane: They certainly are, since they're focusing on how these agents can remember the rules of a job without having to be retaught every single morning.
Lu: It's like giving an AI a permanent professional identity instead of a goldfish's memory.
Lalam: This shift toward shared intelligence could actually change how we build community in digital workspaces.
Tom: That's a big thought, Lalam, so let's look at what this memory actually looks like in practice.
Summary: Tom: We've established that these agents are getting a much better memory, but the paper says they have to be "selective" about it.
Jane: That's the clever part, because they aren't just saving every single word the AI says.
Tom: Right, because if you save everything, doesn't the AI just get confused by its own old mistakes?
Jane: It does, and the paper explains that saving the whole history can actually make the AI perform worse.
Lu: It's like if you tried to learn a new recipe but kept reading all your old, failed attempts at cooking along with the new instructions.
Meng: So instead of a giant pile of old chats, they've broken memory down into four specific buckets.
Jane: They've got task specifications, data schemas, tool configurations, and output constraints.
Tom: That sounds a lot more organized than just a long transcript, Meng.
Meng: It is, and it means the AI knows the rules and the data types without needing to see the messy reasoning steps from yesterday.
Lu: They are essentially teaching the AI to remember the "what" while forgetting the "how" of its previous struggles.
Jane: That "selective forgetting" is such a human way to handle information, isn't it?
Lalam: By discarding the noise, we allow the AI to focus on the essence of the task, which helps preserve the clarity of human intent.
Tom: It sounds like they're building a much cleaner brain for these agents.
Improvements: Tom: Now we have to talk about the actual results, because the numbers in this paper are pretty wild.
Jane: They found that using this selective memory led to a ninety-six percent task completion rate.
Tom: And that's a huge jump compared to the seventy-one percent you get if you just give the AI its entire old history.
Meng: Wait, so giving the AI more information actually made it perform worse?
Jane: Yes, because the old reasoning traces act like distractions that bias the agent toward wrong paths.
Lu: It's a total paradigm shift in how we think about "more context" being "better context."
Tom: And there's this "zero-token data refresh" thing that sounds like magic.
Meng: I was going to ask about that, because how can you update data without talking to the AI again?
Jane: They've designed the AI to write programs that look for data at a specific point, so when the data changes, the program just runs again with the new numbers.
Tom: They're claiming that can reduce the time it takes to finish a task by fourteen times!
Lu: Imagine a dashboard that updates itself instantly without a single new prompt being sent.
Meng: That would save a massive amount of computing power and money in a real production environment.
Jane: They even saw a ninety-seven times reduction in token costs compared to just throwing raw data at the model.
Lalam: This efficiency isn't just about saving resources; it's about making AI interactions feel seamless and invisible in our daily lives.
Tom: It really makes you realize that managing context is just as important as the model itself.
Conclusion: Tom: We've covered a lot of ground today, from the Apple team's new architecture to those incredible efficiency gains.
Jane: It really feels like we're moving away from "chatting with a bot" and toward "working with a teammate."
Tom: The paper "Shared Selective Persistent Memory for Agentic LLM Systems" really sets a new standard for how these systems should behave.
Lu: I can see a future where these workspaces are the primary way humans and AI co-create everything.
Meng: From my side, I'm just excited to see how this makes deploying agents in actual enterprise software much more stable.
Lalam: This technology will help bridge the gap between raw computation and meaningful, shared human culture.
Tom: Well, that's all the time we have for this one, thanks for joining us.
Jane: See you next time!
Apple Inc.
cs.AI, cs.MA, cs.SE
Submitted: 2026-07-10
Updated: 2026-09-15
Comments: 11 pages, 2 figures, 4 tables
Code: https://github.com/langchain-ai/langchain
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 87/100
The gist: This paper introduces a "shared selective persistent memory" architecture designed to solve the statelessness of agentic LLM systems.
Key concepts
- Agentic LLM Systems
- These are AI systems that go beyond simple chatting to perform active tasks like writing code or managing files. They function more like professional teammates capable of working within shared digital spaces rather than just conversational assistants.
- Selective Memory
- Instead of saving entire conversation histories, which can confuse the AI, this method organizes memory into four buckets: task specifications, data schemas, tool configurations, and output constraints. This helps the AI focus on essential rules while discarding distracting past mistakes.
- Zero-token Data Refresh
- This technique allows an AI to update its information without needing new prompts. By writing programs that look for data at specific points, the system can run with updated numbers automatically, reducing task completion time by 14 times and lowering token costs significantly.
Terminology
Summary
This paper introduces a shared selective persistent memory
architecture designed to solve the statelessness of agentic LLM systems. By moving away from naive full-history persistence—which can be actively harmful
due to stale reasoning traces—the authors propose a method to selectively retain reusable context, enabling more efficient, collaborative, and scalable enterprise workflows.
The selective memory architecture
The core innovation is a memory architecture that identifies and retains four specific categories of reusable context while explicitly discarding session-specific reasoning traces. This approach addresses the lost in the middle
phenomenon, where injecting irrelevant context like prior tool-use logs can bias an agent toward previously explored solution paths rather than current tasks. The architecture decomposes knowledge into:
-
Task specifications — custom system prompts that encode domain rules, output format constraints, and generation preferences.
-
Data schemas — precomputed summaries of associated data sources, including column names, types, and statistical distributions.
-
Tool configurations — the set of available external tools, their parameter schemas, invocation patterns, and authentication requirements.
-
Output constraints — structural contracts between the generated artifact and its runtime environment.
Implementation and collaboration
The system is implemented within a collaborative workspace platform
where agents produce git-versioned artifacts such as interactive dashboards, structured reports, and data-driven documents. This environment utilizes a multi-connector integration layer
to access heterogeneous data via CSV, SQL, REST APIs, and MCP servers. To facilitate teamwork and stability, the platform includes:
-
Git-backed versioning with draft isolation that allows users to explore modifications
risk-free.
-
Role-based access control that enables
collaborative reuse of accumulated context
across different users. -
An in-session undo stack that permits users to restore any prior artifact state without re-invoking the model.
Efficiency and data management
A critical component is a zero-token data refresh mechanism
that enforces a strict data-injection contract.
This ensures that generated programs consume data exclusively from a runtime injection point, rather than using hardcoded values. This decoupling enables:
-
Additive schema evolution, where new columns are permitted but missing columns trigger regeneration.
-
A 14× reduction in task time by eliminating LLM re-invocation for recurring data updates.
-
Summary-driven generation that achieves a
97× reduction
in per-invocation token cost compared to raw data injection by using statistical profiling instead of raw tabular data.
Experimental validation
The authors evaluated the architecture through enterprise deployment scenarios and public dataset replications. The results demonstrate that selective memory significantly outperforms both no-memory and full-history conditions, achieving a 96% task completion rate. Notably, naive full-history persistence actively degrades task completion
by biasing the agent with stale reasoning traces, resulting in only 71% completion. Furthermore, summary-driven generation proved highly efficient, reducing token costs by up to 946× on public datasets compared to raw injection.
Improvements for AI systems
1. Implement Structured Selective Memory Decomposition
Replace monolithic conversation history persistence with a four-category structured memory architecture: Task Specifications (domain rules and formatting preferences), Data Schemas (statistical profiles and column metadata), Tool Configurations (API schemas and authentication patterns), and Output Constraints (structural contracts). Explicitly implement Selective Forgetting
to purge all intermediate reasoning traces, tool-use logs, and error-recovery paths from the context window of subsequent sessions.
- Improved System Capability: The system will eliminate
trace anchoring
(where the agent is biased toward failed or stale solution paths) and avoid thelost in the middle
phenomenon. This results in higher task completion rates (targeting >95%) and significantly reduced user turns by preventing redundant re-specification of domain constraints.
2. Transition to Summary-Driven Data Injection
Replace raw data injection with a statistical profiling mechanism. Instead of feeding raw tabular data into the prompt, the system will inject a compact metadata object containing column types, statistical distributions (mean, std, min/max), categorical value catalogs (unique values), and minimal sample rows.
- Improved System Capability: The system can process massive enterprise datasets while maintaining a tiny context footprint (<1K tokens). This achieves a 97x to 946x reduction in token costs compared to raw injection, enabling the agent to reason over large-scale data without hitting context limits or incurring massive latency.
3. Enforce a Runtime Data-Injection Contract
Modify the agentic engine's generation instructions to mandate a strict separation between generated logic and runtime data. The LLM must produce code that consumes data exclusively through a dynamic injection point rather than using hardcoded values or local file paths.
- Improved System Capability: This enables
Zero-Token Data Refresh.
Users can update the underlying data source (e.g., replacing an old CSV with a live SQL database) and re-render interactive dashboards or reports instantly without any LLM re-invocation, reducing task time by up to 14x.
4. Implement Git-Backed Collaborative Workspaces with Draft Isolation
Integrate a version control system (Git) to manage generated artifacts and a metadata store for the four selective memory categories. Implement Draft Isolation,
where user refinements occur in sandboxed sessions that do not affect the published workspace until an explicit commit
or publish
operation is performed.
- Improved System Capability: This enables seamless cross-team collaboration. A lead user can publish a highly-tuned workspace (containing complex task specs and tool configs) as a template; colleagues can load it, swap in their own schema-compatible data, and iterate on refinements risk-free without corrupting the original production artifact.
Sources
- Extending Context Window of Large Language Models via Positional Interpolation
- Active Learning Framework for Cost-Effective TCR-Epitope Binding Affinity Prediction
- Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models
- MemGPT: Towards LLMs as Operating Systems
- Gorilla: Large Language Model Connected with Massive APIs
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection