INMS: Memory Sharing for Large Language Model based Agents
summary
The gist
The gist: The INteractive Memory Sharing (INMS) framework is an asynchronous interaction paradigm for multi-agent systems that establishes a shared conversational memory pool to promote collective
In short
The INMS framework introduces an asynchronous interaction paradigm for multi-agent systems by establishing a shared conversational memory pool. Agents continuously generate and score Prompt-Answer pairs, which are stored and used to dynamically refine a retrieval mediator. This collective knowledge sharing significantly improves agent performance across various complex tasks.
Key concepts
- Interactive Memory Sharing (INMS)
- A system where multiple AI agents share a common memory pool during asynchronous interactions. Agents generate and score their dialogue memories, allowing them to learn from each other's experiences in real-time. This promotes collective self-enhancement.
- Memory Generation and Selection
- This process involves creating Prompt-Answer (PA) pairs as memories. A Large Language Model scorer then evaluates these pairs using a formula based on domain rubrics to decide which memories are valuable enough to add to the shared pool.
- Memory Retrieval and Training
- Agents use a dense retriever, guided by cosine similarity, to fetch relevant past memories from the shared pool when answering new questions. Crucially, every new memory added also trains this retriever using a specific scoring mechanism based on conditional probability.
Terminology used across episodes
This episode discusses
- INMS: Memory Sharing for Large Language Model based Agents · Paper Radio
- GPT-4 Technical Report
- Vector Retrieval with Similarity and Diversity: How Hard Is It?
- SimCSE: Simple Contrastive Learning of Sentence Embeddings
- A Survey on LLM-as-a-Judge
- Solving Math Word Problems by Combining Language Models With Symbolic Solvers
- Unsupervised Dense Information Retrieval with Contrastive Learning
- Mistral 7B
- Can language models learn from explanations in context?
- The Inductive Bias of In-Context Learning: Rethinking Pretraining Example Design
- Think-in-Memory: Recalling and Post-thinking Enable LLMs with Long-Term Memory
- MemoChat: Tuning LLMs to Use Memos for Consistent Long-Range Open-Domain Conversation
- Dr.ICL: Demonstration-Retrieved In-context Learning
- MemGPT: Towards LLMs as Operating Systems
- Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
- REPLUG: Retrieval-Augmented Black-Box Language Models
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT
- BERTScore: Evaluating Text Generation with BERT
- Synapse: Trajectory-as-Exemplar Prompting with Memory for Computer Control
- Large Language Models Are Human-Level Prompt Engineers
The paper
INMS: Memory Sharing for Large Language Model based Agents · Read on arXiv
Hang Gao, Yongfeng Zhang
Rutgers University
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "INMS: Memory Sharing for Large Language Model based Agents".
Jane: The gist:
Tom: First, who's behind it and why it matters.
Title and authors: Tom: Welcome back. We’re looking at a paper today called "INMS: Memory Sharing for Large Language Model based Agents." This framework is about moving past agents that just work in isolation and starting to let them actually talk to each other.
Jane: Right, Tom, so basically, this paper is tackling the problem where LLM agents get stuck working alone because they only look at their own history or a static database. They miss out on the dynamic knowledge exchange that happens when people have a real conversation.
Lu: It’s about creating this shared conversational memory pool. This isn't just storing old answers; it’s about making those memories continuously shared among the agents, which is what they call collective self-enhancement >
Meng: So, if I get an answer in one task, that memory gets stored and then used by other agents later? That sounds like a way to build up a group knowledge base without having to manually feed every piece of information into every single agent.
Tom: Exactly, Meng. The core idea is that this framework uses real-time memory filtering and retrieval to make sure the shared pool stays useful and relevant for everyone involved >
Jane: So, the paper suggests that by letting agents share these memories continuously, they can dynamically refine how they search for information based on what others have already learned >
Lu: That’s a key part. They aren't just dumping everything into one giant bucket; they are using interaction history to make the retrieval mediator smarter over time >
Meng: I wonder how practical this is, Tom. If we add too much noise or irrelevant stuff to the shared pool, will it actually slow down the agents or just confuse them?
Tom: That’s a fair concern, Meng. The authors address that by using a specific scoring formula—Sfinal = one/two Xn i=one Li + Xn i=one Hi! — to decide which memory pairs get added to the pool >
Jane: It seems like they are trying to grade the quality of each memory pair before it gets shared, which sounds like a good way to keep the conversation high quality >
Lu: And then when an agent asks a new question, they use that scoring mechanism again with a formula p(xi, yi) = P(¬Y (xi, yi), X) to find the most relevant memories from the pool >
Title and authors: Tom: That means the system is constantly learning which types of shared experiences are actually helpful for answering future queries >
Jane: It shifts the focus from individual agent performance to collective intelligence, where agents can learn from each other in real time >
Meng: So, what’s the big win here for an actual application? Is this just theoretical stuff for now, or does it have immediate use in building smarter collaborative workflows?
Tom: The experiments show that when agents use these shared memories, their performance across three different kinds of tasks—like creative writing and planning—is significantly improved compared to when they operate alone >
Jane: And the results are pretty encouraging because they show that using three shared memories yields the best performance in their tests >
Lu: They found that traditional search methods, like BM25 or random retrieval, just don't generalize well across different kinds of questions, whereas this approach is much better at adapting to evolving memory contexts >
Meng: So the paper’s main contribution seems to be moving away from those static embedding methods toward a dynamic system that actually learns from the interaction itself >
Tom: It really does. They are showing how continuous, interactive memory generation can lead to agents that evolve from being isolated problem solvers into something more like a collectively intelligent society >
Jane: It suggests that for complex, open-ended tasks, just having a big brain isn't enough; you need this kind of ongoing dialogue mechanism to truly grasp the nuances >
Lu: The implication is that future agents won't just be powerful individual tools but will be systems defined by how well they can maintain and refine their shared knowledge base through constant back-and-forth >
Meng: From an engineering standpoint, it means we have a blueprint for building these more robust, multi-agent environments where the agents aren't just passing data back and forth blindly >
Tom: It’s about modeling that natural asynchronous dialogue between agents so they can truly benefit from each other's specialized training >
Jane: So we’ve seen how this INMS framework handles memory generation and retrieval, but what does this mean for the long-term direction of agent development?
Lu: It opens up possibilities where agents could build upon each other's expertise in ways that weren't possible when they were trained in isolation >
Title and authors: Meng: I see a potential for creating specialized teams of agents where knowledge flows naturally between them, rather than needing a central knowledge repository managed externally >
Tom: Absolutely. It’s about modeling multi-agent communication and collective knowledge sharing in a way that feels like genuine dialogue, not just data transfer >
Jane: So we’ve seen how INMS improves agent performance in creative and planning domains by using this shared memory mechanism, but what are the limitations they mention?
Lu: They point out that the organization of the shared memory pool can still be heavily shaped by the capabilities of the underlying language models themselves >
Meng: That means even with this smart retrieval system, if our foundational LLM isn't good at understanding context, we might just end up with a messy pool of memories >
Tom: Right. They also noted that their current prototype is focused on text-only interactions, so extending it to multimodal inputs like images or audio would be a natural next step for richer context >
Jane: So the paper gives us a clear picture of how we can model this asynchronous interaction, but they’re being cautious about how much control we have over the final memory structure >
Lu: It sets the stage for future work focusing on extending this framework to handle different types of data input and ensuring that the memory filtering remains effective across those modalities >
Meng: I think seeing these results on planning and creation tasks gives us a solid starting point for testing how we integrate this into our current agent architecture >
Tom: It certainly does. Overall, "INMS: Memory Sharing for Large Language Model based Agents" shows a path toward agents that are not just solvers, but participants in a continuous dialogue about knowledge >
Jane: So to wrap up on the INMS framework, it’s an asynchronous way for agents to build a shared conversational ground through filtering and retrieval, which leads to collective self-enhancement >
Lu: It moves us from isolated reasoning toward an interactive, collectively intelligent system driven by continuous dialogue >
Meng: It gives us a practical mechanism for managing how we curate the knowledge base in complex agent systems >
Tom: That’s the essence of it. We’ve seen how this paper tackles multi-agent communication and collective knowledge sharing to improve performance across various tasks >
The paper's summary: Tom: So, we've seen how INMS creates this shared memory pool for agents to talk to each other >
Jane: Exactly, so it’s about moving past agents just looking at their own isolated history and instead making them share a conversational ground >
Tom: The authors are saying this framework sets up a way for agents to constantly exchange information in real time, which helps them collectively improve what they know >
Jane: They’re not just storing old answers; they’re using that shared memory as an active tool to help the agents refine their own search strategy based on what everyone else has learned >
Tom: It sounds like a feedback loop where the interaction history directly shapes how useful the memory pool becomes for future queries >
Jane: That's right, it’s about modeling that natural back-and-forth dialogue so agents can actually build a more robust collective understanding of a task >
Tom: The experiments show that this approach works across different types of tasks—like creative writing and planning—where they actually saw performance bumps >
Jane: And it’s interesting because they found that using three shared memories yielded the best results in their tests, which suggests there's an optimal amount of shared context >
Tom: They also pointed out that this method handles different kinds of knowledge better than traditional search methods like BM25 or simple dense retrievers >
Jane: The big implication here is that instead of every agent having to reinvent the wheel for every query, they can tap into a larger pool of collective experience >
Tom: This moves agents from being just isolated solvers toward something that’s more like a collectively intelligent society driven by continuous knowledge exchange >
Jane: It suggests that future systems won't just be powerful individual tools but will be defined by how well they can maintain and refine this shared knowledge through constant interaction >
Tom: But the authors also admitted their current setup is text-only, so extending it to things like images or audio would be a natural next step for richer context >
Jane: That’s a fair point; they're showing us the core idea now, but there's definitely room to expand that into multimodal interactions down the line >
The paper's improvements: Tom: So, we're looking at how INMS improves things—it’s not just about having one pool of memory, but actually making that pool work >
Jane: Right, so the authors are suggesting ways to make sure that shared memory stays high quality and doesn't just get filled with junk over time >
Tom: They introduced this scoring formula for when a new memory pair gets added, which means they’re actively filtering out the less useful stuff before it becomes part of the collective knowledge base >
Jane: That’s smart because it stops the system from getting bogged down by low-quality interactions, which keeps the overall intelligence higher >
Tom: And when agents retrieve information, they use a scoring mechanism too—a formula that checks how similar a new question is to the existing memories in that pool >
Jane: So it’s not just pulling anything out; it’s intelligently selecting the most relevant experiential context based on what the agents have already exchanged >
Tom: This means even if an agent has a million memories, if its question only matches a few specific types of past interactions, it gets those specific ones prioritized >
Jane: It takes the static embeddings that we talked about before and makes them dynamic by constantly updating how similar things are based on new dialogue >
Tom: The results show that this dynamic updating actually leads to better performance across different domains, like creative writing and planning, compared to older methods >
Jane: And they found that using three shared memories gave the best performance in their tests, which points toward a sweet spot for how much collective experience an agent needs >
Tom: The authors are pushing the idea that this continuous dialogue isn't just a neat trick; it’s essential for agents to evolve from being isolated solvers into something more like a collectively intelligent society >
Jane: It implies that future AI development shouldn't just focus on bigger models, but on building these systems where knowledge flows naturally between them >
Tom: That leads us to think about how this could affect specialized teams of AI agents, where they could genuinely build upon each other’s specific expertise >
Jane: And we need to look at how this continuous learning process helps the whole culture around AI—it moves it from a product you use to a partner that improves itself >
Tom: But there are still guardrails; the authors admit that even with these improvements, the final organization of that memory pool can still be heavily influenced by the underlying language model's own capabilities >
Jane: They also flagged something important: right now, this framework is text-only, so they see extending it to things like images or audio as a key area for future work >
Conclusion: Tom: So we’re wrapping up on INMS: Memory Sharing for Large Language Model based Agents by summarizing how this framework moves AI agents past isolated thinking >
Jane: Basically, they showed us how to create a shared conversational ground that lets agents continuously learn from each other in real time >
Tom: The authors are really pushing the idea that this constant exchange of information is what will allow these systems to evolve into something more like a society with collective intelligence >
Jane: It suggests that the future of AI isn't just about bigger models, but about building these interconnected systems where knowledge flows naturally between them >
Tom: The main point is moving away from agents that are just static problem solvers toward interactive entities driven by ongoing dialogue >
Jane: And they did show us how this system improves performance across different tasks, proving it’s a solid mechanism for collaboration >
Tom: It’s encouraging to see this kind of modeling of multi-agent communication working across things like creative writing and planning without needing a manual database update for every single agent >
Jane: But they did flag that their current version is text-only, so there's definite work ahead to see how this framework handles more complex inputs like images or audio >
Tom: Right, the limitations are real, but the direction is clear—we need to keep pushing these concepts into multimodal spaces >
Jane: It sounds like the next big thing for this research is integrating those different types of data so agents can share a richer context >
Tom: I think we’re going to take a look at how this memory sharing idea compares to other ways agents are learning, maybe looking at things like continuous decoding or alignment methods next >
Jane: That makes sense. We’ve seen the power of shared memory, but now we need to see how it fits into the broader picture of safety and alignment >
Tom: Exactly. Next up, we’re going to look at how AI can steer itself during inference to make sure it stays on track with its goals, which is a totally different kind of challenge >
Jane: It's a big step from just sharing memories to actually steering the model's behavior in real time >
Tom: We’ll talk about that next, so stick around for more deep dives into the world of AI research >
Lu: I'm really excited about the potential for this. Imagine a truly fluid knowledge space where every interaction builds on everything else, it’s like an emergent intelligence system in action >
Meng: From an engineering side, I’m thinking about how we can build that filtering mechanism to be fast and reliable enough for real-world deployment, that's the practical hurdle >
Lalam: If agents can share knowledge this effectively, it could fundamentally improve how we structure information in our entire cultural landscape, making collective understanding a natural part of how things get done >
More episodes
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck
- 2407.14562-Thought-Like-Pro: Enhancing Reasoning of Large Language Models through Self-Bootstrapped Prolog-based Chain-of-Thought