INMS: Memory Sharing for Large Language Model based Agents
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "INMS: Memory Sharing for Large Language Model based Agents".
Jane: The gist:
Tom: First, who's behind it and why it matters.
Title and authors: Tom: Welcome back. We’re looking at a paper today called "INMS: Memory Sharing for Large Language Model based Agents." This framework is about moving past agents that just work in isolation and starting to let them actually talk to each other.
Jane: Right, Tom, so basically, this paper is tackling the problem where LLM agents get stuck working alone because they only look at their own history or a static database. They miss out on the dynamic knowledge exchange that happens when people have a real conversation.
Lu: It’s about creating this shared conversational memory pool. This isn't just storing old answers; it’s about making those memories continuously shared among the agents, which is what they call collective self-enhancement >
Meng: So, if I get an answer in one task, that memory gets stored and then used by other agents later? That sounds like a way to build up a group knowledge base without having to manually feed every piece of information into every single agent.
Tom: Exactly, Meng. The core idea is that this framework uses real-time memory filtering and retrieval to make sure the shared pool stays useful and relevant for everyone involved >
Jane: So, the paper suggests that by letting agents share these memories continuously, they can dynamically refine how they search for information based on what others have already learned >
Lu: That’s a key part. They aren't just dumping everything into one giant bucket; they are using interaction history to make the retrieval mediator smarter over time >
Meng: I wonder how practical this is, Tom. If we add too much noise or irrelevant stuff to the shared pool, will it actually slow down the agents or just confuse them?
Tom: That’s a fair concern, Meng. The authors address that by using a specific scoring formula—Sfinal = one/two Xn i=one Li + Xn i=one Hi! — to decide which memory pairs get added to the pool >
Jane: It seems like they are trying to grade the quality of each memory pair before it gets shared, which sounds like a good way to keep the conversation high quality >
Lu: And then when an agent asks a new question, they use that scoring mechanism again with a formula p(xi, yi) = P(¬Y (xi, yi), X) to find the most relevant memories from the pool >
Title and authors: Tom: That means the system is constantly learning which types of shared experiences are actually helpful for answering future queries >
Jane: It shifts the focus from individual agent performance to collective intelligence, where agents can learn from each other in real time >
Meng: So, what’s the big win here for an actual application? Is this just theoretical stuff for now, or does it have immediate use in building smarter collaborative workflows?
Tom: The experiments show that when agents use these shared memories, their performance across three different kinds of tasks—like creative writing and planning—is significantly improved compared to when they operate alone >
Jane: And the results are pretty encouraging because they show that using three shared memories yields the best performance in their tests >
Lu: They found that traditional search methods, like BM25 or random retrieval, just don't generalize well across different kinds of questions, whereas this approach is much better at adapting to evolving memory contexts >
Meng: So the paper’s main contribution seems to be moving away from those static embedding methods toward a dynamic system that actually learns from the interaction itself >
Tom: It really does. They are showing how continuous, interactive memory generation can lead to agents that evolve from being isolated problem solvers into something more like a collectively intelligent society >
Jane: It suggests that for complex, open-ended tasks, just having a big brain isn't enough; you need this kind of ongoing dialogue mechanism to truly grasp the nuances >
Lu: The implication is that future agents won't just be powerful individual tools but will be systems defined by how well they can maintain and refine their shared knowledge base through constant back-and-forth >
Meng: From an engineering standpoint, it means we have a blueprint for building these more robust, multi-agent environments where the agents aren't just passing data back and forth blindly >
Tom: It’s about modeling that natural asynchronous dialogue between agents so they can truly benefit from each other's specialized training >
Jane: So we’ve seen how this INMS framework handles memory generation and retrieval, but what does this mean for the long-term direction of agent development?
Lu: It opens up possibilities where agents could build upon each other's expertise in ways that weren't possible when they were trained in isolation >
Title and authors: Meng: I see a potential for creating specialized teams of agents where knowledge flows naturally between them, rather than needing a central knowledge repository managed externally >
Tom: Absolutely. It’s about modeling multi-agent communication and collective knowledge sharing in a way that feels like genuine dialogue, not just data transfer >
Jane: So we’ve seen how INMS improves agent performance in creative and planning domains by using this shared memory mechanism, but what are the limitations they mention?
Lu: They point out that the organization of the shared memory pool can still be heavily shaped by the capabilities of the underlying language models themselves >
Meng: That means even with this smart retrieval system, if our foundational LLM isn't good at understanding context, we might just end up with a messy pool of memories >
Tom: Right. They also noted that their current prototype is focused on text-only interactions, so extending it to multimodal inputs like images or audio would be a natural next step for richer context >
Jane: So the paper gives us a clear picture of how we can model this asynchronous interaction, but they’re being cautious about how much control we have over the final memory structure >
Lu: It sets the stage for future work focusing on extending this framework to handle different types of data input and ensuring that the memory filtering remains effective across those modalities >
Meng: I think seeing these results on planning and creation tasks gives us a solid starting point for testing how we integrate this into our current agent architecture >
Tom: It certainly does. Overall, "INMS: Memory Sharing for Large Language Model based Agents" shows a path toward agents that are not just solvers, but participants in a continuous dialogue about knowledge >
Jane: So to wrap up on the INMS framework, it’s an asynchronous way for agents to build a shared conversational ground through filtering and retrieval, which leads to collective self-enhancement >
Lu: It moves us from isolated reasoning toward an interactive, collectively intelligent system driven by continuous dialogue >
Meng: It gives us a practical mechanism for managing how we curate the knowledge base in complex agent systems >
Tom: That’s the essence of it. We’ve seen how this paper tackles multi-agent communication and collective knowledge sharing to improve performance across various tasks >
The paper's summary: Tom: So, we've seen how INMS creates this shared memory pool for agents to talk to each other >
Jane: Exactly, so it’s about moving past agents just looking at their own isolated history and instead making them share a conversational ground >
Tom: The authors are saying this framework sets up a way for agents to constantly exchange information in real time, which helps them collectively improve what they know >
Jane: They’re not just storing old answers; they’re using that shared memory as an active tool to help the agents refine their own search strategy based on what everyone else has learned >
Tom: It sounds like a feedback loop where the interaction history directly shapes how useful the memory pool becomes for future queries >
Jane: That's right, it’s about modeling that natural back-and-forth dialogue so agents can actually build a more robust collective understanding of a task >
Tom: The experiments show that this approach works across different types of tasks—like creative writing and planning—where they actually saw performance bumps >
Jane: And it’s interesting because they found that using three shared memories yielded the best results in their tests, which suggests there's an optimal amount of shared context >
Tom: They also pointed out that this method handles different kinds of knowledge better than traditional search methods like BM25 or simple dense retrievers >
Jane: The big implication here is that instead of every agent having to reinvent the wheel for every query, they can tap into a larger pool of collective experience >
Tom: This moves agents from being just isolated solvers toward something that’s more like a collectively intelligent society driven by continuous knowledge exchange >
Jane: It suggests that future systems won't just be powerful individual tools but will be defined by how well they can maintain and refine this shared knowledge through constant interaction >
Tom: But the authors also admitted their current setup is text-only, so extending it to things like images or audio would be a natural next step for richer context >
Jane: That’s a fair point; they're showing us the core idea now, but there's definitely room to expand that into multimodal interactions down the line >
The paper's improvements: Tom: So, we're looking at how INMS improves things—it’s not just about having one pool of memory, but actually making that pool work >
Jane: Right, so the authors are suggesting ways to make sure that shared memory stays high quality and doesn't just get filled with junk over time >
Tom: They introduced this scoring formula for when a new memory pair gets added, which means they’re actively filtering out the less useful stuff before it becomes part of the collective knowledge base >
Jane: That’s smart because it stops the system from getting bogged down by low-quality interactions, which keeps the overall intelligence higher >
Tom: And when agents retrieve information, they use a scoring mechanism too—a formula that checks how similar a new question is to the existing memories in that pool >
Jane: So it’s not just pulling anything out; it’s intelligently selecting the most relevant experiential context based on what the agents have already exchanged >
Tom: This means even if an agent has a million memories, if its question only matches a few specific types of past interactions, it gets those specific ones prioritized >
Jane: It takes the static embeddings that we talked about before and makes them dynamic by constantly updating how similar things are based on new dialogue >
Tom: The results show that this dynamic updating actually leads to better performance across different domains, like creative writing and planning, compared to older methods >
Jane: And they found that using three shared memories gave the best performance in their tests, which points toward a sweet spot for how much collective experience an agent needs >
Tom: The authors are pushing the idea that this continuous dialogue isn't just a neat trick; it’s essential for agents to evolve from being isolated solvers into something more like a collectively intelligent society >
Jane: It implies that future AI development shouldn't just focus on bigger models, but on building these systems where knowledge flows naturally between them >
Tom: That leads us to think about how this could affect specialized teams of AI agents, where they could genuinely build upon each other’s specific expertise >
Jane: And we need to look at how this continuous learning process helps the whole culture around AI—it moves it from a product you use to a partner that improves itself >
Tom: But there are still guardrails; the authors admit that even with these improvements, the final organization of that memory pool can still be heavily influenced by the underlying language model's own capabilities >
Jane: They also flagged something important: right now, this framework is text-only, so they see extending it to things like images or audio as a key area for future work >
Conclusion: Tom: So we’re wrapping up on INMS: Memory Sharing for Large Language Model based Agents by summarizing how this framework moves AI agents past isolated thinking >
Jane: Basically, they showed us how to create a shared conversational ground that lets agents continuously learn from each other in real time >
Tom: The authors are really pushing the idea that this constant exchange of information is what will allow these systems to evolve into something more like a society with collective intelligence >
Jane: It suggests that the future of AI isn't just about bigger models, but about building these interconnected systems where knowledge flows naturally between them >
Tom: The main point is moving away from agents that are just static problem solvers toward interactive entities driven by ongoing dialogue >
Jane: And they did show us how this system improves performance across different tasks, proving it’s a solid mechanism for collaboration >
Tom: It’s encouraging to see this kind of modeling of multi-agent communication working across things like creative writing and planning without needing a manual database update for every single agent >
Jane: But they did flag that their current version is text-only, so there's definite work ahead to see how this framework handles more complex inputs like images or audio >
Tom: Right, the limitations are real, but the direction is clear—we need to keep pushing these concepts into multimodal spaces >
Jane: It sounds like the next big thing for this research is integrating those different types of data so agents can share a richer context >
Tom: I think we’re going to take a look at how this memory sharing idea compares to other ways agents are learning, maybe looking at things like continuous decoding or alignment methods next >
Jane: That makes sense. We’ve seen the power of shared memory, but now we need to see how it fits into the broader picture of safety and alignment >
Tom: Exactly. Next up, we’re going to look at how AI can steer itself during inference to make sure it stays on track with its goals, which is a totally different kind of challenge >
Jane: It's a big step from just sharing memories to actually steering the model's behavior in real time >
Tom: We’ll talk about that next, so stick around for more deep dives into the world of AI research >
Lu: I'm really excited about the potential for this. Imagine a truly fluid knowledge space where every interaction builds on everything else, it’s like an emergent intelligence system in action >
Meng: From an engineering side, I’m thinking about how we can build that filtering mechanism to be fast and reliable enough for real-world deployment, that's the practical hurdle >
Lalam: If agents can share knowledge this effectively, it could fundamentally improve how we structure information in our entire cultural landscape, making collective understanding a natural part of how things get done >
Hang Gao, Yongfeng Zhang
Rutgers University
cs.CL
Submitted: 2024-04-15
Updated: 2026-10-04
Comments: IJCNLP-AACL 2026 Main
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 83/100
The gist: The gist: The INteractive Memory Sharing (INMS) framework is an asynchronous interaction paradigm for multi-agent systems that establishes a shared conversational memory pool to promote collective
Key concepts
- Interactive Memory Sharing (INMS)
- A system where multiple AI agents share a common memory pool during asynchronous interactions. Agents generate and score their dialogue memories, allowing them to learn from each other's experiences in real-time. This promotes collective self-enhancement.
- Memory Generation and Selection
- This process involves creating Prompt-Answer (PA) pairs as memories. A Large Language Model scorer then evaluates these pairs using a formula based on domain rubrics to decide which memories are valuable enough to add to the shared pool.
- Memory Retrieval and Training
- Agents use a dense retriever, guided by cosine similarity, to fetch relevant past memories from the shared pool when answering new questions. Crucially, every new memory added also trains this retriever using a specific scoring mechanism based on conditional probability.
Terminology
Summary
The gist: The INteractive Memory Sharing (INMS) framework is an asynchronous interaction paradigm for multi-agent systems that establishes a shared conversational memory pool to promote collective self-enhancement and dynamically refine the retrieval mediator based on interaction history.
Introduction
Large Language Model (LLM) based agents excel at complex tasks, but their performance in open-ended scenarios is often constrained by isolated operation and reliance on static databases, missing the dynamic knowledge exchange of human dialogue To bridge this gap, we propose the INteractive Memory Sharing (INMS) framework, an asynchronous interaction paradigm for multi-agent systems By integrating real-time memory filtering, storage, and retrieval, INMS establishes a shared conversational memory pool This enables continuous, dialogue-like memory sharing among agents, promoting collective self-enhancement and dynamically refining the retrieval mediator based on interaction history Extensive experiments across three datasets demonstrate that INMS significantly improves agent performance by effectively modeling multi-agent interaction and collective knowledge sharing
Related Work
Existing methods primarily focus on agents independently utilizing stored historical information to inform responses, which neglects the profound potential of inter-agent interactions and collective memory utilization Current approaches often fail to model the asynchronous dialogue and knowledge exchange that naturally occurs in complex multi-agent environments, missing out on the inherent diversity and complementarity of agents that hold unique conversational histories and specialized training The INMS framework shifts the paradigm from isolated reasoning to an implicit, highly efficient asynchronous dialogue mechanism
The Interactive Memory Sharing (INMS) Framework
The core components of INMS are detailed below, focusing on dynamic memory generation and retrieval
-
Memory Generation and Selection: A memory is essentially a Prompt-Answer (PA) pair, stored in natural language to serve as shared memories After each interaction, the PA pair is scored by a LLM scorer which decides whether to add it to the pool The grading process involves establishing grading rubrics for each domain and using a formula Sfinal = 1/2 Xn i=1 Li + Xn i=1 Hi! (1) where n be the number of rubrics in the scoring criteria
-
Memory Retrieval and Training: Agents retrieve memories from the shared memory pool based on a question with the help of a dense retriever, which are more similar to the target question in terms of cosine similarity Whenever a new PA pair (memory), denoted as (X, Y), is added into the memory pool, it will also be used to train our retriever The scoring mechanism employed is defined as p(xi, yi) = P(¬Y (xi, yi), X), i ∈ 1, …, n (2)
Experiments and Results
The experiments were conducted across three domains: Literary Creation, Unconventional Logic Problem-solving, and Plan Generation Table 2 compares INMS with baselines under both F1 and LLM Judge (LLM-J) metrics Table 3 shows performance across agents with different numbers of PA pairs evaluated by ROUGE-2 and ROUGE-L against gpt-4o Table 4 compares INMS with baselines under both F1 and LLM Judge (LLM-J) metrics Table 5 shows performance across different LLM when equipped with the domain pool and single pool
Main Results
For each agent, we first tested them with using the same backbone and the metric BertScore, that is, in each domain, all memory was generated by agents utilizing the same Large Language Model We can observe that, for all agents among all the tasks, compare to no use of the shared memories, the performance of all the agents has been significantly improved This suggests that the shareable memories from other tasks can help agents get desired answers, rather than interfering with the agents’ learning ability Furthermore, when using three shared memories yields the best performance, all subsequent experiments utilize three shared memories during testing and gpt-4o as the backbone INMS consistently surpasses both sparse and dense retrieval baselines Traditional methods such as BM25 and Random exhibit limited ability to generalize across diverse query types, while contrastive-based dense retrievers achieve moderate gains but remain constrained by static embeddings that fail to adapt to evolving memory contexts
Conclusion
INMS introduces a novel framework modeling real-time memory sharing among agents as asynchronous dialogue via continuous storage and retrieval This dynamically evolving shared memory acts as a conversational ground, enhancing agents’ ability to collaboratively understand nuances and generate high-quality responses Experiments demonstrate its effectiveness in modeling multi-agent communication, successfully overcoming early-stage echo chamber
biases as the pool stabilizes Ultimately, INMS paves the way for agents to evolve from isolated solvers into a collectively intelligent society driven by continuous dialogue and knowledge exchange.
Limitations
Our INMS framework shows encouraging performance, but there remain several directions worth pursuing First, this study emphasizes validating the core idea of Interactive Memory Sharing rather than exhaustive systems tuning Second, although the framework applies memory filtering, the resulting organization can still be shaped by the capabilities of the underlying language models Finally, our current prototype targets text-only interactions; extending it to multimodal inputs, e.g., images or audio, could supply richer context and is a natural avenue for future research
References
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, and 1 others. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 Toufique Ahmed and Premkumar Devanbu. 2022. Few-shot training llms for project-specific codesummarization Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, and 1 others. 2020. Language models are few-shot learners. Advances in neural information processing systems Hang Gao and Yongfeng Zhang. 2024. Vrsd: Rethinking similarity and diversity for retrieval in large language models Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig. 2023. Pal: Program-aided language models Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021. Simcse: Simple contrastive learning of sentence embeddings Jiawei Gu, Xuhui Jiang, Zhichao Shi, Hexiang Tan, Xuehao Zhai, Chengjin Xu, Wei Li, Yinghan Shen, Shengjie Ma, Honghao Liu, and 1 others. 2024. A survey on llm-as-a-judge Joy He-Yueya, Gabriel Poesia, Rose E Wang, and Noah D Goodman. 2023. Solving math word problems by combining language models with symbolic solvers Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bojanowski, Armand Joulin, and Edouard Grave. 2021. Unsupervised dense information retrieval with contrastive learning Gautier Izacard, Patrick Lewis, Maria Lomeli, Lucas Hosseini, Fabio Petroni, Timo Schick, Jane DwivediYu, Armand Joulin, Sebastian Riedel, and Edouard Grave. 2023. Atlas: Few-shot learning with retrieval augmented language models Albert Qiaochu Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stocktvein Le Scao. 2023. Mistral 7b Andrew K Lampinen, Ishita Dasgupta, Stephanie CY Chan, Kory Matthewson, Michael Henry Tessler, Antonia Creswell, James L McClelland, Jane X Wang, and Felix Hill. 2022. Can language models learn from explanations in context? arXiv preprint arXiv:2204.02329 Yoav Levine, Noam Wies, Daniel Jannai, Dan Navon, Yedid Hoshen, and Amnon Shashua. 2021. The inductive bias of in-context learning: Rethinking pretraining example design. arXiv preprint arXiv:2110.04541 Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, and 1 others.
Improvements for AI systems
-
INMS establishes a shared conversational memory pool,
enablingcontinuous, dialogue-like memory sharing among agents.
This allows multi-agent systems to engage incollective self-enhancement
by promotingcollective knowledge sharing.
-
The system can model multi-agent communication by generating Prompt-Answer (PA) pairs which are rigorously evaluated by a dedicated LLM scorer, filtering out content to
maintain a high-quality interaction history.
-
The autonomous retriever is dynamically updated based on the growing pool, acting as an
adaptable mediator of this implicit dialogue,
ensuring thatthe most relevant experiential context is curated for each query.
-
The framework addresses static database limitations by modeling
continuous, interactive memory generation,
facilitating adialogue-driven evolution from isolated agents to collective intelligence.
-
The system can sustain high-quality multi-agent discourse across diverse open-ended scenarios, including
poetry generation, unconventional logical problem-solving, and plan creation domains.
-
Agents can benefit from cross-domain knowledge by utilizing
heterologous shared memories,
where memories generated by agents using different LLM backbones can stillboost the performance for all the agents in answering the open-ended questions.
-
The system mitigates initial bias through its moderation mechanism, as it tracks performance across interaction accumulation phases and demonstrates resilience against an
initial 'echo chamber,' a shared memory pool dominated by biased early interactions.
Abstract
While Large Language Model (LLM) based agents excel at complex tasks, their performance in open-ended scenarios is often constrained by isolated operation and reliance on static databases, missing the dynamic knowledge exchange of human dialogue. To bridge this gap, we propose the INteractive Memory Sharing (INMS) framework, an asynchronous interaction paradigm for multi-agent systems. By integrating real-time memory filtering, storage, and retrieval, INMS establishes a shared conversational memory pool. This enables continuous, dialogue-like memory sharing among agents, promoting collective self-enhancement and dynamically refining the retrieval mediator based on interaction history. Extensive experiments across three datasets demonstrate that INMS improves agent performance by effectively modeling multi-agent interaction and collective knowledge sharing.
Sources
- GPT-4 Technical Report
- Vector Retrieval with Similarity and Diversity: How Hard Is It?
- SimCSE: Simple Contrastive Learning of Sentence Embeddings
- A Survey on LLM-as-a-Judge
- Solving Math Word Problems by Combining Language Models With Symbolic Solvers
- Unsupervised Dense Information Retrieval with Contrastive Learning
- Mistral 7B
- Can language models learn from explanations in context?
- The Inductive Bias of In-Context Learning: Rethinking Pretraining Example Design
- Think-in-Memory: Recalling and Post-thinking Enable LLMs with Long-Term Memory
- MemoChat: Tuning LLMs to Use Memos for Consistent Long-Range Open-Domain Conversation
- Dr.ICL: Demonstration-Retrieved In-context Learning
- MemGPT: Towards LLMs as Operating Systems
- Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
- REPLUG: Retrieval-Augmented Black-Box Language Models
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT
- BERTScore: Evaluating Text Generation with BERT
- Synapse: Trajectory-as-Exemplar Prompting with Memory for Computer Control
- Large Language Models Are Human-Level Prompt Engineers
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering