Collaborative Memory for Multi-Agent VLM Systems
cs.AI
Submitted: 2026-09-15
Updated: 2026-09-18
Comments: First Draft Version: 4 pages, 3 figures
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Vision-language model (VLM) agents combine specialized perception, tools, and reasoning to address complex visual tasks.
Terminology
Abstract
Vision-language model (VLM) agents combine specialized perception, tools, and reasoning to address complex visual tasks. In multi-agent settings, different agents inspect different image regions, video frames, or visual representations, so collaboration extends beyond distributed reasoning to distributed perception. This makes shared visual context a central problem in VLM agent collaboration. In this paper, we frame memory hierarchy, cross-agent sharing, and consistency mechanisms around the need to reconcile interpretations and update dependent reasoning. Effective collaboration requires agents to build on contributions from other agents, recover missing visual context, and reconcile differing interpretations as new evidence emerges. Shared visual memory preserves not only images or textual summaries but also the dependencies among observations, agent interpretations, and subsequent reasoning. Together, these design considerations shape how information flows and evolves across VLM agents. The proposed framework provides a foundation for building reliable and resource-efficient agent teams.
Sources
- Personal Visual Memory from Explicit and Implicit Evidence
- When History Is Multimodal: Rethinking Context Management for Long-Horizon Agents
- Scaling Video Understanding via Compact Latent Multi-Agent Collaboration
- Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems
- Seeing Before Agreeing: Aligning Multi-Agent Consensus with Visual Evidence
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- INMS: Memory Sharing for Large Language Model based Agents
- Self-Evolving Multi-Agent Systems via Decentralized Memory
- Memory in the Age of AI Agents
- MemGPT: Towards LLMs as Operating Systems
- Collaborative Memory: Multi-User Memory Sharing in LLM Agents with Dynamic Access Control
- MemCollab: Cross-Model Memory Collaboration via Contrastive Trajectory Distillation
- Multi-Agent Memory from a Computer Architecture Perspective: Visions and Challenges Ahead
- MARDoc: A Memory-Aware Refinement Agent Framework for Multimodal Long Document QA
- GraphMemix: Query-Aware Evidence Forests for Long-Term Multimodal Agent Memory
- Beyond Retrieval: Analytic Memory for Multimodal Agents
- SlideBank: A Persistent Hierarchical Evidence Bank for Consistent Whole-Slide Reasoning
- MemLens: Benchmarking Multimodal Long-Term Memory in Large Vision-Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection