SRMT: Shared Memory for Multi-agent Lifelong Pathfinding

arXiv:2501.13200 · cs.LG, cs.AI, cs.MA · Submitted 2026-08-24 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "SRMT: Shared Memory for Multi-agent Lifelong Pathfinding".

Jane: The paper was written by Alsu Sagirova, Yuri Kuratov and Mikhail Burtsev from AIRI, Moscow, Russia and Neural Networks and Deep Learning Lab, MIPT, Dolgoprudny, Russia and London Institute for Mathematical Sciences, London, UK.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Summary: Tom: Last time we were talking about how crucial that shared memory concept is, and Jane was explaining it in simple terms. Now we’re looking at the summary of "SRMT: Shared Memory for Multi-agent Lifelong Pathfinding," which really highlights how this system operates over time.

Jane: The key thing I gathered from the summary is that SRMT isn't just good at finding a path; it's *lifelong* pathfinding, meaning it adapts when the goal or the environment changes dramatically.

Tom: And that adaptation capability, Jane, seems to be where the complexity really ramps up. If agents are constantly changing goals or facing unexpected obstacles, how does this shared memory handle that?

Jane: It seems to manage it by maintaining a consistent representation of spatial relationships and historical interactions across all agents in the system.

Lu: The "lifelong" aspect implies that the knowledge gained from an early episode—say, navigating a crowded corridor—doesn't just vanish when the episode ends; it informs how they tackle a completely different task later on.

Meng: From my viewpoint, this is huge because most current pathfinding models are brittle; if you change one parameter in the simulation, they often fail entirely. SRMT suggests robustness through continuity of knowledge.

Lalam: That continuity translates into culture because real-world teams don't wipe the slate clean every Monday morning; they build on previous successes and failures.

Tom: So, to circle back to that summary point: the agents can learn from an entire sequence of events, which is fundamentally different from just solving a single maze problem.

Jane: Exactly! They’re not just finding A to B; they’re learning how the *group* moves through the environment over time, which is much richer data.

Lu: And this ability to generalize across vastly different scenarios suggests a leap towards true common-sense reasoning in AI, not just pattern matching.

Meng: If this system can maintain state and adapt goals robustly, we could see it applied to complex logistical planning, like emergency response teams coordinating movements.

Lalam: That capability of persistent learning makes the resulting AI feel more like a team member and less like a single tool—it improves our reliance on AI as a true partner.

Improvements: Tom: We've talked about the conceptual framework, and now we're looking at the improvements suggested by "SRMT: Shared Memory for Multi-agent Lifelong Pathfinding." This section seems to bring in some really concrete evidence, particularly around scalability.

Jane: The visual data, like Figure ten showing scalability on MovingAI maps from POGEMA benchmark, is really striking. It clearly demonstrates how the system handles increasing numbers of agents.

Tom: Looking at that figure with the throughput—going from four point zero zero to two point zero zero to one point zero zero as the number of agents increases—it seems they are quantifying exactly how much overhead is managed by this shared memory approach compared to other methods.

Jane: It suggests a very efficient scaling curve, which is critical because real-world deployments never have a fixed, small number of agents; they fluctuate wildly.

Lu: And let's not forget Figure eleven which tackles the memory representation itself. The alignment between SRMT distances in memory and the actual Euclidean distances on the map is mathematically compelling.

Meng: That alignment isn't just a cool graph; it proves that the internal memory structure—the cosine distance—is actually reflecting physical reality in a meaningful, measurable way.

Lalam: That reliable mapping between abstract knowledge and physical space is what gives the AI its grounding, making its decisions predictable and trustworthy for human operators.

Tom: So, to build on that memory alignment idea from Figure eleven: when the agents are moving in a straight line or keeping a constant distance, the cosine distance decreases steadily. What does that tell us about how SRMT models motion?

Jane: It shows that the shared memory is successfully encoding relative spatial positioning—it knows not just *where* an agent is, but *how far* it is from others in a predictable manner.

Lu: And when they face each other, the memory representation captures that moment of direct interaction, which must be a highly complex piece of information for the network to encode accurately.

Meng: The fact that it can model separation—

Paper discussion segment 3: Tom: So, we’ve seen that SRMT is a big step forward in decentralized pathfinding by allowing agents to share information implicitly through this shared memory architecture. The real question now is what kind of practical improvements does this offer over other methods?

Jane: Essentially, the biggest improvement is how robust the system is. Instead of failing when things get complicated or adapting to a massive map, SRMT maintains its performance even when the corridor length gets huge.

Meng: That robustness translates directly into reliability for my team. If we’ are deploying this kind of AI in a real-world warehouse, we can't have it crash just because the operational space is bigger than what was used during training.

Lu: And I see that resilience as a gateway to completely new creative solutions, Meng. Because the knowledge persists across multiple sessions, it suggests that we can design complex dynamic scenarios where agents learn from historical interactions and solve entirely different problems using that same learned knowledge base.

Lalam: It’s not just about solving a single task; it's about building a shared understanding of continuous movement and adaptation. This ability to keep learning across time fundamentally improves how we view collaboration, moving from one-off instructions to genuine, persistent teamwork.

Tom: Lalam makes that sound powerful—the idea of continuous teamwork. The paper shows that even when rewards are extremely sparse or negative, which usually causes other models to give up, SRMT keeps its performance high.

Jane: That’s the generalization aspect in action. It isn't just memorizing a path; it's understanding the *principle* of maintaining coordination regardless of the environment.

Meng: From an engineering standpoint, that means less need for complex, manual reward-shaping in real life. The system has learned to find its own functional rewards internally.

Lu: And when we consider Lifelong MAPF—the lifelong part—we’re talking about a continuous learning loop where the AI is not just solving problems but evolving its internal strategy over time.

Lalam: That evolution, that persistent growth, means our future systems won't just execute code; they will genuinely improve their collaborative ability with every single step.

Tom: It certainly opens up possibilities far beyond the initial scope of a simple maze. Let's look at how this shared memory translates into real-world coordination challenges...

Conclusion: Tom: So, wrapping up our discussion on SRMT: Shared Memory for Multi-agent Lifelong Pathfinding, it's clear that this method is really pushing how we think about coordination and memory in complex environments.

Jane: Exactly. What I keep remembering is how much the shared memory aspect helps maintain performance even when the environment changes or gets much longer, which is such a huge leap over previous models.

Lu: But Jane's right, because this isn't just about pathfinding anymore; think about any scenario where multiple autonomous units need to cooperate over a long duration and adapt to varying conditions—this architecture could be fundamental.

Meng: I agree with Lu on the fundamentals, but practically speaking, if you want to deploy something like this in a real industrial setting, we need to know how efficiently that shared memory scales with hundreds of distinct agents simultaneously.

Tom: That's a crucial point, Meng; the stability and efficiency of that shared memory must be robust enough for massive real-world deployments.

Jane: And what’s amazing is how it handles those transitions, like when one agent finishes its goal and the others have to adjust their whole strategy immediately.

Lu: It suggests a paradigm shift away from pre-scripted, brittle coordination rules and toward genuinely emergent, shared understanding among agents.

Meng: Emergent understanding is great for theory, but I'm more focused on the hardware abstraction—if we could map this memory representation onto something like a distributed ledger or a highly optimized edge computing fabric, the commercial impact would be massive.

Lalam: Thinking about the wider cultural impact of this capability, it suggests that future AI systems won't just be tools for individual tasks; they'll become collaborative intelligence layers, improving how humans interact with complex automated infrastructure.

Tom: It really shows the power of combining robust memory structures with multi-agent learning, doesn't it?

Jane: It’s definitely a major piece of work that sets a new bar for cooperative AI planning.

Lu: I can't wait to see what kind of creative limitations we can push with this level of shared understanding in future research.

Meng: We definitely need to follow up on the computational complexity details; that's where the engineering rubber meets the road.

Lalam: And it inspires a whole new chapter in how we design human-AI partnerships, moving toward true co-intelligence.

Tom: All right, folks, what a discussion! Before we sign off on "SRMT: Shared Memory for Multi-agent Lifelong Pathfinding," you can tell our listeners that the concept of shared memory is fundamentally changing the trajectory of multi-agent AI. We'll be back next week to talk about…

Alsu Sagirova, Yuri Kuratov, Mikhail Burtsev

AIRI, Moscow, Russia · Neural Networks and Deep Learning Lab, MIPT, Dolgoprudny, Russia · London Institute for Mathematical Sciences, London, UK

cs.LG, cs.AI, cs.MA

Submitted: 2026-08-24

Updated: 2026-08-25

Comments: 16 pages, 11 figures

Code: https://github.com/Aloriosa/srmt

License: http://creativecommons.org/licenses/by-sa/4.0/

Importance score: 89/100

The gist: The study focuses on Multi-agent Pathfinding (MAPF) and evaluates the performance and scalability of the SRMT architecture.

Key concepts

Lifelong Pathfinding
This concept means the system does not just solve a single problem and forget it. Agents learn from an entire sequence of events, allowing knowledge gained in one task to inform future actions across vastly different scenarios.
Shared Memory
This architecture allows all agents in the system to implicitly share information about spatial relationships and historical interactions. It maintains a consistent, collective representation of the environment that is used by every agent.
Robustness and Scalability
SRMT maintains high performance even when dealing with extremely large environments or a fluctuating number of agents. The system does not fail simply because the operational space is larger than what was used during initial training.

Terminology

Summary

The study focuses on Multi-agent Pathfinding (MAPF) and evaluates the performance and scalability of the SRMT architecture.

Training Details and Methodology:

The environments used for testing were created with the POGEMA framework, and policy model training utilized the Sample Factory codebase. The training parameters for all tested methods are detailed in Table 1. For model training, a single Tesla P100 was employed for approximately one hour per policy model. Results derived from models trained with Sparse and Dense reward functions were averaged over 10 runs using different random seeds, while results from policies trained with Directional and Directional Negative rewards were averaged over 5 runs due to their observed lower variation during training. Furthermore, each run evaluation was initially averaged over 10 different evaluation procedure random seeds. The authors conducted a grid search for the SRMT training entropy coefficient (ranging from [0.00001, 0.0003]) and the learning rate (ranging from [0.01, 0.05]).

Reward Functions Tested:

The evaluation included testing multiple reward functions for the Bottleneck MAPF task, listed in Table 2:

  • Directional: Rewards were set to +1 for achieving the goal, +1 for moving toward the goal, and +1 for other actions.

  • Sparse: The rewards were set to +1 on goal, and specific values (e.g., 0.005) for movement towards the goal, with-0.01 otherwise.

  • Dense: Rewards included a high value (+1) on goal, a smaller positive value (+0.005) for moving toward the goal, and -(-0.01) for other actions.

  • Directional Negative: This function assigned +1 on goal, -0.01 for moving towards the goal, and 0 otherwise (with specific negative values listed for movement/holding).

Performance Evaluation and Scalability:

The evaluation results provide performance metrics for Dense, Directional, and Directional Negative reward functions. The findings indicate that SRMT has superior or comparable performance compared to the baselines. Specifically:

  • When trained with a Dense reward function, SRMT consistently outperforms baselines both in success rates and in the time needed to solve the task, particularly as the corridor length increases, as shown in Figure 7.

  • Training with Directional reward functions resulted in all tested methods preserving the scores for all tested corridor lengths, according to Figure 8.

  • Using Directional Negative rewards demonstrated that Vanilla attention fails to scale at corridor lengths of more than 400, compared to the SRMT which preserves the highest scores, which is presented as proof of the sufficiency of the proposed SRMT architecture (Figure 9).

  • Regarding scalability on MovingAI maps from the POGEMA benchmark, Figure 10 shows that both SRMT models consistently outperform cooperative baselines (MAMBA and QPLEX), even when evaluated with a mixture of 64 and 128 agents, compared to the baseline methods.

Memory Analysis:

An exploration was conducted regarding the relationship between the SRMT agents’ memory representations and the spatial distances between agents on the map. Figure 11 illustrates that SRMT distances between memory representations are aligned with distances between agents for different corridor lengths. The analysis details how, at the start of an episode, the agents move closer to each other quickly, and the respective cosine distances decrease significantly. Subsequently, when agents face each other in the environment (marked by a triangle), they move in the same direction while maintaining a constant spatial distance. Finally, after one agent reaches its goal and disappears (marked by a star), the remaining agent moves away to reach its goal, which is depicted as an increasing memory distance at the end of the episode.

Improvements for AI systems

1. Dynamic, Predictive Multi-Agent State Estimation via Memory Embedding:

The system should be improved by explicitly leveraging the demonstrated fidelity between internal memory representation distances (cosine distance) and physical Euclidean distances (Fig. 11). Instead of treating the memory vector solely as an observation encoding, it must be integrated into a predictive module.

  • Improvement: Implement a Geometric State Predictor Module (GSPM) that takes the current memory state m t and outputs not just a predicted next location p t+1, but also an estimated covariance matrix t+1 describing the expected spatial uncertainty of the entire agent group at t+1.

  • Improved Capability: The AI system can move beyond reactive pathfinding to proactive, uncertainty-aware planning. It can calculate optimal trajectories that minimize the predicted collision probability (based on t+1) even when agents are not visible or when environmental dynamics introduce noise. This is critical for deployment in real-world, partially observable environments like autonomous vehicle platooning or complex warehouse logistics.

2. Adaptive Reward Function Synthesis and Meta-Learning:

The observed superior performance across vastly different reward structures (Sparse, Dense, Directional, Directional Negative) suggests that the core RL policy should not be brittle to reward engineering changes.

  • Improvement: Develop a Reward Policy Meta-Learner (RPML). This module operates before the main RL training loop and is trained on a diverse set of simulated reward landscapes (e.g., varying weights for goal achievement vs. path adherence). RPML learns to synthesize an optimal, composite reward signal R composite that maximizes sample efficiency and stability given the specific task constraints provided at runtime.

  • Improved Capability: The system can be deployed in Zero-Shot Reward Adaptation. When faced with a novel coordination task (e.g., rescue operations where the goal is ambiguous or rewards are sparse), the RPML analyzes initial low-fidelity feedback and rapidly generates a highly effective, composite reward structure that guides the agents to optimal behavior without requiring manual hyperparameter tuning or extensive retraining.

3. Hierarchical Scalable Coordination with Agent Role Decomposition:

The scalability demonstrated in Figure 10 (maintaining performance when agent numbers increase) needs to be formalized into a hierarchical control structure suitable for massive deployments.

  • Improvement: Integrate a Role Assignment and Decomposition Layer (RADL) above the core MAPF policy. This layer dynamically partitions the N agents into functional groups based on the current task state (e.g., Scout Group, Support Group, Primary Pathfinders). Each group is assigned a specialized, reduced-dimensional sub-task objective, and the main policy learns to coordinate the inter-group communication and handoffs between these roles rather than coordinating every individual action simultaneously.

  • Improved Capability: The AI system can manage Hyper-Scale Coordination in Dynamic Swarms. Instead of simply scaling performance up to N agents, it scales the complexity of the coordination problem. For instance, in disaster relief involving hundreds of drones, the system automatically delegates local obstacle avoidance to low-level controllers while reserving the high-dimensional central memory network only for critical strategic decisions (e.g., identifying optimal search sectors or resource allocation).

4. Causality-Informed Policy Regularization:

The success of SRMT suggests strong internal representations, but true robustness requires ensuring that the learned correlations are causal, not merely correlational artifacts of the training environment.

  • Improvement: Introduce a Causal Intervention Regularizer (CIR) during policy optimization. This regularizer penalizes the policy gradient when it relies on features that are statistically correlated with success but cannot causally influence the outcome (e.g., relying too heavily on a stable background element rather than the agent's own action). The CIR encourages the model to learn invariant causal mechanisms of coordination.

  • Improved Capability: The system achieves Extreme Domain Transferability. By disentangling correlation from causation, the resulting policy can be transferred to entirely new physical domains (e.g., moving from simulated indoor corridors to real-world traffic networks) with minimal fine-tuning, as the underlying principles of efficient interaction are robustly encoded, not just the pixel patterns of the training map.

Abstract

Coordination in decentralized multi-agent reinforcement learning (MARL) necessitates that agents share information about their behavior and intentions. Existing approaches rely on communication protocols with domain or resource constraints or centralized training that poorly scales to large agent populations. We introduce the Shared Recurrent Memory Transformer (SRMT), which enables coordination through unconstrained communication. SRMT provides a global memory workspace where agents broadcast their learned working memory states and query others' memory representations to exchange information and coordinate while maintaining decentralized training and execution. We evaluate SRMT on the Partially Observable Multi-Agent Pathfinding (PO-MAPF) problem, where coordination is vital for optimal path planning and deadlock avoidance. We demonstrate that shared memory enables emergent coordination even when the reward function provides minimal or no guidance. On the specifically constructed Bottleneck task that requires negotiation, SRMT consistently outperforms communicative and memory-augmented baselines, particularly under sparse reward signals, and successfully generalizes to longer corridors unseen during training. On POGEMA maps, SRMT scales with the increasing agents' population and map size, achieving competitive performance with recent MARL, hybrid, and planning-based methods while requiring no domain-specific heuristics. These results demonstrate that a transformer with shared recurrent memory enhances coordination in decentralized multi-agent systems. The source code for training and evaluation is available on GitHub: https://github.com/Aloriosa/srmt.

Sources

Related papers