delta-mem: Efficient Online Memory for Large Language Models

arXiv:2605.12357 · cs.AI · Submitted 2026-05-12 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "delta-mem: Efficient Online Memory for Large Language Models".

Jane: δ-mem proposes a lightweight memory mechanism that augments frozen full-attention backbones with a compact online state of associative memory to dynamically maintain and steer historical information during generation.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, we're diving into the paper today, "delta-mem: Efficient Online Memory for Large Language Models." It sounds like they’ve tackled the big problem of how to keep track of long-term information in these large models without just bloating the context window every time.

Jane: Exactly, Tom. The authors are Jingdi Lei and their team, and what's interesting is their focus on making this memory mechanism really lightweight by augmenting a frozen full-attention backbone rather than replacing it entirely.

Lu: I find it fascinating that they are working with a fixed-size state matrix to encode historical context, which suggests a very structured way of handling associative memories, and the authors are coming from institutions like Nanyang Technological University and Fudan University.

Meng: From an engineering standpoint, focusing on augmenting a frozen backbone sounds smart because it means we don't have to re-train the massive attention layers every time we add memory capabilities; it's more modular.

Lalam: I’m really looking forward to seeing how this structure translates into something that can genuinely improve the culture of how these assistants interact with users, making their long-term memory feel more persistent and useful.

The paper's summary: Tom: The core idea behind "delta-mem: Efficient Online Memory for Large Language Models" is to compress past information into a compact online state of associative memory, which they call OSAM, and then use that state to generate adjustments directly into the backbone’s attention computation while the model is generating new text.

Jane: To put it simply, instead of feeding everything back in every time, the system continuously updates this fixed-size matrix by learning delta rules as new tokens arrive, allowing it to steer how the model pays attention to past information.

Lu: The paper characterizes existing memory methods into three types: textual memory mechanisms that use context injection but suffer from context limits and noise, outside-channel modules that introduce integration complexity, and parametric mechanisms which are efficient but lack adaptability.

Meng: So they are proposing a new way to handle the state—it's not just storing text or putting things in external modules; it’s a dynamic online state updated by delta-rule learning that directly influences the attention calculation itself.

Lalam: That sounds incredibly elegant because if you can modify the attention computation during generation, you are fundamentally changing how the AI reasons about history in real time, which is what we want for truly intelligent agents.

The paper's improvements: Tom: The paper highlights a few specific improvements they’ve achieved with delta-mem. They show that with just an eight by eight online memory state, the average score improves by a factor of one point one zero times over the frozen backbone and even better results against the strongest non-delta-mem baseline at one point one five times.

Jane: That quantitative improvement is striking, Tom; it shows that this compact representation isn't just theoretically sound but actually delivers tangible performance gains on memory-heavy benchmarks like HotpotQA and LoCoMo.

Lu: What I find particularly interesting is how they address the steering aspect by projecting the read signal into both query-side and output-side corrections, modifying the query vector before it goes through the attention mechanism.

Meng: That’s a clever way to inject information without requiring a massive architectural overhaul; it uses low-rank corrections to steer attention, which keeps things computationally manageable for deployment.

Lalam: I think that ability to recover part of missing multi-hop evidence even when explicit context is absent, as they showed on HotpotQA and LoCoMo, means this system could make long-horizon agents much more reliable in real-world scenarios.

Conclusion: Tom: So, to wrap up the discussion on "delta-mem: Efficient Online Memory for Large Language Models," the main point is that this mechanism successfully augments a frozen full-attention backbone with a compact online state of associative memory that dynamically maintains and steers historical information during generation.

Jane: It really boils down to using delta-rule learning to update this fixed-size matrix, which provides an effective interface for accessing history without the huge overhead of explicit context extension or needing completely new backbone architectures.

Lu: The paper addresses a lot of existing limitations by offering a specific memory state and steering mechanism that keeps the memory compact and dynamically updated across interactions, which is a key challenge in this area.

Meng: From an engineering perspective, maintaining a fixed-size representation with very low trainable parameter counts makes it practical for deployment in long-running assistants without significant latency penalties during inference.

Lalam: I think the biggest impact here is that we are moving toward systems where historical context isn't a luxury but something that the AI can continuously access and refine its reasoning based on its entire history.

Nanyang Technological University

cs.AI

Submitted: 2026-05-12

Updated: 2026-09-28

Importance score: 92/100

The gist: δ-mem proposes a lightweight memory mechanism that augments frozen full-attention backbones with a compact online state of associative memory to dynamically maintain and steer historical information

Key concepts

δ-mem
A mechanism that augments a frozen full-attention backbone with a compact online state of associative memory. It dynamically maintains and steers historical information during text generation by using the memory's readout to correct the model's attention computation.
αεAM
Online State of Associative Memory (OSAM). This is a fixed-size matrix representation that continuously updates as new tokens arrive. It stores historical information in a compact format, allowing for efficient retrieval and steering of the model's reasoning without needing to store all past text explicitly.
δε-rule learning
The process by which the online state is continuously updated. New information is incorporated via delta-rule learning, which allows the model to maintain useful historical data in a fixed-size matrix representation. This ensures that past context remains relevant and integrated into the current generation.
δε steering attention
The process where associative memory signals influence backbone reasoning. The memory state is projected into query-side and output-side corrections, which are then used to modify the model's query and compute a corrected attention output. This steers the model's focus based on stored historical information.

Terminology

Summary

δ-mem proposes a lightweight memory mechanism that augments frozen full-attention backbones with a compact online state of associative memory to dynamically maintain and steer historical information during generation. This approach is significant because it enables effective context utilization in long-term assistants and agent systems without the high costs of explicit context extension or backbone replacement.

The gist

δ-mem compresses past information into an online state of associative memory (OSAM), which is continuously updated via delta-rule learning, and uses its readout to generate low-rank corrections to the backbone’s attention computation during generation.

Memory Mechanism Characterization

Existing memory mechanisms are characterized by two dimensions: memory state, defining how historical information is stored, and memory steering, determining how stored information influences backbone reasoning. The paper categorizes prior methods into three paradigms:

  1. Textual Memory Mechanisms (TMMs): These store memory as text and inject it through the input context, but suffer from context-window limits, retrieval noise, and inevitable compaction loss.

  2. Outside-channel Memory Mechanisms (OMMs): These keep memory in external modules and interact with the backbone via retrieval or encoding on outside pathways, introducing overhead, integration complexity, and potential misalignments with the backbone.

  3. Parametric Memory Mechanisms (PMMs): These encode memory into parameters of prefixes or adapters, which are efficient but have a static nature [that] limits adaptation to dynamically evolving information.

δ-mem Architecture and Dynamics

δ-mem maintains a compact and dynamically updated memory state alongside a frozen full-attention backbone. The core dynamics are governed by the following processes:

  1. Online State Update: The state is continuously updated via delta-rule learning as new tokens arrive, allowing the model to maintain useful historical information in a fixed-size matrix representation of associative memories.

  2. Reading from State: Before writing current information, δ-mem reads from the old state using a projection where the read vector is defined as rt = St−1q m t, noting that this step's cost is independent of the history length.

  3. Steering Attention: The associative memory signals steer attention through two lightweight linear mappings. The read signal is projected into query-side and output-side corrections, where the query is modified as q˜t = q 0 t + α r ∆qt, and the attention output is computed using this corrected query: at = Attn(q˜t, K≤t, V≤t).

  4. Writing into State: The current information is written using a dimension-wise gated delta-rule update: St = Diag(λt)St−1 + Diag(βt) (v m t − St−1k m t)(k m t)⊤, where λt controls retention and βt controls the strength of the residual write.

Writing Granularity Strategies

The mechanism supports three writing strategies to balance fidelity and stability:

  1. Token-State Write (TSW): Updates the online state at each token position, preserving the finest-grained information but being sensitive to format symbols, repeated expressions, and short-term noise.

  2. Sequence-State Write (SSW): Averages hidden states within a message segment first (x¯(j) = 1/M(j) Σ X t∈M(j) xt) and updates the state once per message to reduce redundant writes and smooth the state evolution.

  3. Multi-State Write (MSW): Decomposes memory into multiple parallel sub-states, allowing different types of information to accumulate in separate states, which helps in reducing mutual interference within a single state.

Empirical Evaluation and Contributions

δ-mem was evaluated on memory-heavy benchmarks including HotpotQA, LoCoMo, and MemoryAgentBench. The results showed that with only an 8 × 8 online state of associative memory, δ-mem improves the average score by 1.10× over the frozen backbone and outperforms the strongest non-δ-mem baseline by 1.15×. Furthermore, in a no-context setting, δ-mem consistently improved performance on HotpotQA and LoCoMo, suggesting that the online state can recover part of the missing multi-hop evidence even when explicit context is unavailable. The paper concludes that this compact online state provides an effective associative memory interface without relying on extending explicit context or heavy external retrieval modules.

Contributions Summary

The main contributions include:

**: Proposing δ-mem, a mechanism that augments a frozen full-attention backbone with a compact online state of associative memory to enable historical information to be dynamically maintained and directly coupled with the backbone’s attention computation. 2.

Improvements for AI systems

As a fastidious researcher, I have analyzed the provided paper, δ-mem: Efficient Online Memory for Large Language Models. The core innovation lies in introducing a lightweight, dynamic online state of associative memory (OSAM) that directly steers the frozen attention backbone via low-rank corrections.

Here are the specific improvements and capabilities this system enables:


The improved AI system leverages δ-mem to achieve the following advancements:

  1. A robust mechanism for maintaining long-term, dynamically updated historical context without requiring expensive full fine-tuning or massive context window expansion.

  2. Efficient, low-overhead memory steering that directly influences the attention computation during generation, ensuring memory is utilized effectively at test time.

The improved AI system can specifically:

  1. Perform complex, multi-turn reasoning tasks (e.g., long-horizon agent systems) where the model must continuously accumulate and reuse historical information across extended interactions (as demonstrated by the performance gains on MemoryAgentBench).

  2. Recover context-relevant information even when explicit historical context is removed from the input, demonstrating resilience against context rot or memory degradation when dealing with very long sequences.

  3. Execute sophisticated knowledge-intensive Question Answering (QA) and instruction-following tasks by dynamically steering the attention mechanism using learned associative signals derived from a fixed, compact online state (the 8x8 matrix).

  4. Maintain high performance across different backbone model sizes, showing that the memory mechanism scales effectively with varying reasoning capacities (e.g., benefiting smaller models through Multi-State Write (MSW) and larger models through Sequence-State Write (SSW)).

  5. Operate with minimal computational overhead: δ-mem introduces very low trainable parameter counts (0.12% to 0.48% of backbone parameters) and maintains comparable GPU memory usage to vanilla or other context-augmented methods, making it practical for deployment in long-running assistants without significant latency penalties during inference.

  6. Handle scenarios requiring varying levels of memory granularity: the system can be configured (via TSW, SSW, or MSW) to prioritize either fine-grained token updates, segment averaging over messages, or multi-state organization for different types of historical data (e.g., facts vs. preferences).

Sources

Related papers