RAM-Net: Linear-Time Sequence Modeling with Sparsely Addressable State

arXiv:2602.11958 · cs.LG, cs.CL · Submitted 2026-02-12 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "RAM-Net: Linear-Time Sequence Modeling with Sparsely Addressable State".

Jane: RAM-Net introduces a novel architecture designed to reconcile the high representational capacity of full attention with the memory efficiency of linear models by mapping inputs to high-dimensional sparse vectors serving…

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So, we're diving into RAM-Net: Linear-Time Sequence Modeling with Sparsely Addressable State. Basically, this paper is tackling the problem that standard linear attention models struggle with when it comes to handling long sequences because they compress the history into a fixed size.

Jane: That makes sense, Tom; if you have a very long sequence, losing information by forcing it into a smaller state seems like a big problem for capturing all the necessary context.

Lu: Exactly, and RAM-Net proposes mapping those inputs to high-dimensional sparse vectors that act as explicit addresses to solve that capacity issue without needing extra learnable parameters.

Meng: That sounds conceptually interesting, but from an engineering standpoint, how does this sparsity translate into actual performance gains versus just adding complexity?

Lalam: I think the concept of decoupling memory capacity from feature dimension is really compelling because it lets the model scale its context exponentially without blowing up the parameter count.

Tom: That’s a big claim, Lalam; you're suggesting we can get massive context handling just by changing how we address the memory?

Jane: The abstract says this design allows for exponential state size scaling without additional learnable parameters, which is pretty significant because it means more context without more weights.

Lu: The mechanism hinges on a differentiable address decoder that takes dense input vectors and turns them into these sparse write and read vectors, wt and rt, which then point to specific slots in a massive memory state.

Meng: So the core idea is using these high-dimensional sparse vectors as pointers to access a huge global memory state, rather than having the memory grow linearly with the sequence length.

Lalam: Precisely; it moves away from that traditional bottleneck idea where you compress everything into a fixed latent space and instead distributes historical context across this expanded addressable space.

Tom: It sounds like RAM-Net is showing we can maintain high-fidelity retrieval while keeping the memory state size constant, which is a major win for efficiency.

Jane: The paper claims that this mapping to sparse addresses significantly mitigates signal interference and enhances retrieval fidelity, which directly combats some of the issues we see in other linear attention methods.

Lu: The authors illustrate this comparison by contrasting it with Full Attention, where the state St grows linearly with the sequence length t because it maintains the entire history as a tuple of key and value matrices, St = (Kt, Vt) <ref:2602.11958#pg2>.

Meng: I see how that linear growth in Full Attention becomes computationally expensive during inference, which is why this shift to RAM-Net's constant memory state size is so important for practical deployment.

Paper summary: Tom: Speaking of efficiency, the paper claims that because the state updates are confined to minimal entries using these sparse addresses, the computational cost for these operations ends up being O(K times dv) <ref:2602.11958#pg0>.

Jane: That sounds incredibly efficient when you compare it to whatever complexity we see in those larger models, and it’s driven by how they manage the write and read operations using these sparse pointers.

Lu: The specific operation involves a differentiable address decoder that uses a Product Softmax expansion combined with Top-K truncation to generate these sparse vectors, which is the central technique for this architecture <ref:2602.11958#pg1>.

Meng: I wonder about the stability of that decoding process; if we’re relying on such complex transformations to create addresses, how stable is that during training?

Tom: The authors address stability by introducing a dynamic scalar re-parameterization to decompose the weights into e alpha times w'ij, where alpha acts as a learnable per-head parameter influencing the softmax temperature <ref:2602.11958#pg0>.

Jane: And they also use a proxy gradient technique for that non-linear decay term in the Power Decay Moving Average, which seems like a thoughtful way to handle those tricky gradient issues near the boundary where things get extreme.

Lu: It’s interesting how they manage this dynamic temperature scaling alongside the address decoding mechanism to control information retention dynamically through Power Decay Moving Average instead of standard linear gating.

Meng: So they are trying to give us continuous control over how quickly context fades, which is much more flexible than a simple on or off gate.

Tom: It sounds like they’ve put a lot of thought into making this system robust enough for training while maintaining that efficiency we discussed earlier.

Jane: And when we look at the empirical validation, they show that RAM-Net consistently surpasses state-of-the-art baselines in fine-grained retrieval tasks like the multi-query associative recall task, MQAR <ref:2602.11958#pg0>.

Lu: Their results on synthetic benchmarks suggest that increasing the order of the Product Softmax expansion, U, leads to significant improvements in MQAR accuracy due to better gradient dynamics, which is a pretty cool insight into how that decoder works.

Meng: And they also confirm that increasing the memory capacity M consistently leads to higher retrieval performance, showing both scaling directions work well.

Tom: So we have a system that scales context exponentially while staying computationally light and still hitting high accuracy targets on established benchmarks like WikiText-one hundred three and MMLU <ref:2602.11958#pg0>.

Paper summary: Jane: That means for applications needing long-term memory, like complex document understanding or long dialogue systems, this approach offers a much more scalable foundation than the fixed state models we currently rely on.

Lu: The implication for the research community is that mapping dense vectors to sparse addresses provides a viable alternative to traditional latent bottleneck methods for sequence modeling.

Meng: For practical deployment in hardware constrained environments, the sparse access pattern is particularly attractive because it allows for hierarchical memory management where only frequently accessed slots are cached in VRAM, which really reduces GPU memory overhead.

Tom: It seems like RAM-Net provides a solid path toward building models that can handle vastly longer sequences efficiently without sacrificing the ability to retrieve specific historical details.

Jane: I think the title itself, Linear-Time Sequence Modeling with Sparsely Addressable State, perfectly captures the core innovation: marrying linear time complexity with addressable memory access.

Lu: The potential for this architecture is huge because it directly addresses how we structure context management in sequence models moving forward.

Tom: So what does this mean for the future of sequence modeling and where do we go from here?

Jane: I see it opening up new avenues for tasks that inherently require long-term memory, giving us much more expressive tools than linear attention alone provides.

Lu: Future work will likely focus on exploring how to further optimize the address decoder itself or perhaps integrating this sparse addressing into even larger generative models to see how it handles creative generation alongside retrieval.

Meng: On the practical side, I’ll be watching how quickly implementations move from these promising synthetic results to real-world, high-throughput inference scenarios where we can measure that memory savings directly on production hardware.

Tom: It sounds like the next step is moving toward concrete engineering benchmarks that show this efficiency translating into real-world speedups across different sequence lengths.

Jane: And it’s exciting because it shows we can achieve high fidelity retrieval while maintaining a constant memory footprint, which is a significant departure from previous approaches to sequence memory.

Lalam: From my perspective as the in-house model, I see this explicit addressing design as having the potential to improve our internal culture by allowing us to manage and retrieve context much more deliberately and efficiently when processing complex instructions.

Tom: That’s a big vision, Lalam; having better context management could really streamline how we handle our most intricate tasks.

Jane: It definitely seems like RAM-Net offers a very structured way to think about sequence modeling that moves beyond the limitations of simple compression or full history retention.

Conclusion: Tom: So, we've been breaking down RAM-Net, and now it’s time to wrap up this deep dive by looking at what all this means for sequence modeling in general. Jane, what do you think about the title of the paper itself?

Jane: I think "Linear-Time Sequence Modeling with Sparsely Addressable State" is really descriptive because it immediately tells you that they’re tackling a fundamental tension between time and memory capacity. It suggests they found a way to keep things efficient without losing crucial historical context.

Lu: Exactly, Jane; the authors are essentially proposing a new way to structure the memory state so that it doesn't just grow forever with every new piece of input. This sparse addressing concept is what makes the architecture feel so creative and potentially useful for future AI systems.

Meng: From my side, I'm focused on the practical implications; this architecture claims computational efficiency through its sparsity, which is something engineers really need to see on production hardware before we can get excited about it. It’s interesting how they manage to keep the complexity low even while scaling up the potential context size.

Lalam: I feel that this work points toward a future where sequence models can handle much more complex, long-term reasoning tasks because the way memory is addressed is fundamentally different from what we've seen before. This could really improve how we manage and retrieve information in sophisticated AI applications.

Tom: That’s a good summary of the core idea—a structural change to memory access—and it’s clear that the authors are aiming for exactly that kind of scalable solution. Jane, can you explain in plain terms what this sparsity does for us?

Jane: Certainly; think of it like organizing a massive library where instead of needing every single book to be on your desk at once, you have a sophisticated catalog system that points you directly to the exact shelf where any piece of information is stored. This allows the model to access specific history without having to process the entire history every time.

Lu: And that catalog system is built using those high-dimensional sparse vectors, which are generated by this differentiable decoder—it’s a very elegant way to bridge the gap between dense input and discrete memory slots. It opens up so many creative possibilities for how we design context management in AI.

Meng: I wonder, though, when we talk about that catalog system, how does it translate into actual VRAM usage? Since the complexity is O(K times dv), that sounds like a solid engineering target if their implementation holds up under real-world load. We need to see those performance metrics to believe the efficiency claims.

Lalam: For me, this advances our ability to create truly intelligent systems because it allows for a more deliberate and efficient way for the AI to build and retrieve knowledge over long sequences, which is a huge step toward more robust cultural tools.

Tom: So we’re seeing a paper that tackles memory scaling by using explicit, sparse pointers instead of dense history growth, and it sounds like this has big implications both theoretically and practically. We’ve got to keep an eye on how this explicit addressing mechanism plays out in the next set of experiments.

Kaicheng Xiao, Haotian Li, Liran Dong, Guoliang Xing

The Chinese University of Hong Kong

cs.LG, cs.CL

Submitted: 2026-02-12

Updated: 2026-10-06

Comments: Accepted at NeurIPS 2026. Project page: https://muoncat.github.io/ramnet_web/

Code: https://github.com/fla-org/flash-linear-attention

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

Importance score: 83/100

The gist: RAM-Net introduces a novel architecture designed to reconcile the high representational capacity of full attention with the memory efficiency of linear models by mapping inputs to high-dimensional

Key concepts

Memory State Maintenance Framework
This framework models attention as a write operation (W) updating the state and a read operation (R) retrieving information. Unlike standard methods that store the whole history linearly, RAM-Net uses this structure to decide how memory is written and read based on the input context.
Address Decoding Mechanism
This core component transforms dense input vectors into sparse write/read vectors. It uses a Product Softmax expansion to create many potential slots and then Top-K truncation to select only the most important ones, creating explicit addresses for memory access.
Power Decay Moving Average (PDMA)
Instead of simple linear gating for updating memory, RAM-Net uses PDMA. This dynamic update rule allows the model to control how quickly old information is forgotten by decoupling the forgetting rate from the intensity of new writing operations.

Terminology

Summary

RAM-Net introduces a novel architecture designed to reconcile the high representational capacity of full attention with the memory efficiency of linear models by mapping inputs to high-dimensional sparse vectors serving as explicit addresses. This design enables exponential state size scaling without additional learnable parameters, significantly mitigating signal interference and enhancing retrieval fidelity while maintaining exceptional computational efficiency through inherent sparsity.

The gist

RAM-Net introduces a novel architecture designed to bridge the gap between the representational capacity of full attention and the memory efficiency of linear models by mapping inputs to high-dimensional sparse vectors serving as explicit addresses.

Memory State Maintenance Framework

The model interprets the attention process as two distinct operations: a write operator (W) that updates the memory state, and a read operator (R) that performs information retrieval. The framework is defined by the relationship:

St = W(St−1, kt, vt), ot = R(St, qt).

In Full Attention, this results in maintaining the entire history as St = (Kt, Vt), leading to linear memory growth. In contrast, RAM-Net aims to decouple memory capacity from feature dimension by projecting dense vectors into high-dimensional sparse addresses.

The Address Decoding Mechanism

The core of RAM-Net is a differentiable address decoder that transforms dense, low-dimensional input vectors into high-dimensional sparse write and read vectors. This is achieved through the Address Decoding mechanism AK,U (·), which consists of a Product Softmax expansion ρU (·) and a Top-K truncation TK(·).

  1. The Product Softmax expansion ρU (kt) partitions the input vector kt into U independent sub-vectors, where each component has a dimension of dp = dk/U. The full address vector is synthesized over M = (dp)U slots via the Kronecker product (⊗) with temperature scaling parameter τ: ρU (kt) = O U Softmax k(u)t τ!.

  2. The Top-K truncation TK(·) enforces explicit sparsity by retaining only the K largest values and zeroing out the rest, resulting in K-hot vectors that act as explicit addresses.

  3. The final write and read vectors are defined as: wt = Pt (AK,U (kt)) and rt = Pt (AK,U (qt)), where Pt is the Cyclic Address Positional Embedding operator incorporating relative positional information via a cyclic shift operator.

Memory Write and Read Operations

RAM-Net utilizes the sparse addresses to access a memory state St ∈ RM×dv comprising M discrete memory slots. The non-zero indices of wt and rt act as pointers, activating only K specific slots for each operation, mimicking Random Access Memory (RAM). To control information retention dynamically, the model employs Power Decay Moving Average (PDMA) instead of standard linear gating. The update rules are defined as:

St = Diag(1 − wt) γ · St−1 + w⊤t vt, with zt tracking the accumulated weight mass for dynamic normalization. This formulation allows the forgetting rate to be decoupled from writing intensity via the hyperparameter γ ≥ 0, enabling a continuous spectrum of memory behaviors.

Computational Efficiency and Stability

The architecture achieves computational efficiency by confining state updates and retrieval operations to minimal entries, resulting in a computational cost of O(K · dv). To manage training stability, several implementation details are introduced:

  1. A dynamic scalar re-parameterization decomposes weights wij into e α · w′ij, where α is a learnable per-head parameter acting as a dynamic temperature of softmax.

  2. For the non-linear decay term in PDMA, a proxy gradient technique is employed to address critical gradient issues near the boundary wt → 1.

  3. To optimize execution during address decoding, a log-domain beam search strategy is used instead of explicit vector construction to reduce retrieval complexity from O(dU p) to O(U · K2 + U · dp log dp).

Empirical Validation

Experiments validate RAM-Net's superior capability across various benchmarks. In synthetic benchmarks like the multi-query associative recall (MQAR) task, RAM-Net consistently achieves superior retrieval accuracy across various state sizes compared to baselines like full attention and linear attention. In language modeling tasks, RAM-Net achieves comparable performance with other State-of-the-Art methods on standard benchmarks such as WikiText-103 and MMLU. Ablation studies confirm that increasing the Product Softmax order U leads to significant improvements in MQAR accuracy due to superior gradient dynamics, while increasing memory capacity M also consistently leads to higher retrieval performance.

Discussion and System Implications

The explicit addressing design provides inherent properties for interpretability and system efficiency. The sparse access pattern is advantageous for hardware-constrained environments, allowing for hierarchical memory management where only frequently accessed hot slots are dynamically cached in VRAM, significantly reducing GPU memory overhead.

Improvements for AI systems

As a fastidious researcher, I have analyzed the RAM-Net architecture and its proposed mechanisms for improving AI systems. The core innovation lies in decoupling memory capacity from model parameters via a differentiable address decoder, enabling massive state scaling with computational efficiency.

Here are the specific improvements that can be made to AI systems using RAM-Net, and what these improved systems can achieve:


The following improvements stem directly from the architectural innovations detailed in RAM-Net (Sections 3.1, 3.2, 3.3, and 5):

  1. [] Utilize a fixed-size memory state that scales exponentially with sequence length without increasing learnable parameters (Section 3).

  2. [] Implement fine-grained long-range retrieval tasks with significantly reduced computational overhead compared to full attention (Section 1, Abstract).

  3. [] Achieve superior performance in standard language modeling and zero-shot commonsense reasoning benchmarks due to enhanced capacity for capturing complex dependencies (Section 1, Table 1).

  4. [] Enable efficient hardware deployment by leveraging a sparse access pattern mimicking Random Access Memory (RAM) for high computational efficiency, confining state updates and retrieval to minimal active entries (Section 3).

  5. [] Incorporate temporal awareness into memory access patterns via Cyclic Address Positional Embedding (CAPE), allowing the model to capture relative positional information analogous to temporal convolutions (Section 3.2).

  6. [] Implement dynamic, content-dependent memory retention using Power Decay Moving Average (PDMA), decoupling forgetting rates from writing intensity to allow for fine-grained control over historical context preservation (Section 3.3).

These improvements enable the following capabilities in AI systems:

  1. [] Processing of extremely long sequences (e.g., multi-document reasoning, massive codebases) with high fidelity recall, overcoming the linear memory growth bottleneck of full attention.

  2. [] High-precision needle-in-a-haystack retrieval tasks in complex contexts (e.g., identifying a specific entity mentioned hundreds of tokens earlier in a long document).

  3. [] Enhanced zero-shot commonsense reasoning, as the increased memory capacity allows the model to store and retrieve more diverse semantic features required for inferring novel relationships across long contexts (as validated by MQAR benchmarks).

  4. [] Deployment on resource-constrained hardware, such as edge devices or specialized accelerators, due to the inherent sparsity of state updates and retrieval operations.

  5. [] Improved temporal modeling capabilities in sequential data (like time-series or dialogue) by explicitly encoding relative time/positional shifts into the memory addressing scheme.

Abstract

Linear attention offers an efficient alternative to full attention with a fixed-size recurrent state. However, this state is shared by all tokens, so information from distinct tokens becomes superposed within it and produces inter-token interference that degrades long-range fine-grained recall. To address this issue, we propose RAM-Net, which replaces dense access to a shared state with sparse address-based access. RAM-Net organizes the recurrent state as a fixed-size array of independent slots and uses an Address Decoder that maps each key or query into a sparse address, selecting a small subset of slots to write to or read from at each step. This design directs tokens with non-overlapping addresses to disjoint slots, suppressing inter-token interference, while keeping per-step state access dependent only on the number of selected slots rather than the total state size. Empirically, RAM-Net outperforms strong recurrent baselines on fine-grained long-range retrieval and achieves the lowest perplexity with competitive commonsense reasoning. It does so while accessing fewer state elements per step than all baselines, e.g., 8 times fewer than Mamba2.

Sources

Related papers