MARCH: Scaling Recurrent Memory with Content-Routed State Anchors
Ming Zhang, Kaisen Yang, Shu Yu, Ermo Hua, Ning Ding, Xia Hu, Bowen Zhou, Chaochao Lu, Youbang Sun
Shanghai AI Laboratory · Tsinghua University · Fudan University
cs.LG
Submitted: 2026-08-12
Updated: 2026-08-14
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 82/100
The gist: MARCH: Scaling Recurrent Memory with Content-Routed State Anchors Abstract Transformers owe much of their strong long-context retrieval capability to a token-level memory that grows with context
Terminology
Summary
MARCH: Scaling Recurrent Memory with Content-Routed State Anchors
Abstract
Transformers owe much of their strong long-context retrieval capability to a token-level memory that grows with context length. This flexibility, however, incurs a quadratic computation complexity during training and a key–value cache that grows linearly during autoregressive inference. Recurrent alternatives offer efficient decoding by compressing the entire history into a fixed-size state, but often underperform on recall-intensive tasks since earlier associations usually get overwritten by subsequent updates, and only the most recent contextual information is retained. In this paper, we introduce Memory-Anchor Routing across Context History (MARCH), a network architecture that effectively scales state-space models beyond a fixed-size dimension, while maintaining computational efficiency over long-sequences. MARCH periodically caches cumulative recurrent-state checkpoints as state anchors and associates each anchor with a compact, content-conditioned anchor key. This lets MARCH maintain a memory bank, which can grow as context length increases, providing a controllable trade-off between historical resolution and memory cost. At each token, MARCH produces an anchor query to attend all causally available state anchors, and the output is calculated as an attention-style aggregation over all historical anchors along the current state. We show that after standard pretraining, MARCH consistently outperforms multiple linear attention variants across commonsense reasoning, LongBench, and in-context retrieval. These results demonstrate that content-routed state caching substantially strengthens recurrent long-range memory while preserving its native computation path.
1. Introduction
Large language models (LLMs) have demonstrated remarkable capabilities across a wide range of language understanding and generation tasks. However, many real-world applications—including long-document understanding, multi-turn interaction, and in-context learning—require models to integrate information distributed across extended sequences. Supporting such contexts involves more than simply increasing the number of input tokens: models must retain relevant information across long intervening spans and reliably retrieve it when needed. Effective long-context modeling therefore hinges on a model’s ability to manage memory—determining what information to preserve, how to represent it, and when to retrieve it. A useful perspective is to view a sequence model as a memory system with two basic operations: writing, which incorporates each new input into memory, and reading, which retrieves information relevant to the current input. Under this view, standard self-attention maintains a growing token-level memory in its key–value cache: it writes by appending each new key–value pair without compressing the existing cache, and reads by matching the current query against all stored keys and combining their associated values. This uncompressed, token-level memory provides a direct path to every preceding token, enabling accurate recall of fine-grained details and distant dependencies. Its flexibility, however, entails quadratic computation during training and a key–value cache that grows linearly with sequence length during autoregressive inference, making self-attention increasingly costly as context windows expand.
Linear attention and modern recurrent sequence models make the opposite trade-off. They compress the causal prefix into a fixed-size, matrix-valued recurrent state, update this state with each new input, and read from it by applying the current query. This design enables constant-memory recurrent decoding, but the same compressed write that makes it efficient also limits its long-range memory. At each step, information from the new token is written into a state already shared by the entire history. Much recent work has therefore focused on improving the write operation. Selective state-space models such as Mamba introduce input-dependent state transitions and forgetting, whereas DeltaNet and Gated DeltaNet use data-dependent delta-rule updates to revise existing associations before incorporating new information. These mechanisms improve state tracking and mitigate indiscriminate accumulation, but they do not eliminate the underlying fixed-state bottleneck: the entire history must still share one evolving state, and only its latest version remains available for reading. Indeed, although recent linear recurrent models can match or surpass softmax attention in short-context settings, their performance often degrades as the evaluation context grows. Once an earlier association has been weakened by forgetting or modified by subsequent writes, the model has no direct path to its earlier representation and cannot recover it from the latest state alone.
Recent works have relaxed the fixed-state bottleneck along two broad directions. One approach increases the memory capacity available at each step through partitioned sparse states, large memories with sparse reads and writes, or routed mixtures of independent states. A second line preserves a temporally structured collection of compressed states through logarithmic hierarchies, adaptive state construction and merging, or recurrent-state caching. Collectively, these approaches demonstrate that expanding memory capacity or temporal coverage can improve long-range recall. However, with the exception of certain instances in [Behrouz et al., 2026], all of the existing works still maintain a finite or upper-constrained state space dimension, though the dimensionality of which is increased. Moreover, as multiple states become available, the primary bottleneck shifts from memory construction to memory retrieval: the model must determine which state retains the information most relevant to the current token. In particular, how to construct context-dependent representations for historical checkpoints that facilitate effective and efficient query-dependent retrieval remains underexplored.
In this work, we introduce Memory-Anchor Routing across Context History (MARCH), a memory-augmented recurrent architecture that enables selective retrieval from earlier versions of recurrent memory. Without modifying the underlying recurrence, MARCH periodically preserves cumulative states as state anchors, giving later tokens access to earlier versions of the evolving memory. Each anchor is associated with a compact learned descriptor, allowing the model to route each token to relevant historical states when additional context is needed and combine their contents with the current-state readout. MARCH thereby complements efficient recurrent processing with selective access to preserved historical memory. The resulting mechanism remains causal and is trained end to end with the standard language-modeling objective.
Main contributions:
-
Routable memory-state bank. We introduce MARCH, which expands fixed-state recurrence into a growing bank of historical memory states which increases its capacity as context grows, alleviating the single-state memory bottleneck without modifying the underlying recurrence.
-
Content-conditioned historical retrieval. MARCH brings attention-style content routing to recurrent memory by applying a standard softmax over compact keys for a temporally sparse set of state anchors rather than token-level key–value pairs. A learned null route allows the model to suppress the historical branch when the current recurrent state is sufficient, while residual fusion preserves the original recurrent path and supports end-to-end training.
-
Extensive empirical validation. We demonstrate consistent improvements over strong recurrent baselines across commonsense reasoning, LongBench, in-context retrieval, and NIAH evaluations, including robust extrapolation beyond the training context length.
2. Preliminaries
Full and Linear Attention as Memory. Let xt ∈ Rd denote the hidden representation at position t. The corresponding query, key, and value vectors are obtained through learned linear projections: qt = Wqxt, kt = Wkxt, vt = Wvxt, where Wq, Wk ∈ Rdk×d and Wv ∈ Rdv×d, such that qt, kt ∈ Rdk and vt ∈ Rdv. Following the memory-system perspective adopted in prior work, we view a causal sequence mixer as an online memory system. Let Mt denote the memory state after processing the first t tokens, with M0 denoting its initial state. At each position, the current key–value pair is first written into memory, after which the updated memory is queried using the current query: Mt = Write(Mt−1; kt, vt), ot = Read(Mt; qt), where ot ∈ Rdv denotes the memory readout at position t. For causal softmax attention, the memory explicitly retains all projected key–value pairs observed up to position t: Mt = (K≤t, V≤t), where K≤t ∈ Rt×dk and V≤t ∈ Rt×dv stack the keys and values row-wise, respectively. Writing appends (kt, vt) to these matrices, whereas reading performs content-based retrieval: ot = V⊤≤t softmax(K≤tqt/√dk). This explicit storage keeps individual tokens directly retrievable and enables fine-grained retrieval from the entire causal prefix. However, processing a sequence of length T requires O(T2) query–key interactions, while autoregressive decoding maintains a key–value cache of size O(T(dk+dv)) per attention head.
Linear attention instantiates the memory state Mt as a fixed-size matrix St ∈ Rdv×dk. Its write and read operations are given by St = St−1 + vtkt⊤, ot = Stqt. Each write therefore adds a rank-one key–value association to the shared matrix, while each read retrieves a query-dependent superposition of the stored values. This enables constant-memory recurrent decoding, but introduces interference as the compressed history grows.
Gated DeltaNet (GDN). To mitigate the interference caused by the additive write rule of linear attention, GDN retains the same state-based read operation but introduces input-dependent retention and a targeted delta-rule write: St = alphatSt−1 + betat(vt − alphatSt−1kt)kt⊤, ot = Stqt, where alphat ∈ (0,1) is an input-dependent retention gate, and betat ∈ [0,1] modulates the strength of the targeted delta update. Despite its more adaptive state dynamics, GDN still compresses the entire causal history into a single fixed-size recurrent state. Because all associations share this evolving state, information weakened or overwritten by subsequent updates has no direct retrieval path, limiting reliable long-context recall.
Scaling recurrent memory. Recent work has sought to relax the fixed-state bottleneck along two broad directions. Capacity-expansion methods enlarge the current recurrent state, whereas temporal-expansion methods retain multiple versions of the state along its trajectory: M cap t:= S̃t = [St(1) · · · St(P)] ∈ Rdv×Dmem, Dmem = Σp=1..P dp; M temp t:= St,1,..., St,Mt, St,m ∈ Rdv×dk, 1 ≤ taut,1 < · · · < taut,Mt ≤ t. In the capacity formulation, P is the fixed number of state partitions and Dmem is their total memory dimension. Such methods increase Dmem while using sparse access to keep computation tractable. In the temporal formulation, Mt is the number of retained state representations, each associated with a temporal boundary taut,m. These methods preserve states from distinct temporal regions or earlier stages of the recurrent trajectory. MARCH follows the latter direction by retaining cumulative snapshots of a continuously evolving recurrent state.
3. Method
Figure 2 illustrates MARCH, a content-routed recurrent memory framework that enables selective retrieval from earlier versions of an evolving recurrent state. Rather than routing over predefined state indices or temporal scales, MARCH matches each query against individual historical states based on their contents. As tokens are processed, MARCH periodically checkpoints the cumulative recurrent state, producing a bank of state anchors. Each checkpoint is paired with an occurrence of a shared learned anchor token, whose hidden representation yields a compact routing key. For each text token, a routing query scores all causally visible anchors alongside a learned null option, allowing the model to use historical memory only when useful. The resulting routing probabilities define a weighted combination of the visible anchor states, which is read using the token’s standard recurrent query. This historical readout is added to the current-state readout, preserving the native recurrent path while introducing a content-dependent route to earlier memory. Together, state anchoring and content-routed retrieval turn the otherwise transient state trajectory into a persistent source of long-range memory.
3.1. Continuous Recurrent-State Anchoring
Anchor placement. Let T = [t1,..., tL] be a sequence of L text tokens. An anchoring policy specifies an ordered set of text boundaries B = bm m=1..M, where 0 = b0 < b1 < · · · < bM ≤ L. We insert an anchor position after each boundary: T̃ = ∥m=1..M [tbm−1+1,..., tbm] ∥ [xim] ∥ [tbM+1,..., tL], where ∥ denotes sequence concatenation, and xim is the m-th occurrence of a shared learned anchor embedding xi. Text and anchor positions serve different computational roles. Text positions apply the base recurrent update, allowing the matrix-valued state to evolve continuously across anchor boundaries. Immediately after processing tbm, MARCH checkpoints the resulting cumulative state to form the m-th state anchor. The following anchor position xim does not modify the recurrent state; instead, its hidden representation provides the routing metadata associated with that checkpoint. Thus, each anchor boundary produces two coupled objects: a snapshot of the recurrent memory and a compact representation through which that snapshot can later be retrieved.
Cumulative recurrent-state checkpointing. At layer l, MARCH leaves the underlying recurrent update unchanged and carries the state St(l) ∈ Rdv×dk continuously across anchor boundaries. At each boundary bm, it snapshots the current state: A(m,l) = Sbm(l) ∈ Rdv×dk, m = 1,..., M. Because the recurrence is not reset between anchor boundaries, A(m,l) encodes the cumulative prefix up to position bm, rather than only the segment since the preceding boundary. We therefore refer to it as a state anchor. The ordered bank A(1,l),..., A(M,l) traces the temporal evolution of a single recurrent memory, preserving earlier versions before subsequent decay and delta updates attenuate or modify their contents.
Content-conditioned anchor metadata. Let um(l) denote the normalized input representation of anchor position xim at layer l. The anchor position reads only its aligned state checkpoint: qm(l) = Wq(l)um(l), om(l) = A(m,l)qm(l). The same input representation is projected into a compact routing key: kappam(l) = Wk(l)um(l) ∈ Rdr. The aligned readout is incorporated into the anchor position through the standard output projection and residual pathway. Consequently, um(l+1) depends on A(m,l), and the routing key produced at layer l+1 becomes conditioned on the content retained by the aligned state anchor. Thus, although all anchor positions share the same learned input embedding, they acquire distinct, state-dependent representations after the first layer. This cross-layer construction makes routing explicitly dependent on what each state anchor contains, rather than only on its temporal index.
3.2. Content-Routed Historical Reading
Content-based routing. For a text token at position t, the causally available state anchors are indexed by Vt = m ∈ 1,..., M bm < t. MARCH projects the normalized hidden state xt into a routing query and scores it against the key of each visible anchor: rhot = WRxt, at,m = rhot⊤kappam, m ∈ Vt. To allow the model to bypass historical memory, we augment the visible anchor set with a null option ∅, whose payload is fixed to zero, A(∅) = 0. Its query-dependent logit is nt = w∅⊤xt + b∅. Let Ṽt = Vt ∪ ∅ denote the augmented candidate set. We define the logit of each candidate j ∈ Ṽt as st,j = at,j if j ∈ Vt, and st,j = nt if j = ∅. The routing probabilities are then pit,j = exp(st,j) / Σr∈Ṽt exp(st,r). Since the selected routing probabilities directly weight the historical state readouts, their scores remain jointly optimized by the language-modeling objective. The routing query rhot determines which anchors to retrieve, whereas the state-read query qt reads their matrix-valued contents. If no anchor is visible, the null option receives all probability mass.
The aggregation formulation readily admits a sparse variant by restricting aggregation to the K highest-scoring visible anchors (Top-K). This sparse approach exhibits natural connections with the hierarchical sparse attention approaches, while preserving dense token-level processing rather than relying on hard token-level pruning. Ablation studies show that sparse routing substantially reduces aggregation cost with minimal performance degradation.
Historical retrieval and residual fusion. Given the routing probabilities, the causally visible state anchors are aggregated into a query-dependent historical state, which is read using the same state-read query as the current state. The resulting historical readout is then added to the current-state readout: ot = Stqt + Σj∈Ṽt pit,j A(j)qt. This additive formulation preserves the original recurrent path and introduces historical retrieval as an auxiliary residual branch, without modifying the underlying recurrent update. Since the routing probabilities directly affect the layer output, the routing queries and anchor-derived keys are optimized end-to-end with the language-modeling objective.
3.3. Implementation
MARCH is implemented as a two-stage producer–reader computation. Following the hardware-efficient chunkwise formulation of Gated DeltaNet, the producer processes recurrent updates in blocks amenable to tensor-core acceleration, computes each token’s current-state output, and checkpoints the recurrent state at each anchor boundary. The resulting state anchors are consumed by the historical reader. Inspired by the I/O-aware principles of FlashAttention, the reader jointly tiles query tokens and state anchors, reuses each anchor tile across a block of queries, and fuses routing-score computation, online softmax updates, and the accumulation of weighted state readouts into a streaming reduction. This fused schedule avoids materializing either the dense token-to-anchor routing matrix or the substantially larger tensor of per-anchor candidate readouts, thereby reducing intermediate storage and the associated HBM traffic. As shown in Figure 4, despite the cost of historical retrieval, the fused dense implementation exceeds FlashAttention-2 in throughput at 64K and above and incurs lower core runtime from 32K onward.
4. Experiments
MARCH is designed to extend the long-range memory of recurrent models while preserving their general language capabilities. The paper verifies the effectiveness of MARCH by training from scratch, and evaluates across a diverse suite of benchmarks spanning zero-shot commonsense reasoning, long-context understanding, and in-context retrieval. Across these tasks, MARCH consistently outperforms existing recurrent baselines, with particularly strong gains on retrieval-intensive and long-context benchmarks.
4.1. Experimental Setup
Training configuration. Following the academic-scale protocol used by Log-Linear Attention, the models are pretrained from scratch on 50B tokens from the Long-Data-Collections dataset, using a sequence length of 16K. The main configurations use 21 layers and a hidden size of 1536. The Transformer (693M) uses 16 attention heads and a RoPE base of 500K, while Gated DeltaNet (793M) and its variants use six value heads. To control for parameter count in addition to model depth, a 24-layer Transformer with (778M) parameters is also included, closely matching the size of the Gated DeltaNet. For MARCH, the routing dimension is set to dr = 64 and a periodic anchoring interval of C = 512 text tokens is used. All models are trained with a global batch size of approximately 4.2M tokens using the fused AdamW optimizer, with beta1 = 0.9, beta2 = 0.95, epsilon = 10−8, and a weight decay of 0.1. The peak learning rate is set to 4 × 10−4 with a warmup-stable-decay schedule. All models use the same training data, token budget, context length, and optimization configuration.
Baselines. Primary comparisons are against standard GDN and GDN augmented with Log-Linear Attention. MARCH and these two baselines use matched architectural configurations and the same pretraining setup, enabling a controlled comparison of their memory mechanisms. To contextualize their performance against full attention, two Transformer baselines are additionally included: a 21-layer model matched in depth to the recurrent models and a 24-layer model approximately matched to them in parameter count.
Evaluation tasks. For short-context generalization, eight zero-shot commonsense benchmarks are used: LAMBADA, PIQA, HellaSwag, WinoGrande, ARC-Easy and ARC-Challenge, OpenBookQA, and CommonsenseQA. Long-context understanding is evaluated on LongBench, covering single-document QA, multi-document QA, summarization, and few-shot learning. The long-context retrieval evaluation covers six single-needle and multi-needle tasks from RULER at 4K, 8K, and 16K context lengths. Finally, the in-context retrieval suite contains SQuAD, TriviaQA, SWDE, FDA, Natural Questions, and DROP. The evaluation protocol of prior work is followed, using the LM-Evaluation-Harness.
4.2. Main Results
Commonsense reasoning. As shown in Table 1, MARCH consistently outperforms both the vanilla and Log-Linear variants of Gated DeltaNet across all eight zero-shot commonsense reasoning benchmarks. It improves the average accuracy from 40.1 and 40.0 to 41.5, respectively, with the largest gain over the vanilla baseline observed on OpenBookQA (+2.8 points), aligning with findings in retrieval tasks. Moreover, MARCH achieves a higher average score than both Transformer baselines, surpassing the standard Transformer on six of eight tasks and the 24-layer Transformer on four. These results indicate that MARCH consistently strengthens the Gated DeltaNet backbone while remaining competitive with comparable full-attention models on short-context language understanding tasks.
Needle-in-a-haystack retrieval. The long-context associative retrieval is evaluated using the needle-in-a-haystack (NIAH) suite from RULER, where a model must recover values associated with keys embedded among irrelevant context. All models are trained with a maximum context length of 16K; Figure 3 reports results from 4K to 32K, making 32K a zero-shot length-extrapolation setting. Across the 24 task–length combinations, MARCH outperforms the stronger recurrent baseline in 19 settings and matches it in the remaining five. On the multi-needle tasks, it wins in 11 of 12 settings. At 32K, MARCH achieves the best result on all six tasks, retaining perfect accuracy on S-NIAH-1 and nonzero accuracy on the remaining tasks, whereas both Transformer variants and Log-Linear Gated-DeltaNet score zero throughout. This contrast is consistent with RoPE extrapolation in the Transformers and the state-index-dependent coefficients of Log-Linear Gated-DeltaNet. MARCH instead shares the same content-based router across all anchors, allowing longer contexts to introduce additional anchors without requiring new anchor-specific routing parameters.
Long-context understanding. Table 2 reports results across four LongBench task categories. MARCH consistently outperforms both vanilla Gated DeltaNet and its log-linear variant on all twelve tasks. The improvements are particularly pronounced on multi-document QA: relative to the stronger of the vanilla and log-linear Gated DeltaNet baselines, MARCH raises the 2WikiMultihopQA score from 8.7 to 11.5 and the MuSiQue score from 3.3 to 4.8, corresponding to relative gains of 32% and 45%, respectively. The benefits also extend to summarization, where the QMSum score increases from 13.2 to 17.4 (32%), and to all three few-shot learning tasks. These results show that content-routed state anchors improve long-context understanding across diverse task formats, rather than benefiting only retrieval-oriented question answering.
In-Context Retrieval. Following prior work, in-context retrieval is evaluated on six real-world, recall-intensive benchmarks. As shown in Table 3, MARCH consistently outperforms both vanilla Gated DeltaNet and its Log-Linear variant across all tasks. Relative to the stronger of the vanilla and log-linear Gated DeltaNet baselines on each benchmark, MARCH yields relative improvements ranging from 8% on SQuAD to 23% on TriviaQA and raises the average accuracy from 20.5 to 23.3, corresponding to a 14% relative improvement. These consistent gains across heterogeneous retrieval tasks demonstrate that MARCH improves the retrieval capability of the Gated DeltaNet backbone beyond a particular dataset or input format. Together, these results establish content-routed state anchors as an effective mechanism for strengthening fine-grained retrieval in recurrent models.
Training efficiency. Figure 4 compares the end-to-end throughput and core forward–backward runtime of FlashAttention-2, Gated DeltaNet, dense MARCH, and its Top-4 implementation. Sparse routing becomes increasingly beneficial as the context grows. At 128K tokens, Top-4 MARCH more than doubles the training throughput of dense MARCH and reduces its core runtime by roughly an order of magnitude. It also achieves higher throughput than FlashAttention-2 at this length, although vanilla Gated DeltaNet remains faster because it incurs no historical-retrieval overhead.
5. Ablation Studies
Effect of chunk size. The training chunk size C determines how frequently MARCH checkpoints the recurrent state. Smaller chunks create denser candidate anchors and offer finer temporal resolution, at the cost of a larger anchor cache and higher historical-routing overhead. C is varied from 256 to 2048 while keeping all other model and training settings fixed. Table 4 reports performance on three in-context retrieval benchmarks and the average over the six NIAH tasks at context lengths of 4K, 8K, and 16K. As an additional inference test, the MARCH state bank is organized according to the Fenwick tree scheme used by Log-Linear Attention. This scheme arranges state anchors hierarchically and retains O(log T) anchors as the context grows. Performance close to Log-Linear Attention is reported, demonstrating that the learned router in MARCH shows great generalizability and flexibility across various state bank organization schemes.
In Panel (a), C = 512 provides the best overall balance between retrieval quality and anchor count. Smaller chunks improve some long-context results but incur higher memory and routing costs, whereas larger chunks generally degrade retrieval because the resulting checkpoints are too sparse. C = 512 is therefore adopted as the default. Panel (b) shows that changing the chunk size at inference provides a flexible accuracy–memory trade-off: denser anchors generally improve retrieval at higher cost, while overly sparse anchors lead to substantial degradation. The Fenwick tree row additionally evaluates hierarchical organization of the state bank at inference.
Routing design. The router’s query–key dimension dr, routing sparsity, and learned null option are ablated. Table 5 reports aggregate results across general language understanding, long-context benchmarks, and NIAH. The default uses dense routing with dr = 64 and includes the null option. Increasing dr to 192 yields higher retrieval capability but reduces general performance across other tasks, making dr = 64 a more balanced choice overall. Top-4 nearly matches dense routing on commonsense and retrieval but trails on NIAH, making it an efficiency-oriented operating point when considered alongside Figure 4. Removing the null option degrades every aggregate, confirming the benefit of bypassing irrelevant historical states.
6. Related Work
Efficient Attention Mechanisms. Efficient attention reduces the quadratic cost of full self-attention through local windows, kernelization, or systems optimization. Local sliding-window attention limits each query to a bounded neighborhood. Performer, Nystromformer, and Linear Attention replace the softmax kernel with feature maps and exploit associativity for linear-time computation. FlashAttention-2, sequence parallelism, and chunkwise algorithms instead improve hardware efficiency without changing the dense attention pattern.
Sparse Attention. Sparse attention retains content-based softmax retrieval but restricts each query to a small subset of token-level key–value pairs. Early methods rely on predefined connectivity: Sparse Transformer factorizes the attention pattern, while Longformer and BigBird combine local windows with global or random links. Later methods make the sparse pattern input dependent: Routing Transformer clusters tokens by content, H2O evicts low-utility cache entries, and Quest selects KV-cache pages conditioned on the current query. More recent trainable designs route queries to relevant blocks, as in MoBA, or combine compressed, selectively retrieved, and local branches with hardware-aligned kernels, as in Native Sparse Attention. These approaches reduce attention-score computation or memory traffic, but their accuracy hinges on token or block selection and they still store or manipulate token-level KV memories. In contrast, MARCH routes over compact keys associated with historical recurrent-state snapshots, retrieving compressed prefix states rather than sparsifying token-to-token attention.
State Space Models and Gated Linear Recurrences. State space models (SSMs) and linear recurrent networks compress the prefix into a recurrent state. Linear Attention and its kernelized variants share this view through decayed outer-product updates and query-based reads. S4 uses structured linear dynamics, while Mamba and Mamba-2 use selective transitions; RetNet, RWKV, HGRN, and LRU combine associative memories with structured recurrences. GLA introduces input-dependent decay, DeltaNet and GDN use delta-rule corrections, and GDN-2 decouples erase and write through channel-wise gates. Despite these advances, most models retain one fixed-capacity state, whose dimension directly controls update and read cost; this bottleneck contributes to the retrieval gap with Transformers. MARCH preserves efficient recurrence while expanding memory into selectively accessed states.
State Expansion and Associative Memory. Long-context studies identify recurrent state capacity as a central limitation of Linear Attention models. Multi-State RNNs, HGRN2, and Log-Linear Attention expand or hierarchically organize recurrent states. Mixture-of-Memories, Sparse State Expansion, Product Key Memory, and Fast-weight Product Key Memory use memory experts or sparse banks, while Sparse Delta Memory sparsifies GDN reads and writes. Context-compression methods instead retrieve at the token or chunk level, using learned summary tokens or selective chunk reopening. MARCH treats recurrent states as retrieval units, expands total capacity, and reads only selected states, decoupling capacity from dense per-token updates.
7. Limitations and Future Work
MARCH adopts periodic checkpointing and organizes all historical states in a single homogeneous anchor bank. Although this design is simple and efficient, fixed-interval anchoring does not account for the non-uniform evolution of recurrent memory: it may create redundant anchors in stable regions while providing insufficient resolution when the state changes rapidly. A natural extension is to develop adaptive anchoring mechanisms according to state novelty or update magnitude and to consolidate or evict redundant anchors given a memory budget. More broadly, MARCH improves access to earlier states but does not explicitly increase or specialize the capacity of the underlying memory. Future work could combine state anchoring with larger-capacity memory and multiple memory partitions specialized for different temporal scales or information types. For example, short-term context, salient episodic events, and slowly consolidated knowledge could be maintained through distinct write, retention, and forgetting mechanisms, while a hierarchical router determines both which memory partition and which stored state should serve each query. In addition, MARCH has the potential to support external memory modules for optimized performance over specific downstream tasks, knowledge consolidation from experience to parametric information, and other memory manipulation mechanisms, opening up new scaling directions for test-time training and continual learning.
8. Conclusion
We introduce MARCH, a novel attention architecture which augments recurrent models with content-routed state anchors. By preserving cumulative state checkpoints, MARCH enables selective access to earlier recurrent states without modifying the underlying recurrence. It consistently outperforms strong recurrent baselines across commonsense reasoning, LongBench, in-context retrieval, and NIAH. Ablations further show a controllable retrieval–efficiency trade-off through checkpoint density and sparse routing. These results establish historical-state retrieval as a practical approach to scaling recurrent memory beyond a single evolving state.
Improvements for AI systems
Based on the paper, here are the specific improvements you can make to AI systems and what the improved system can do:
Improvement 1: Content-Routed Historical State Retrieval
-
What to implement: Add a memory bank of periodic recurrent-state checkpoints (state anchors), each paired with a learned, content-conditioned routing key. At each token, compute a routing query that scores all causally visible anchors plus a learned null option, then aggregate the top or all anchors’ states weighted by softmax probabilities, and add this historical readout to the current-state readout.
-
What the improved AI system can do: Retrieve earlier versions of its compressed memory that would otherwise be overwritten or forgotten. This enables accurate recall of distant facts, multi-hop reasoning across long contexts, and fine-grained retrieval from earlier parts of a sequence—without quadratic attention cost or a linearly growing KV cache. It can also bypass historical memory when the current state is sufficient, preserving efficiency.
Improvement 2: Scalable, Content-Adaptive Memory Capacity
-
What to implement: Replace the fixed-size recurrent state with a growing bank of state anchors that increases in number as context length grows. Use periodic checkpointing (e.g., every 512 tokens) and allow the bank to expand without bound, while keeping per-token computation sublinear via sparse routing (e.g., Top-4 anchors).
-
What the improved AI system can do: Maintain long-context performance beyond its training length (e.g., extrapolate from 16K to 32K tokens) without new parameters or retraining. It can handle arbitrarily long documents, multi-turn dialogues, and in-context learning tasks where earlier information must be retrieved after long intervening spans, while keeping memory and compute costs controllable via checkpoint density and routing sparsity.
Improvement 3: Hardware-Efficient Fusion of Routing and State Aggregation
-
What to implement: Use a two-stage producer–reader computation with chunkwise recurrent updates and I/O-aware tiling. Fuse routing-score computation, online softmax, and weighted state accumulation into a streaming reduction, avoiding materialization of dense routing matrices or per-anchor readout tensors.
-
What the improved AI system can do: Achieve higher training throughput than FlashAttention-2 at sequence lengths of 64K and above, and lower core runtime from 32K onward. This makes long-context training and inference practical on existing GPUs, enabling deployment in resource-constrained settings without sacrificing retrieval quality.
Improvement 4: Adaptive Anchoring and Hierarchical Memory Organization
-
What to implement: Extend the fixed-interval checkpointing to adaptive anchoring based on state novelty or update magnitude (e.g., checkpoint more frequently when the state changes rapidly, consolidate redundant anchors in stable regions). Optionally, organize the anchor bank hierarchically (e.g., Fenwick tree) to retain O(log T) anchors as context grows.
-
What the improved AI system can do: Optimize memory usage dynamically, reducing redundancy and improving temporal resolution where it matters most. This yields a flexible accuracy–memory trade-off at inference, allowing the system to adapt its memory footprint to the task’s demands (e.g., denser anchors for retrieval-heavy tasks, sparser for general reasoning) without retraining.
Improvement 5: Null-Route Suppression and Residual Fusion
-
What to implement: Include a learned null option in the router that can suppress the historical branch entirely, and add the historical readout as a residual to the current-state readout, preserving the original recurrent path.
-
What the improved AI system can do: Automatically decide when historical memory is unnecessary, avoiding noise or interference from irrelevant past states. This improves robustness on short-context tasks and prevents performance degradation, while maintaining end-to-end trainability with standard language-modeling objectives.
Improvement 6: Multi-Scale and Specialized Memory Partitions
-
What to implement: Combine state anchoring with multiple memory partitions (e.g., short-term, episodic, and consolidated knowledge), each with distinct write, retention, and forgetting mechanisms, and use a hierarchical router to select both the partition and the stored state for each query.
-
What the improved AI system can do: Handle heterogeneous information types (e.g., recent dialogue, salient events, long-term facts) with specialized memory, improving performance on tasks requiring both immediate context and distant knowledge. This enables more human-like memory consolidation and supports continual learning or test-time training by updating only relevant partitions.
Sources
- Titans: Learning to Memorize at Test Time
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- Linear Transformers Are Secretly Fast Weight Programmers
- Dynamic Linear Attention
- Scaling Linear Attention with Sparse State Expansion
- Sparse Delta Memory: Scaling the State of Linear RNNs through Sparsity
- MoM: Linear Sequence Modeling with Mixture-of-Memories
- Memory Caching: RNNs with Growing Memory
- MoBA: Mixture of Block Attention for Long-Context LLMs
- Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes
- RATTENTION: Towards the Minimal Sliding Window Size in Local-Global Attention Models
- Short window attention enables long-term memorization
- Linear Attention Sequence Parallelism
- Generating Long Sequences with Sparse Transformers
- Longformer: The Long-Document Transformer
- Retentive Network: A Successor to Transformer for Large Language Models
- Longhorn: State Space Models are Amortized Online Learners
- Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks