Memory-Efficient Looped Transformer: Decoupling Compute from Memory in Looped Language Models

arXiv:2605.07721 · cs.CL, cs.AI, cs.LG · Submitted 2026-05-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Memory-Efficient Looped Transformer".

Jane: Recurrent LLM architectures have emerged as a promising approach for improving reasoning, as they enable multi-step computation in the embedding space without generating intermediate tokens.

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So, wrapping up our discussion on "Memory-Efficient Looped Transformer: Decoupling Compute from Memory in Looped Language Models," we’ve covered how MELT proposes a novel architecture that keeps a single KV cache per layer shared across loops, updated by a learnable gating mechanism. Jane, what are your final thoughts on the title and authors?

Jane: I think the title perfectly captures the main idea: decoupling compute from memory consumption. The authors have done a lot of work showing how to retain that multi-step computation capability while solving the problem of memory growth in looped transformers like LoopLM one.

Lu: The implications are significant because it suggests that we can build recurrent reasoning systems that are much deeper than what was practically feasible before due to memory constraints. It really pushes the boundary on what's possible with these types of architectures.

Meng: From my side, the practical implication is that this architecture could make complex AI agents much more capable in tasks requiring sustained internal deliberation without requiring prohibitively large amounts of dedicated memory during those reasoning steps.

Lalam: For culture, I see this as enabling a new class of AI systems where complex planning and nuanced decision-making can be built into the core structure rather than relying solely on massive external context windows or deep prompt engineering.

Tom: It’s about building capability into the architecture itself, which is a fundamental shift in how we approach reasoning in LLMs. I really think this work lays out a path forward for more efficient and powerful AI systems overall.

Jane: Exactly, and it shows that even when dealing with complex recurrent structures, we can find ways to manage the memory footprint effectively through clever design choices like the gated momentum mechanism they describe.

Lu: The stability proofs mentioned in the paper also add weight to this; they suggest that this decoupling isn't just an empirical observation but something structurally sound for optimization across many loops.

Meng: It’s interesting how they combine architectural novelty with training stabilization techniques, which is often a challenge in AI research, and it seems they’ve navigated that space well here.

Lalam: If we can make these complex reasoning systems more efficient, it means the next generation of AI could be deployed where intelligence isn't just about size, but about intelligent state management.

Conclusion: Tom: So we've seen how MELT tackles the memory challenge in looped transformer architectures by sharing KV caches across loops, and now we're looking at what all this means for the title and who wrote this paper.

Jane: I think the title really does a good job of explaining that while keeping things simple, it hints at a clever trick—decoupling the heavy lifting from how much memory you use during reasoning.

Lu: From my perspective, it’s fascinating because they’ve essentially found a way to make deep reasoning possible without the memory requirements ballooning linearly with every step in the chain.

Meng: I'm wondering how practical this is for real-world applications; can we actually deploy a system that runs this efficiently on existing hardware without needing specialized setups?

Lalam: For me, the implication is huge because it suggests we can build AI systems that sustain long, complex internal thought processes much more reliably and affordably than before.

Tom: That's what I'm hearing—it’s about making those deep reasoning capabilities accessible to more people without needing a massive infrastructure just to keep the memory moving.

Jane: Exactly, it means we can focus on the quality of the reasoning steps rather than constantly worrying about running out of room for the intermediate calculations.

Lu: It opens up whole new avenues for creative AI structures where memory management isn't a bottleneck, allowing us to explore much more intricate logical pathways.

Meng: But I still need to see how robust this constant-memory design is when we start scaling these models up to truly massive parameter counts.

Lalam: And that's where the real impact lies; if we can manage that state efficiently, the next generation of AI could be deployed in areas requiring sustained, deep strategic planning.

Victor Conchello Vendrell, Arnau Padrés Masdemont, Niccolò Grillo, Jordi Ros-Giralt Arash Behboodi, Fabio Valerio Massoli

Qualcomm AI Research

cs.CL, cs.AI, cs.LG

Submitted: 2026-05-08

Updated: 2026-09-28

Code: https://github.com/huggingface/trl

Importance score: 91/100

The gist: Recurrent LLM architectures have emerged as a promising approach for improving reasoning, as they enable multi-step computation in the embedding space without generating intermediate tokens.

Key concepts

KV Cache Sharing
Instead of storing separate Key and Value states for every loop iteration, MELT maintains one shared KV cache per layer. This prevents memory usage from growing linearly with the number of reasoning steps, making it constant regardless of how deep the computation goes.
Latent State Evolution
MELT uses a separate latent state 'h' that evolves across iterations. Keys and Values are derived from this evolving state using learned projections, rather than being directly updated at every step. This preserves semantic integrity while decoupling memory updates from the attention retrieval process.
Learnable Gating Mechanism
A learnable gating mechanism controls how the latent state is updated across loops. It uses a formula to mix information from the current input and previous states, allowing the model to selectively retain or discard relevant information, optimizing memory usage during reasoning.

Terminology

Summary

Recurrent LLM architectures have emerged as a promising approach for improving reasoning, as they enable multi-step computation in the embedding space without generating intermediate tokens.

The gist: MELT decouples reasoning depth from memory consumption by maintaining a single KV cache per layer that is shared across reasoning loops and updated via a learnable gating mechanism, achieving constant-memory iterative reasoning without sacrificing LoopLM performance.

Architecture and Memory Efficiency

MELT is a novel architecture designed to decouple reasoning depth from memory consumption by maintaining a single KV cache per layer that is shared across reasoning loops. This contrasts with standard looped transformers where memory grows linearly with the number of loops due to Key-Value (KV) states, which in LoopLM scales as ∝ O(L × T). Instead of appending a new state at every loop step, MELT maintains a single KV entry per token and layer and updates it across loops via a learnable gating mechanism. This design leads to a memory complexity of M MELT ∝ O(N × L), effectively recovering the footprint of non-looped transformers.

Latent State Evolution

A key design choice in MELT is to maintain a separate latent state h that evolves across iterations, from which keys and values are derived through learned projections (WK, WV), rather than directly updating the KV cache at each loop step. This approach preserves semantic integrity and decouples memory updates from attention retrieval. The latent state is updated via a learnable gated momentum mechanism defined by Equation 1:

z(l)t = σ(x(l)t W(l)z + h(l)t−1U(l)z + b(l))

h(l)t = z(lt)(x(lt))⊙ h(lt−1) + (1 − z(lt))(x(lt))

Training Procedure

To enable stable and efficient training, MELT is trained using a chunk-wise training in a two phase procedure. The first phase involves an interpolated transition from the LoopLM starting model to MELT. This is achieved by computing two KV pairs in parallel—KVbase from the hidden states as in a standard LoopLM and KVMELT from the MELT architecture—and linearly combining them: KV = α KVMELT + (1 − α) KVbase, where α increases linearly from 0 to 1 during training. The second phase involves attention-aligned distillation, using the frozen LoopLM as a layer-wise teacher and applying an attention-alignment loss to stabilize training and consolidate learned representations.

Performance and Ablation

Empirically, MELT models fine-tuned from pretrained Ouro parameters outperform standard LLMs of comparable size while maintaining a memory footprint comparable to those models. Ablation studies confirm the necessity of both training phases: removing attention-aligned distillation causes a notable performance drop, and removing interpolated transition degrades performance. The paper also investigates the gating mechanism, showing that the element-wise gating mechanism consistently achieves the best performance compared to variants like Mean or Last. Furthermore, analysis of untrained KV-cache sharing strategies on Ouro demonstrates that they obtain zero performance on several reasoning benchmarks, highlighting the necessity of MELT's constant-memory design for long reasoning traces.

Theoretical Stability

The architecture is theoretically supported by stability proofs. Proposition E.1 shows that in the saturated-gate regime (where zt → 1), the Jacobian Jt converges to the identity matrix (Jt ≈ I), ensuring that gradients are preserved across loops, establishing a Gradient Superhighway where error signals traverse arbitrary depths, effectively alleviating the vanishing gradient problem and allowing optimization of deeper looped transformer models. This structural stability is further supported by Remark E.2, which notes that the continuous gate achieves Effective Spectral Regulation.

Conclusion

MELT establishes a retrieval hierarchy that decouples positional addressing from feature extraction, using the Value projection WV to act as a spectral filter to isolate the specific feature subspace required for computation. It is presented as the first architecture to exceed the performance of standard models with the same memory footprint, demonstrating that improved reasoning capability can be achieved without increasing memory. Despite challenges like sequential KV updates during training, chunk-wise training provides a controllable trade-off between inference fidelity and training efficiency. Future work includes exploring GQA integration to further reduce KV cache overhead.

References

[1] Rui-Jie Zhu et al., Scaling latent reasoning via looped language models, 2025.

[2] Mostafa Dehghani et al., Universal transformers. ICLR, 2019.

[3] Jason Wei et al., Chain-of-thought prompting elicits reasoning in large language models. NeurIPS, 2022.

[8] Nikunj Saunshi et al., Reasoning with latent thoughts: On the power of looped transformers.

Improvements for AI systems

As a fastidious researcher, I see several high-impact avenues for improving AI systems based on the Memory-Efficient Looped Transformer (MELT) architecture described in this paper.

Here are the specific improvements and capabilities of an AI system utilizing MELT:


) Improvements to AI Systems via MELT Architecture:

  1. The primary improvement is the ability to perform deep, multi-step reasoning with a fixed, manageable memory footprint, directly addressing the linear memory growth limitation of standard looped transformers (like Ouro).

  2. The architecture decouples reasoning depth from memory consumption by replacing the append-only Key-Value (KV) cache with a single KV entry per token and layer, updated via a learnable gating mechanism. This reduces complexity from linear in reasoning depth to constant per layer, resulting in a memory footprint comparable to non-looped models (e.g., Qwen3-1.7B), but with superior reasoning performance.

  3. The system can be fine-tuned from pre-trained Looped Language Models (LoopLM) using a novel, data-efficient two-phase training procedure:

4 a) Interpolated Transition: A linear transition between the LoopLM dynamics and MELT dynamics to ensure smooth architectural adaptation.

5 b) Attention-Aligned Distillation: Using a frozen LoopLM as a teacher with an attention-alignment loss to stabilize representations during the transition phase.

  1. The internal gating mechanism is designed to allow each token, at every time step, to attend to keys and values that integrate information across all preceding time steps of preceding tokens (not just the current step), enabling richer contextual integration without expanding memory.

  2. The training utilizes chunk-wise training, processing sequences in fixed-length chunks in parallel within each chunk and sequentially across chunks, balancing inference fidelity (smaller chunks) with computational efficiency (larger chunks).

  3. The system benefits from a Gradient Superhighway effect (Proposition E.2), where the gated update mechanism ensures that error signals traverse arbitrary depths during backpropagation, effectively alleviating the vanishing gradient problem in deep recurrent models and enabling optimization of deeper loops.

) Capabilities of the Improved AI System:

An AI system built on MELT would excel in complex, iterative cognitive tasks requiring sustained thought over long reasoning chains without incurring prohibitive computational or memory costs. Specifically:

  1. Deep, Multi-Step Mathematical Reasoning: The system can tackle complex problems like MATH-500 and AIME competitions (e.g., AIME26), maintaining high accuracy (up to 94.4% on MATH-500) while utilizing a memory footprint comparable to much smaller, non-looped models.

  2. Advanced Code Reasoning and Problem Solving: It can perform complex coding tasks and algorithmic problem-solving (e.g., AMC23), achieving top performance on benchmarks like HumanEval (Humaneval accuracy 81.7%).

  3. Long-Horizon Contextual Reasoning: Unlike standard models where reasoning depth is limited by memory, MELT enables the model to maintain coherent reasoning over significantly longer sequences or iterative steps without memory exhaustion, making it suitable for complex scientific hypothesis generation or multi-hop logical deduction tasks that require sustained internal state tracking.

  4. Efficient Adaptation and Knowledge Transfer: The two-phase training procedure allows developers to efficiently adapt massive, pre-trained looped models (like Ouro) to new architectures (MELT) with minimal fine-tuning data, ensuring the resulting model retains the strong reasoning capabilities of the original architecture while benefiting from its memory efficiency.

  5. Stable Performance Under Memory Constraints: The system demonstrates a superior performance–efficiency trade-off, achieving results that outperform similarly sized standard Transformer baselines while operating under a strictly constrained memory budget (constant memory per layer).

Sources

Related papers