Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior
cs.LG, cs.CL
Submitted: 2026-05-26
Updated: 2026-09-01
Code: https://github.com/karpathy/nanochat
License: http://creativecommons.org/licenses/by/4.0/
The gist: We study Latent Recurrent Transformer (LRT), a lightweight augmentation of autoregressive transformers that reuses a high-level source-layer hidden state from the previous token as recurrent memory
Terminology
Abstract
We study Latent Recurrent Transformer (LRT), a lightweight augmentation of autoregressive transformers that reuses a high-level source-layer hidden state from the previous token as recurrent memory for the next token. Because this state is already computed during ordinary decoding, LRT introduces a cross-token, cross-layer latent pathway while preserving the standard attention mechanism, KV-cache interface, and one model forward per generated token. To pretrain this recurrence without sequentially unrolling the full sequence, we introduce interleaved parallel training: one full-sequence initialization forward constructs a shared buffer, followed by sequential refinement of disjoint position subsets with parallel computation within each subset. This provides every token with recurrent-memory-aware supervision at approximately 2x ideal token compute. Across 1.3B- and 2.1B-parameter nanochat-style backbones and a wide range of training budgets, LRT improves both BPB and CORE under matched effective compute. Additionally, LRT outperforms two-forward PonderLM-2 and matches a three-loop Transformer in BPB, while retaining one-forward-per-token decoding with 9% latency overhead over the standard Transformer.
Sources
- Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models
- Addressing Some Limitations of Transformers with Feedback Memory
- Hungry Hungry Hippos: Towards Language Modeling with State Space Models
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- Training Large Language Models to Reason in a Continuous Latent Space
- Thinking Tokens for Language Modeling
- TransformerFAM: Feedback attention is working memory
- Jamba: A Hybrid Transformer-Mamba Language Model
- Landmark Attention: Random-Access Infinite Context Length for Transformers
- Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention
- RWKV: Reinventing RNNs for the Transformer Era
- Let's Think Dot by Dot: Hidden Computation in Transformer Language Models
- Compressive Transformers for Long-Range Sequence Modelling
- Simplified State Space Layers for Sequence Modeling
- Learning to (Learn at Test Time): RNNs with Expressive Hidden States
- Retentive Network: A Successor to Transformer for Large Language Models
- Gemma 2: Improving Open Language Models at a Practical Size
- Memorizing Transformers
- Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks