LSTM-UT and Recurrent-Depth Transformers on Cellular Automata
cs.LG
Submitted: 2026-09-17
Updated: 2026-09-17
License: http://creativecommons.org/licenses/by/4.0/
The gist: Recurrent-depth Transformers apply shared computation repeatedly, but differ in how they retain information across steps.
Terminology
Abstract
Recurrent-depth Transformers apply shared computation repeatedly, but differ in how they retain information across steps. We compare a Block Universal Transformer (BUT), which carries only its current hidden state; CoTFormer, which also retains an expanding attention cache; and a new LSTM Universal Transformer (LSTM-UT) with bounded gated memory. On Rule 30 cellular automata, BUT extrapolates to unseen recurrent depths more reliably than CoTFormer, although its accuracy eventually degrades. State and cache interventions show that CoTFormer's failure depends on their interaction: correcting the current state can temporarily restore accuracy, while retained history can undermine that correction. In a delayed-recall task, BUT also outperforms CoTFormer despite lacking direct access to past states; CoTFormer does not reliably select the requested cached representation. LSTM-UT improves both depth extrapolation and delayed recall over these baselines. The results support bounded gated memory as an effective inductive bias for repeated computation and later retrieval in these tasks.
Sources
- Universal Transformers
- Reasoning with Latent Thoughts: On the Power of Looped Transformers
- Loop, Think, & Generalize: Implicit Reasoning in Recurrent-Depth Transformers
- Why Limit the Residual Stream to Layers and Not Tokens? Persistent Memory for Continuous Latent Reasoning
- Dream to Control: Learning Behaviors by Latent Imagination
- Stabilizing Recurrent Dynamics for Test-Time Scalable Latent Reasoning in Looped Language Models
- Fixed-Point Reasoners: Stable and Adaptive Deep Looped Transformers
- Scaling Latent Reasoning via Looped Language Models
- Chain-of-Thought and Compressed Looped Transformers: A Memory-Budget Separation
- Controlling Transient Amplification Improves Long-horizon Rollouts
- Compression-based investigation of the dynamical properties of cellular automata and other systems
- Addressing Some Limitations of Transformers with Feedback Memory
- Attention Residuals
- The Residual Stream Is All You Need: On the Redundancy of the KV Cache in Transformer Inference
- Gated Recurrent Transformers: Expressive Depth through Recurrent Modulation
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks