The Orthogonalized Read Is a Removable Training Scaffold for Recurrent Memory
cs.LG
Submitted: 2026-07-02
Updated: 2026-09-14
Comments: 17 pages, 8 figures. Code, per-seed results, and checkpoints: https://github.com/no-way-labs/recurrent-memory-scaffold
Code: https://github.com/no-way-labs/recurrent-memory-scaffold
License: http://creativecommons.org/licenses/by/4.0/
The gist: Orthogonalizing the mLSTM memory matrix at read time with five differentiable Newton-Schulz iterations improves noisy associative recall.
Terminology
Abstract
Orthogonalizing the mLSTM memory matrix at read time with five differentiable Newton-Schulz iterations improves noisy associative recall. We replicate this effect and investigate its mechanism. Training on MAD noisy recall exhibits a long chance-level plateau followed by a sharp increase in accuracy. The orthogonalized read improves conditioning during this plateau and can be removed after escape. Ablations support three findings. First, the benefit requires a self-consistent read and gradient: an exact recursive least-squares read (the Mesa layer) yields a similar benefit, while straight-through variants, delta-rule writes, frozen random keys, and Frobenius normalization show no improvement over baseline. Second, across a learning-rate x task-difficulty grid, orthogonalization multiplies escape hazard roughly six-fold, with no detectable dependence on difficulty, and widens the range of learning rates that produce successful runs. Third, adding orthogonalization at inference leaves chance-level failures unresolved, while removing it gradually after escape yields standard mLSTMs at near-perfect accuracy. Schedule changes alone recover much of the reported gain. A batch-size x learning-rate analysis separates the effects of per-step learning rate and gradient noise on escape hazard (elasticities +3.0 and-1.65, respectively), linking the original vocab-96 result to its large-batch training regime. Direct decoding of the memory state recovers roughly half of the associations in behaviorally failed models, indicating a readout-learning limitation despite substantial stored information. These results show that fixed-budget recall benchmarks are sensitive to trainability and provide a tractable setting for investigating abrupt behavioral transitions through measurements of internal representations.
Sources
- Zoology: Measuring and Improving Recall in Efficient Language Models
- Simple linear attention language models balance the recall-throughput tradeoff
- xLSTM: Extended Long Short-Term Memory
- Old Optimizer, New Norm: An Anthology
- Noise-Driven Escape from Metastable Phases explains Grokking in Deep Neural Networks
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- An Embarrassingly Simple Way to Optimize Orthogonal Matrices at Scale
- Variational Linear Attention: Stable Associative Memory for Long-Context Transformers
- RWKV: Reinventing RNNs for the Transformer Era
- Mechanistic Design and Scaling of Hybrid Architectures
- Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets
- Uncovering mesa-optimization algorithms in Transformers
- MesaNet: Sequence Modeling by Locally Optimal Test-Time Training
- Test-time regression: a unifying framework for designing sequence models with associative memory
- Muon Outperforms Adam in Tail-End Associative Memory Learning
- Parallelizing Linear Transformers with the Delta Rule over Sequence Length
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks