Anatomy of Associative Recall in Fixed-State Recurrences: A Matched-State Decomposition, an Interference Wall, and a Curriculum That Breaks It
cs.LG, cs.CL
Submitted: 2026-09-14
Updated: 2026-09-14
Comments: 14 pages, 2 figures, 6 tables. Preprint of preliminary results; code and result JSONs at https://github.com/JIBSIL/dualgoose
Code: https://github.com/JIBSIL/dualgoose
License: http://creativecommons.org/licenses/by/4.0/
The gist: Fixed-state recurrences--linear attention and state-space models--are reported to lag behind attention on associative recall, but whole-architecture comparisons cannot say which ingredient is
Terminology
Abstract
Fixed-state recurrences--linear attention and state-space models--are reported to lag behind attention on associative recall, but whole-architecture comparisons cannot say which ingredient is responsible. We decompose masked multi-query recall at a fixed state budget along three single-knob axes: a short causal convolution, the transition structure (rank-1 delta rule vs. diagonal), and decay. The convolution dominates (+0.5 recall in both families under matched training): comparisons that pit convolution-free cells against a convolution-equipped Mamba measure the missing convolution, not the recurrence. The rank-1 transition beats its diagonal ablation by +0.19/+0.32 at 16/32 pairs, but the margin shrinks to +0.03 once both cells carry the convolution, and a state-matched Mamba-2 ties the unarmed rank-1 cell: no class claim survives. Cells that solve 32-pair recall degrade gracefully with load yet fall to chance retrieving 4 pairs from a distractor haystack--flat across lengths and transitions. Interference under sparse supervision, not capacity: a distance curriculum takes the unchanged architecture from 0.021 to 1.000. Training is a lock-in lottery--a seed either locks in or does not--and the curriculum is the lever. Lock-in rises from 1/10 to 7/10 (p=0.02); dense supervision adds nothing; at L=256 a shaped ramp reopens a boundary the uniform curriculum cannot (4/5 vs. 0/9); and at L=512, where the ramp collapses (0/6), gating it on measured accuracy locks in 6/6 (p=0.001). Bidirectional denoiser cells, reading the query before the haystack, show no measurable advantage over causal training (ten seeds), and collision-key retrieval needs two layers. Arming for recall is free on an S 5 state-tracking guardrail--the armed cell is significantly better at every depth (p<=0.0044). These replace "recurrent models are bad at recall" with a measured decomposition and two cheap interventions.
Sources
- Mechanistic evaluation of Transformers and state space models
- Zoology: Measuring and Improving Recall in Efficient Language Models
- Simple linear attention language models balance the recall-throughput tradeoff
- Just read twice: closing the recall gap for recurrent language models
- Structured Denoising Diffusion Models in Discrete State-Spaces
- The pitfalls of next-token prediction
- DeciMamba: Exploring the Length Extrapolation Potential of Mamba
- Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
- Kimi Linear: An Expressive, Efficient Attention Architecture
- Triplet-Block Diffusion RWKV
- Understanding and Improving Length Generalization in Recurrent Models
- Transformers Learn Shortcuts to Automata
- Stuffed Mamba: Oversized States Lead to the Inability to Forget
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
- The Illusion of State in State-Space Models
- Neural Networks and the Chomsky Hierarchy
- Revisiting associative recall in modern recurrent models
- Unlocking State-Tracking in Linear RNNs Through Negative Eigenvalues
- Quantifying Memory Utilization with Effective State-Size
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks