Consolidator: Learning Persistent Routed Memory Across Context Boundaries
Sungwoo Goo, Hwi-yeol Yun, Sangkeun Jung
Chungnam National University
cs.LG, cs.AI
Submitted: 2026-08-12
Updated: 2026-08-13
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: The paper introduces Consolidator, a shared slot-local operator in a Phasor Memory Network (PMNet) that transforms routed short-term memory (STM) before accumulating it into long-term memory (LTM),
Terminology
Summary
The paper introduces Consolidator, a shared slot-local operator in a Phasor Memory Network (PMNet) that transforms routed short-term memory (STM) before accumulating it into long-term memory (LTM), without replaying source tokens. After each consolidation, the KV cache and STM are cleared, but retained LTM can still be read and is fed into the hierarchical router, conditioning which explicit-memory slots subsequent inputs access. The mechanism is evaluated on a two-segment modulo-10 mapping task where the second segment updates the mapping at the same memory address. With only 12.35K Consolidator parameters trainable (0.041% of a 29.95M model), direct LTM routing raises updated-mapping recall from 44.38 ± 1.94% to 87.02 ± 1.76% (+42.64 ± 1.10 percentage points), while immediate STM recall remains 89.90% in both conditions. Learned consolidation outperforms forced identity accumulation by 21.40 ± 1.91 percentage points without routing and 68.70 ± 1.76 with routing.
The paper addresses two linked questions: can a separately trained operator convert useful routed STM into LTM, and can that LTM guide later slot selection after the KV cache and STM are cleared?
The authors note that "merely detaching and copying STM into a slower buffer would provide carryover, but it would leave two mechanistic questions unresolved. Does a fixed pretrained memory interface require a learned transition to revise conflicting content, rather than raw accumulation? Does the retained state only supply values to the read path, or does the router also use it to select slots for subsequent writes and reads?"
For an occupied STM slot S ∈ R dm, the Consolidator defines z(S) = [cos S; sin S], then applies a gated MLP shared across all blocks, groups, and slots:
-
rψ(z) = Wd[SiLU(Wg z) ⊙ Wu z] + bd
-
Cψ(S) = atan2(sψ(S), cψ(S)) via element-wise complex multiplication in paired cosine/sine coordinates
The transform is slot-local and shares 12.35K parameters regardless of tree capacity. Zero output weights and unit-phasor bias make Cψ(S) = S in Equation (5), so learning begins from exact identity.
The boundary update is:
-
L+ b,g,j = (L b,g,j + Cψ(S b,g,j)) mod 2π when occupancy mask O b,g = 1
-
L+ b,g,j = L b,g,j when O b,g = 0
STM, occupancy, and KV state are then cleared while LTM remains.
The identity control replaces Cψ(S b,g,j) with raw S b,g,j, making it raw copy-and-accumulate, subject only to phase wrapping.
The base routing state is augmented with LTM: a t,b,g,j = u t,b,g,j + e b,g,j + L b,g,j. This creates the recurrent path: St → Cψ → Lt → a t+1 → S t+1, so a fixed router can make experience-dependent slot selections because its non-parametric LTM input changes. This path makes LTM an access state rather than only retrievable content.
Each memory episode contains two context segments and one active memory address selected from four address tokens. Rule families are ADD10 (y = (x + k) mod 10) and AFFINE10 (y = (ax + b) mod 10), with parameters resampled per episode. Each segment provides eight demonstrations and one held-out query. The final prediction therefore cannot receive credit for copying an observed answer or retaining only the stale rule.
-
Phase 1 (STM-pretraining): Trains same-segment rule induction from routed STM; Consolidator is frozen.
-
Phase 2 (Consolidation-training): Processes two demonstration segments sequentially, consolidates STM into LTM after each segment, then clears KV cache and STM.
Only retained LTM can provide the function parameters sampled for that episode.
-
Learned full: Entire model trainable (29.95M parameters)
-
Identity full: Raw STM accumulation, all except Consolidator trainable (29.94M)
-
Consolidator only: Only 12.35K Consolidator parameters trainable (0.041%)
-
Consolidator only, routing off: Same as above but removes direct LTM term in routing
-
Memory + Consolidator: Memory read/write/routing + Consolidator (1.526M, 5.095%)
-
Learned full, dual objective: Entire model with additional STM supervision
"Identity accumulation transfers the initial mapping but fails after the second same-address write. The learned transform instead reaches 87.02 ± 1.76% updated-mapping LTM recall, a same-checkpoint gain of 68.70 ± 1.76 pp over forced identity while training 0.041% of the model. The initial mapping after first consolidation shows 50.50 ± 3.42% for learned vs. 86.86 ± 0.05% for identity, reflecting
supervision only on the final updated mapping, not a claim about general retention."
"Direct routing raises updated-mapping LTM recall from 44.38±1.94% to 87.02 ± 1.76%, a paired gain of 42.64 ± 1.10 pp (95% CI [41.27, 44.01], p = 1.07 × 10-7), while immediate STM recall remains exactly 89.90% in both conditions. Without direct routing,
learned consolidation still exceeds forced identity by 21.40 ± 1.91 pp, showing that LTM remains useful through the read path; direct routing provides the larger additional gain."
Updated-mapping LTM recall from the same checkpoint falls from 87.02% with the correct experience to 9.30 ± 0.24% with a mismatched experience and 11.00% with fresh memory.
The mismatch preserves address and rule family while changing function parameters, demonstrating that recall depends on the consolidated content rather than a fixed address association or the mere presence of memory.
"Learned full reaches 93.70% and independently trained identity full reaches 91.34%; their +2.36 ± 3.30 pp difference is not statistically resolved. We therefore do not claim that learned consolidation dominates a fully plastic identity system on this task." Training only the memory path and Consolidator reaches 90.70 ± 0.51%.
Adding pre-consolidation STM supervision raises mean immediate recall across both segments from 11.39 ± 0.47% to 95.76 ± 0.72%, while updated-mapping LTM recall reaches 95.58 ± 0.75%.
The +1.88 ± 2.69 pp change in updated LTM recall is not statistically resolved (p = 0.193). The result therefore shows that one parameter set can support both capabilities under a suitable objective, not that both were jointly read in a single pass.
"Direct LTM routing improves both families: the simpler ADD10 family approaches saturation (99.01 ± 0.60% on vs. 52.52 ± 1.37% off), while AFFINE10 retains a paired gain of 38.56 ± 1.29 pp (74.80 ± 3.07% on vs. 36.24 ± 2.61% off)."
-
A shared slot-local transform that consolidates routed latent STM without replay or topology-dependent parameter growth.
-
Direct LTM-conditioned slot routing as an architectural inductive bias that lets retained state guide later writes and reads through a frozen router.
-
A sequential same-address update task that separates four memory functions: carrying state across a reset, revising an existing memory, retrieving retained content, and using LTM to guide slot selection.
-
Memory episodes contain two short context segments, one active address, and modular-arithmetic rules.
-
The five main consolidation-training seeds share one selected STM-pretraining representation, so their variance excludes variation from the first training stage.
-
Training retains gradients across the two consolidation boundaries, leaving detached or truncated long-horizon training untested.
-
Identity is the principal same-checkpoint control, but we do not compare alternative learned overwrite, EMA, linear, or gated recurrent operators.
-
LTM is initialized for each memory episode; persistence across unrelated sessions, serialization, and deployment restarts is not evaluated.
-
STM, LTM, and consolidation denote computational timescales, not a biological model of memory or sleep.
"Training only its 12.35K parameters while freezing the rest of PMNet yields 87.02% updated-mapping recall, compared with 18.32% when the same checkpoint uses identity accumulation; replacing the retained experience removes this gain. A paired routing ablation further reduces recall from 87.02% to 44.38% while leaving immediate STM recall unchanged, showing that consolidated LTM supports later computation both as retrievable content and as an input to slot selection. These results establish a controlled forward-state adaptation mechanism, not yet a general long-term memory system; detached or truncated long-horizon training, natural language, and scale-up remain open."
Improvements for AI systems
Improvements to AI Systems:
-
Add a learned, slot-local consolidation operator between short-term and long-term memory that transforms routed latent states via a gated MLP with identity initialization, rather than raw copy-and-accumulate. This enables conflict resolution when new information overwrites existing memory at the same address, improving recall of updated content by 68.70 percentage points over forced identity accumulation.
-
Use long-term memory as an input to the routing mechanism, not just as retrievable content. By adding the retained LTM vector to the router's state (a t = u t + e t + L t), the system can make experience-dependent slot selections for subsequent writes and reads, even after the KV cache and STM are cleared. This yields a 42.64 percentage point gain in updated-mapping recall over routing without LTM conditioning.
-
Clear the KV cache and STM after each consolidation boundary while retaining LTM, forcing the system to rely on consolidated state rather than raw token replay. This separates the mechanisms of carrying state across resets, revising existing memories, retrieving retained content, and using LTM to guide future slot selection—enabling more robust long-horizon operation without source token access.
-
Initialize the consolidation transform as identity (zero output weights, unit-phasor bias) so learning begins from exact copy behavior, then learns to modify only when necessary. This reduces training difficulty and provides a safe default when no conflicting update is present.
-
Add a dual-objective training regime that supervises both immediate STM recall and final LTM recall. This allows a single parameter set to support both short-term rule induction and long-term consolidation, raising immediate recall from 11.39% to 95.76% while maintaining high updated-mapping LTM recall (95.58%).
-
Use a shared, topology-independent operator (12.35K parameters regardless of model capacity) so the consolidation mechanism scales to larger networks without parameter growth proportional to memory slots or depth.
What the improved AI system can do:
-
Maintain and update associative memories across context resets without replaying source tokens, correctly revising prior mappings when new conflicting information arrives at the same memory address.
-
Use its consolidated memory state to actively guide which memory slots to access next, enabling experience-dependent attention and retrieval in sequential tasks.
-
Operate with a frozen pretrained backbone while only training a tiny consolidation module (0.041% of parameters), making it feasible to add persistent memory to existing large models without full fine-tuning.
-
Distinguish between content-dependent memory retrieval and mere address association—recall drops from 87% to 9.3% when the consolidated content is mismatched, showing the system truly stores and uses specific learned information.
-
Support both simple and complex rule families (e.g., additive and affine modular arithmetic) with significant gains from LTM-conditioned routing, while remaining robust to task complexity.
Abstract
Copying short-term memory (STM) into a slower store can preserve state across a context boundary, but persistence alone does not ensure that the retained state influences subsequent memory access. We test this distinction in a Phasor Memory Network (PMNet) using Consolidator, a shared slot-local operator that transforms routed STM before accumulating it into long-term memory (LTM), without replaying the source tokens. After each consolidation, the KV cache and STM are cleared. The retained LTM can still be read and is also fed into the hierarchical router, thereby conditioning which explicit-memory slots subsequent inputs access. We evaluate this mechanism on a two-segment modulo-10 mapping task in which the second segment updates the mapping at the same memory address. Following a second consolidation and reset, a held-out query must recover the updated mapping from LTM. The backbone and memory interface are frozen, leaving only 12.35K Consolidator parameters trainable (0.041% of a 29.95M model). Across five paired runs from the same STM-pretraining checkpoint, direct LTM routing raises updated-mapping recall from 44.38 plus or minus1.94% to 87.02 plus or minus1.76% (+42.64 plus or minus1.10 percentage points), while immediate STM recall remains 89.90% in both conditions; both train separate Consolidators and retain the same LTM read paths. Learned consolidation outperforms forced identity accumulation by 21.40 plus or minus1.91 percentage points without routing and 68.70 plus or minus1.76 with routing. Thus, on this task, consolidated LTM serves as both retrievable content and an access state that shapes subsequent slot selection.
Sources
- Attention Is All You Need
- Neural Turing Machines
- Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context
- Compressive Transformers for Long-Range Sequence Modelling
- Learning to (Learn at Test Time): RNNs with Expressive Hidden States
- Titans: Learning to Memorize at Test Time
- Phasor Memory Networks: Stable Backpropagation Through Time for Scalable Explicit Memory
- Recurrent Memory Transformer
- Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention
- Memorizing Transformers
- Using Fast Weights to Attend to the Recent Past
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- Context Distillation as Latent Memory Management
- Is One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL Training
- Replay in Deep Learning: Current Approaches and Missing Biological Elements
- Auto-Dreamer: Learning Offline Memory Consolidation for Language Agents
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks