MMLA: Memory-Mediated Learning Architecture for Predictive Dual-State Adaptation
cs.CL, cs.LG
Submitted: 2026-06-27
Updated: 2026-09-14
Comments: 222 pages, 64 figures, 81 tables. Substantially expanded v4: updated architecture terminology; five integrated theory parts; additional post-training readout studies, including negative results; corrected proofs and causal statements; expanded related work. Project: https://github.com/MMLA-org/mmla-memory
Code: https://github.com/MMLA-org/mmla-memory
License: http://creativecommons.org/licenses/by/4.0/
The gist: Memory-Mediated Learning Architecture (MMLA) separates slow base parameters theta, a bounded numerical policy carrier Phi, and a bounded authoritative memory M.
Terminology
Abstract
Memory-Mediated Learning Architecture (MMLA) separates slow base parameters theta, a bounded numerical policy carrier Phi, and a bounded authoritative memory M. Predictive Dual-State Adaptation (PDSA) lets feedback update Phi while one problem remains active and lets a trusted lifecycle atomically commit one typed row or exact NULL. Later reasoning may read both states, but their writers, resets, rollback domains, and ledgers remain distinct. Realized futures supervise values only during training; deployment is causal and future-blind. We give conditional theory and falsifiable contracts for reasoning-time updates, completed-segment consolidation, predictive admission, authoritative memory, and dual-state attribution. Assumptions, counterexamples, capacity and cost ledgers, recovery duties, and identifying experiments are explicit; these are not implementation guarantees. Controlled studies show exact lifecycle execution on 300/300 held-out records for each of three seeds, calibrated retrieval gains over frozen-hidden dense and BM25 baselines, and exact typed anchor-filler transport on 240/240 held-out records per seed. These validate restricted components, not natural-language memory management or complete PDSA. Post-training studies retain positive and negative evidence. Later protocols obtain restricted readout progress, but a nine-trajectory comparison finds that matched latent readout, a full-width bridge, and bridge plus frozen text-teacher alignment all fail continuous-event qualification across both task families. A separately adapted text reference and restoration checks pass. At the September 13, 2026 evidence cutoff, no strict policy-only reasoning-time-training effect, predictive-admission oracle margin, learned future-blind admission policy, or policy-by-memory factorial advantage is established.
Sources
- Deep Variational Information Bottleneck
- Relational inductive biases, deep learning, and graph networks
- Memory Layers at Scale
- Improving language models by retrieving from trillions of tokens
- Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context
- The Llama 3 Herd of Models
- Neural Turing Machines
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- World Models
- MemOps: Benchmarking Lifecycle Memory Operations in Long-Horizon Conversations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- SnapKV: LLM Knows What You are Looking for Before Generation
- Landmark Attention: Random-Access Infinite Context Length for Transformers
- MemGPT: Towards LLMs as Operating Systems
- RWKV: Reinventing RNNs for the Transformer Era
- Hyena Hierarchy: Towards Larger Convolutional Language Models
- Qwen2.5 Technical Report
- Compressive Transformers for Long-Range Sequence Modelling
- Retentive Network: A Successor to Transformer for Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering