Stochastic Engrams for Efficient Continual Learning
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Stochastic Engrams for Efficient Continual Learning".
Jane: The ability to learn continuously in artificial neural networks (ANNs) is often limited by catastrophic forgetting, a phenomenon in which new knowledge becomes dominant.
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So, to summarize this paper, "Stochastic Engrams for Efficient Continual Learning," they are addressing the fundamental issue of catastrophic forgetting in artificial neural networks when they try to learn sequentially. The core thesis is that by drawing inspiration from memory encoding in neuroscience—specifically engrams—they propose a novel approach using stochastically-activated engrams as a gating mechanism within metaplastic binarized neural networks.
Jane: They claim this method leverages the computational efficiency inherent in binarized neural networks while simultaneously employing the robustness of probabilistic memory traces to actively mitigate forgetting and keep the model reliable across different tasks. This is designed to be an energy-efficient architecture that combines those specific elements.
Lu: What’s particularly compelling is their formulation of context-dependent, engram-inspired gating that operates directly on a binarized latent space, which they say unlocks robust memory retention without sacrificing the computational advantages of using binarized networks.
Meng: They are aiming to provide a framework where the network can manage its synaptic state in a way that is both stable and plastic, essentially balancing the need to remember old information with the need to learn new information in a constrained environment.
Lalam: This research suggests that we might be able to build AI systems that don't just memorize data for one task but develop a more cohesive and persistent knowledge structure across many different experiences.
Tom: And they’ve done some strong validation, claiming this specific strategy is the only one capable of achieving average accuracies over seventy percent in both class-incremental and domain-incremental MNIST benchmarks, matching the performance of full-precision state-of-the-art methods <ref:2503.21436#pg0,capable of achieving average accuracies over 70% in both class-incremental and>.
Jane: That result is significant because it shows that their method isn't just theoretically sound; it actually delivers high performance metrics on standard test sets while achieving a notable reduction in peak GPU and RAM usage compared to other methods.
Lu: The focus there seems to be on demonstrating that this integration of engram gating directly within the latent space is a viable path forward for architectures that need to perform well in sequential learning scenarios.
Meng: If they can maintain high accuracy while keeping the computational footprint small, then it moves from a theoretical concept to something practically implementable for real-world edge devices where memory and processing power are limited.
Lalam: For the future, this points toward AI that can be deployed more widely because it solves one of the most persistent problems in making sequential learning practical across various platforms.
Conclusion: Tom: So, wrapping up our discussion on "Stochastic Engrams for Efficient Continual Learning," we have to look at how this work by Aguilar, Herbozo Contreras, and Kavehei impacts the field of continual learning. It’s about moving beyond simply trying to patch forgetting with standard techniques.
Jane: They essentially show that by using these stochastically-activated engrams as a gating mechanism in their metaplastic binarized neural networks, they can achieve solid performance metrics, like those over seventy percent accuracy on MNIST tasks, without the prohibitive hardware costs associated with more complex learning strategies <ref:2503.21436#pg0,stochastically-activated engrams as a gating mechanism>.
Lu: The implication here is that we can start designing more efficient architectures where memory management is intrinsically linked to the way the network processes information in its latent space, which could lead to fundamentally different kinds of AI systems.
Meng: For practical deployment, it means we might see AI models that are much more robust when they encounter new data streams because their internal mechanisms for remembering and forgetting are better managed by this gating system.
Lalam: I think the bigger impact is on the overall development trajectory of AI; if we can make foundational learning more sustainable across tasks, it opens up possibilities for truly adaptive and continuously evolving intelligence in a way that feels more natural.
Tom: It really shows that inspiration from biology isn't just academic; it’s a concrete way to engineer better solutions for the practical limitations we face when training large AI models sequentially.
Jane: Precisely, the paper demonstrates how to merge computational efficiency with robust memory traces, offering a tangible strategy to maintain model reliability in dynamic learning environments.
Lu: This work provides a solid foundation for future research exploring how different types of gated mechanisms can be adapted to other continual learning challenges, pushing the boundaries of what we think is possible for sequential AI.
Meng: I’ll be keeping an eye on how engineers start applying this concept to more complex, real-world datasets where resource constraints are even tighter than MNIST.
Lalam: It’s exciting because it shows that deep inspiration from other fields can lead to tangible improvements in the stability and adaptability of AI systems we build every day.
School of Biomedical Engineering, The University of Sydney
cs.LG
Submitted: 2025-03-27
Updated: 2026-09-28
Code: https://github.com/NeuroSyd/engramBNN
Project page: https://vlomonaco.github.io/core50
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 76/100
The gist: The ability to learn continuously in artificial neural networks (ANNs) is often limited by catastrophic forgetting, a phenomenon in which new knowledge becomes dominant.
Key concepts
- Catastrophic Forgetting
- This is the problem where an AI model, when trained on new information, rapidly loses its ability to perform tasks it previously learned. New knowledge overwhelms old knowledge, causing performance on previous tasks to drop sharply.
- Metaplasticity Function
- This is a mechanism within the network that emulates biological synaptic plasticity. It controls how the network updates its weights based on new data, allowing it to stabilize and consolidate memory for older tasks while remaining sensitive enough to learn new ones.
- Engram Gating Block
- This block acts as a filter for incoming information. It uses stochastic activation and sparsity constraints (like top-k masking) to decide which parts of the network's representation should be kept active, effectively gating the memory retention process.
- Binarized Neural Networks (BNNs)
- These are neural networks where all weights and activations are restricted to binary values, typically 0 or 1. This approach is used here because it offers computational efficiency while still allowing for complex learning capabilities.
Terminology
Summary
The ability to learn continuously in artificial neural networks (ANNs) is often limited by catastrophic forgetting, a phenomenon in which new knowledge becomes dominant. By taking mechanisms of memory encoding in neuroscience (i.e., engrams) as inspiration, we propose a novel approach that integrates stochastically-activated engrams as a gating mechanism for metaplastic binarized neural networks (mBNNs). This method leverages the computational efficiency of mBNNs combined with the robustness of probabilistic memory traces to mitigate forgetting and maintain the model’s reliability.
The gist
Our method is the only strategy capable of achieving average accuracies over 70% in both class-incremental and domain-incremental MNIST benchmarks, matching full-precision state-of-the-art methods, while achieving a significant reduction in peak GPU and RAM usage.
Bio-Inspired Architecture and Mechanism
The proposed architecture is called engramBNN, which integrates binarized neural networks (BNNs) with neuro-inspired metaplasticity and dynamic engram gating. The core novelty lies in the formulation of context-dependent, engram-inspired gating that operates directly on a binarized latent space,
unlocking robust memory retention without sacrificing computational advantages. This architecture is designed to be an energy-efficient architecture that combines binarized networks with neuro-inspired metaplasticity and dynamic engram gating.
Engram Gating Block
The engram gating block is introduced to promote task-specific representations and mitigate interference during continual learning. For a given input sample, the process involves:
-
A single shared two-layer full-precision encoder projecting the input into a lower-dimensional latent space to extract essential spatial context.
-
Separate linear layers projecting outward from this shared latent space to generate
distinct, independent gating vectors (masks) for each corresponding hidden layer of the metaplastic BNN.
-
These projection layers apply a stochastic activation and a top-k mask to generate discrete 0, 1 gating values, with the parameter k set to retain
20% of the units
(e.g., k = 409 for a 2048-dimensional layer), corresponding to an80% sparsity constraint.
-
A
Bernoulli sampling is applied element-wise to˜p to yield a discrete binary mask M ∈ [0, +1]
during training, while non-top-k units are strictly zeroed out.
Metaplasticity Function
Artificial metaplasticity is implemented as a second continual learning component within the BNN backbone to emulate synaptic plasticity and mitigate catastrophic forgetting. This mechanism is governed by the function:
fmeta(m, Wh) = 1 − tanh2(m · Wh) (4).
The hyperparameter 'm' controls the network’s tendency to consolidate weights towards a task, influencing its sensitivity to new data. The update rule involves applying the standard weight update (UW) and then modulating it via fmeta based on the model’s current synaptic state. This mechanism presents a stability-plasticity trade-off,
where a higher 'm' value helps strengthen retention of older tasks but may restrict learning on newer tasks.
Performance and Efficiency Metrics
The performance of engramBNN is evaluated using several metrics, including:
-
Average Accuracy (ACC): The average of the final test accuracy across all tasks seen.
-
Weighted Accuracy (wACC): A metric that
prioritizes long-term memory stability and acts as a stress test rather than a balanced metric,
by scaling the final accuracy of each task by a weight wi = eλti, where ti is theage
of the task relative to the end of training. -
Forward Transfer (FWT): Evaluates forward knowledge transfer from past tasks to future ones.
-
Backward Transfer (BWT): Quantifies how prior task knowledge is retained or forgotten during subsequent task learning, defined as BWT = 1/(T − 1) Σ[AT,i - Ai,i].
-
Remembering (REM): A modified metric based on BWT, defined as REM = 1 − min(BWT, 0).
The results demonstrate that engramBNN achieves the highest overall average test accuracy (∼ 73%) while maintaining a minimal computational footprint (∼ 5% peak GPU usage),
validating it as a strategy to effectively mitigate catastrophic forgetting without incurring prohibitive hardware overhead. Furthermore, in the Split MNIST domain-incremental experiments, engramBNN achieved the highest ACC and weighted ACC of 75.34% and 79.47%, respectively.
Ablation and Analysis
Ablation experiments on the class-incremental Split MNIST setting revealed that engramBNN exhibited the most well-defined clusters
in feature visualization (Fig. 5), achieving the "highest score of 0.
Improvements for AI systems
Based on the provided scientific paper, here are specific improvements that can be made to existing AI systems, as well as a description of what these improved systems could achieve:
Primary Improvements:
-
Improve Catastrophic Forgetting Mitigation via Stochastic Context-Dependent Gating (EngramBNN):
-
Integrate Metaplasticity into Binarized Architectures for Enhanced Synaptic Stability:
-
Reduce Computational and Memory Footprint by Leveraging Binarization with Engram Routing:
Specific System Capabilities of the Improved AI Model (engramBNN):
Detailed Breakdown of System Improvements and Capabilities:
---Detailed Breakdown of System Improvements and Capabilities:
---Specific System Capabilities of the Improved AI Model (engramBNN):
Abstract
The ability to learn continuously in artificial neural networks (ANNs) is often limited by catastrophic forgetting, a phenomenon in which new knowledge becomes dominant. By taking mechanisms of memory encoding in neuroscience (i.e., engrams) as inspiration, we propose a novel approach that integrates stochastically-activated engrams as a gating mechanism for metaplastic binarized neural networks (mBNNs). This method leverages the computational efficiency of mBNNs combined with the robustness of probabilistic memory traces to mitigate forgetting and maintain the model's reliability. Previously validated metaplastic optimization techniques have been incorporated to further enhance synaptic stability. Compared to baseline binarized models and benchmark fully connected continual learning approaches, our method is the only strategy capable of achieving average accuracies over 70% in both class-incremental and domain-incremental MNIST benchmarks, matching full-precision state-of-the-art methods. Furthermore, we achieve a significant reduction in peak GPU and RAM usage, under 5% and 20%, respectively, as well as an 8x reduction in memory footprint compared to full precision counterparts. Our findings demonstrate (A) an improved stability vs. plasticity trade-off, (B) reduced memory intensiveness, and (C) enhanced performance in binarized architectures. By uniting principles of neuroscience and efficient computing, we offer new insights into the design of scalable and robust deep learning systems.
Sources
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks