Divisive Normalization Shapes Low-Rank Slow Manifolds for Continuous Working Memory

arXiv:2608.01947 · q-bio.NC, cs.AI, cs.NE · Submitted 2026-08-22 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Divisive Normalization Shapes Low-Rank Slow Manifolds for Continuous Working Memory".

Jane: The paper was written by Zhaotian Gu, Jie Su, Weiwei Wang, Chang Liu, Tianyi Qian et al. from School of System Science, Beijing Normal University and State Key Laboratory of Cognitive Neuroscience and Learning, Beijing Normal University and Qiyuan Laboratory.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Paper discussion segment 1: Tom: So, we’ve just heard the title of this paper, "Divisive Normalization Shapes Low-Rank Slow Manifolds for Continuous Working Memory," and it sounds incredibly dense. Jane, could you help us break down what the authors are suggesting here in simple terms?

Jane: Absolutely. At its core, the paper is tackling a fundamental problem in AI: how do we make machine memory persistent? When current models try to remember something over time, they often drift or lose focus—a phenomenon that mimics forgetting. The authors propose a mathematical framework to keep that 'memory' stable and consistent.

Lalam: It’s about building stability into the very structure of the AI, rather than just hoping the training process solves it organically. That implies a deep understanding of how information needs to be managed over long stretches of time.

Meng: If I understand correctly, they are using concepts from manifold theory—which is really just a fancy way of describing a curved space—to show that certain mathematical constraints can keep the data living on a specific, predictable path.

Tom: A 'slow manifold,' they call it. What does that actually mean for the function of the AI? Is this just theoretical math, or does it have practical implications for building smarter systems?

Jane: It means that instead of wandering randomly through possibilities and forgetting what it was tracking, the data is mathematically guided to stay on a stable track. This allows the AI to maintain continuous variables—like remembering where an object was moving from hour to hour—without losing coherence.

Lu: From a computational neuroscience perspective, this really speaks to how biological brains manage ongoing tasks. They aren't just storing discrete facts; they are tracking continuous states, and the paper provides a way to model that stability computationally.

Meng: It’s essentially proposing an architectural feature that acts like an internal governor, preventing the system from becoming overly noisy or unstable when processing complex sequences of information.

Lalam: And this goes beyond just being able to remember *what* happened; it's about remembering *how* things are changing—the rates and dynamics. That level of continuous awareness is what defines sophisticated cognition in humans.

Tom: So, if we can successfully implement this mechanism, the AI could move past simple recall and into something that feels genuinely persistent. But how do we get from this mathematical blueprint to an actual working system? Jane, are there further details about the core mechanics of how they achieve this stability?

Jane: Yes. The key mechanism they identify is Divisive Normalization. While many AI models use simpler forms of normalization, the authors argue that by making it 'divisive' and constraining it to a 'low-rank' structure, they create exactly the kind of stable path needed for continuous memory tracking. This brings us to understanding how this mechanism works in practice, which is covered in the next section of the paper.

Paper discussion segment 2: Tom: We’ve established that Divisive Normalization Shapes Low-Rank Slow Manifolds for Continuous Working Memory gives us a solid mathematical framework for stability. Jane, could you summarize how the authors detail this mechanism and what that means for improving standard AI architectures?

Jane: Building on the concept of stability, the paper shows exactly *how* this normalization layer must interact with the existing RNN structure. They aren't suggesting an add-on; they are specifying a deep architectural change. It forces the internal calculations to adhere to those low-rank constraints, thereby smoothing out temporal variability and ensuring consistency over long sequences.

Meng: The 'low-rank' aspect is critical because it makes the system computationally manageable while still providing significant stability gains. It’s a clever way of constraining complexity without sacrificing too much capacity.

Lalam: It sounds like they are modeling a kind of resource limitation on the AI’s thought process, which is actually highly reflective of human cognition—we don't have infinite computational power or perfect memory bandwidth.

Lu: And what this implies architecturally is that the system learns to prioritize and consolidate information along the most efficient, stable pathways. It’s an optimization of cognitive resources within the model itself.

Tom: So, rather than just improving performance generally, they are fundamentally changing *how* the model processes time and dependency. Jane, could you elaborate on why this specific normalization technique is superior to other methods we might already be using?

Jane: Because it addresses the source of instability in standard RNNs: the accumulation of noise over time. The divisive nature acts like a continuous self-correction mechanism. Instead of letting small errors compound and cause drift, the system constantly normalizes its state relative to a stable, low-dimensional subspace.

Meng: If we think about it in terms of data flow, this normalization is keeping the information tightly bound to the 'slow' manifold—meaning the important variables change slowly and predictably—while filtering out high-frequency noise or temporary fluctuations.

Lalam: It’s moving us toward a system that doesn't just react to every little piece of data; it maintains a coherent, steady understanding of the underlying reality, which is key to deep understanding.

Lu: This conceptual shift is huge. We are moving from models that are excellent at short-term prediction to models that can maintain a consistent internal state for extended periods, mimicking true cognitive persistence.

Tom: It really does feel like we’ve found a blueprint for making AI think with more reliability and less jitteriness. But if this architecture is so effective, what happens next? Jane, does the paper suggest improvements beyond just implementing this normalization layer?

Jane: Yes, it absolutely does. The authors recognize that while this RDNN structure provides massive stability gains, the human brain doesn't learn via simple feed-forward calculations. They point to other complex biological learning mechanisms that we need to incorporate for a truly general AI. This brings us into the exciting realm of improving the learning rules themselves in segment three.

Paper discussion segment 3: Tom: We’ve covered how Divisive Normalization Shapes Low-Rank Slow Manifolds for Continuous Working Memory and established its role in stabilizing memory, but the authors point out that BPTT isn't how brains learn. Jane, what are the key architectural suggestions or additions they recommend to make the system more biologically plausible?

Jane: They are suggesting we move beyond standard backpropagation through time (BPTT). The paper suggests incorporating mechanisms inspired by biological plasticity and learning rules like Hebbian principles—the idea that 'neurons that fire together, wire together.' This shifts the focus from merely simulating learned dynamics to designing architectures that naturally *evolve* toward those stable manifolds.

Lu: That suggests a huge theoretical challenge: we need to design a learning rule, not just an architecture. We need the system to guide itself toward stability over time, much like natural selection guides evolution.

Meng: And I’m particularly interested in the practical implication of combining this divisive normalization with other inhibitory mechanisms. They study subtractive inhibition—which is different from divisive—and how those two forces interact within the same network structure.

Lalam: It feels like we are moving towards a much more integrated model of thought, where the AI isn't just storing a stable continuous memory but is also capable of juggling that data alongside highly discrete categories or immediate sensory inputs.

Tom: Juggling is such a good word for it, Lalam; it implies managing multiple streams of information at once—a multi-modal input stream. Jane, can you explain how these different types of inhibition interact within the system?

Jane: Essentially, the divisive normalization handles the broad, continuous background state. Then, introducing subtractive inhibition allows the system to perform more focused error detection or fine-tuning on specific inputs. The paper suggests investigating how these two complementary mechanisms reinforce each other.

Lu: We could explore if that core stabilizing effect

Conclusion: Tom: So, we’ve spent a lot of time unpacking the mechanics of Divisive Normalization and its impact on memory stability. To wrap up our discussion on this fascinating paper by Gu et al., I think the most important thing to take away is that we have a powerful blueprint for building truly persistent AI systems.

Jane: Exactly, Tom; it shows us how to architect a system so that its continuous variables stay stable and coherent, preventing that dreaded drift or shattering effect seen in standard AI models. It's all about creating the right environment for the memory to thrive.

Lu: From a theoretical standpoint, I find this to be such a beautiful bridge between natural processes and artificial computation; it really demonstrates how specific biophysical constraints can guide us toward an optimal mathematical solution that mimics robust cognitive function.

Meng: I agree with Lu, but from an engineering view, the impact is huge because we are finding a path to build far more reliable AI that doesn't fail under sustained operation—this leads to better performance on real-world tasks and more efficient hardware.

Lalam: And I think the long-term cultural implication of this research is equally significant; by modeling continuous working memory, we are essentially creating digital tools that support a higher level of sustained focus in our own lives.

Tom: Sustained focus is a great way to put it, Lalam; it moves us beyond simple processing and into continuous flow.

Jane: It’s not just about having a mechanism; it’s about ensuring that the mechanism works correctly under real-world conditions.

Lu: We are looking at moving toward complex models that mimic the entire range of neural behaviors, not just one specific function like integration.

Meng: And if we can successfully implement these multi-component systems, we could potentially deploy AI in environments where continuous state tracking is mission-critical.

Lalam: This moves us toward creating a form of digital persistence that mirrors our own capacity for sustained engagement with the world around us.

Tom: It truly sounds like the right time to conclude this discussion, after all these insights into "Divisive Normalization Shapes Low-Rank Slow Manifolds for Continuous Working Memory."

Jane: It’s clear that by building in this specific mechanism, the AI isn't just storing information; it’s storing it in a way that is structurally robust.

Lu: I believe this work provides the mathematical language we needed to design thinking systems more than just react to external stimuli.

Meng: We need to explore how these stable patterns can be scaled up and optimized for future AI deployment.

Lalam: And I hope that by moving forward with this architecture, we are building tools that help us achieve a more continuous and thoughtful engagement with our digital lives.

Tom: It sounds like the perfect moment to transition into our next paper, which will explore how these foundational concepts apply to different types of learning paradigms.

Zhaotian Gu, Jie Su, Weiwei Wang, Chang Liu, Tianyi Qian, Dahui Wang

School of System Science, Beijing Normal University · State Key Laboratory of Cognitive Neuroscience and Learning, Beijing Normal University · Qiyuan Laboratory

q-bio.NC, cs.AI, cs.NE

Submitted: 2026-08-22

Updated: 2026-08-25

Comments: Published in Transactions on Machine Learning Research (08/2026)

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 86/100

The gist: The study investigates how divisive normalization contributes to shaping low-rank, slow manifolds essential for continuous working memory tasks.

Key concepts

Continuous Working Memory
This refers to the ability the AI has to track variables—like an object's location—from hour to hour without losing coherence. The paper provides a mathematical way for data to stay on a stable, predictable path rather than wandering randomly or forgetting what it was tracking.
Divisive Normalization
This is the key mechanism of the paper. It is a specific normalization layer that forces internal calculations to adhere to strict constraints. This acts as a continuous self-correction mechanism, preventing small errors from compounding and causing the system to drift over time.
Low-Rank Manifolds
Manifolds describe curved spaces where data lives. By making the structure 'low-rank,' the system limits complexity while providing stability. This allows the AI to maintain a coherent, steady understanding of underlying data dynamics without becoming overly noisy or unstable.
Hebbian Principles
The authors suggest moving beyond standard backpropagation (BPTT). They incorporate biological learning rules, such as 'neurons that fire together, wire together.' This shifts the focus from simply simulating learned dynamics to designing architectures that naturally evolve toward stability.

Terminology

Summary

The study investigates how divisive normalization contributes to shaping low-rank, slow manifolds essential for continuous working memory tasks. The authors first establish theoretical limitations in existing architectures; for instance, they note that certain attention mechanisms can smear the non-stationary step transitions and cumulative updates required for path integration, and that their auto-correlation attention is structurally biased toward discovering global periodicity rather than local, causal temporal integration. Furthermore, the analysis points out a fundamental mathematical limitation regarding Transformers: "we note that the continuoustime dynamical systems analysis is mathematically inapplicable to Transformer-based models. Because these feedforward attention architectures lack a recurrent state equation h t = F(h t-1, x t), they do not possess a recursive hidden state space in which localized fixed points and tangent flows can be defined."

In evaluating long-horizon stability, the authors test trained networks on tasks like the angular velocity integration and memory-guided saccade up to 2048 time steps. The resulting long-term error dynamics are shown to exhibit a canonical trade-off in attractor networks, closely aligning with the theoretical framework of Ságodi et al. (2024). Specifically, while the short behavioral timescale is bounded by the uniform norm of the flow field, the long-term asymptotic error is dominated by the topology of the dynamics. In comparing models on these tasks, it was found that "both the RDNN and the Subtractive Network display nearly identical error degradation up to 2000 steps... confirming that under zero-input conditions, both models form functionally equivalent continuous ring attractors and thus suffer from the same long-term diffusion."

To test generalization to more complex systems, the researchers evaluated networks on the Double Angular Velocity Integration task, which requires tracking two independent circular variables on a two-dimensional torus manifold. The results demonstrated that the proposed RDNN maintains a superior performance compared to standard baselines, demonstrating its robustness in high-dimensional continuous tracking. This advantage is attributed to the underlying topological properties of the learned manifolds: "Our analysis (Figure A32, top) reveals that the RDNN is predominantly governed by stable and uncertain limit cycles, reflecting a fluid, continuous flow on the torus surface. In contrast, the GRU and LSTM are overwhelmingly dominated by stable fixed points, indicating that they solve the double integration by discretizing the 2D torus into a rigid grid of point-attractor basins. The authors conclude that These results indicate that the biological divisive normalization bottleneck is critical for stabilizing higherdimensional continuous attractors."

Finally, concerning computational overhead, the analysis compared several architectures. While all are state-space networks with an identical single-step asymptotic time complexity of O(H 2), their leading-order operations differ significantly. The authors detail that "the RDNN requires 4H squared recurrent floating-point operations (FLOPs) per step... which is lower than the TaskGRU (6H squared), TaskLSTM (8H squared), and ORGaNICs (12H squared, due to its multiple gating and modulation matrices)."

Improvements for AI systems

1. Implementation of Recurrent Divisive Normalization Network (RDNN) for State Maintenance

We propose replacing traditional gated units (LSTM/GRU) with a Divisor-Gated architecture, specifically the RDNN structure derived from this paper. This involves implementing an auxiliary inhibitory pool (G) that dynamically tracks overall activity and dividing the excitatory state (R) by (eta + G).

  • Capability:

  • Robust Continuous State Representation: The system will be able to represent continuous variables (e.g., position, accumulated velocity, or spatial coordinates) with high fidelity without suffering from the shattering or discretization inherent in standard gated RNN architectures.

  • Stable Long-Term Integration: It can maintain and update continuous variables over prolonged zero-input delays (autonomous maintenance) by converging to a stable slow manifold characterized by marginal fixed points, preventing state drift.

2. Utilization of Activity-Dependent Gradient Attenuation for Optimization

We exploit the inherent mechanism of divisive normalization—the inverse scaling of gradients by the dynamic divisor (eta + G) —as a form of implicit regularization during Backpropagation Through Time (BPTT).

  • Capability:

  • Implicit Low-Rank Inducement: The system naturally restricts parameter updates to a subset of marginally active states. This activity-dependent attenuation guides the network toward an emergent, low-dimensional subspace (sub-linear effective rank), achieving strong self-compression without requiring explicit, computationally expensive low-rank factorization.

  • Optimized Training Stability: By avoiding the non-convex optimization pathologies associated with forced matrix factorization, training stability is enhanced. The system achieves faster and more reliable convergence in complex tasks involving continuous state dynamics.

3. Application of Tripartite Spectral Analysis for Continuous Control Systems

We utilize the observed spectral properties of RDNN to ensure stability in real-time control loops that require continuous integration (e.g., autonomous vehicle navigation or robotic trajectory planning).

  • Capability:

  • Guaranteed Normal Hyperbolicity: The architecture ensures a structured Jacobian spectrum (tripartite: fast dissipation modes, slow maintenance modes, and off-axis conjugate pairs) that maintains the required level of normal hyperbolicity. This guarantees that the continuous state space can be smoothly traversed without sudden collapse into discrete attractors when integrating time-varying external inputs.

  • Robust Path Planning: The system can fluidly drive its state along a continuous manifold, enabling precise and stable path execution in dynamic environments where traditional discrete decision boundaries are insufficient.


The improved AI system will possess the following core competencies:

  1. Precise Analog Memory: Store, maintain, and update complex continuous states (e.g., theta in [0, 2 pi)) with high precision (about 10-4 drift), far exceeding the stability of traditional discrete-state models.

  2. Continuous Trajectory Planning: Execute integration tasks (e.g., angular velocity accumulation) by smoothly driving its internal state along a continuous attractor manifold, ensuring no shattering or abrupt jumps in the output path.

  3. Efficient Learning: Achieve optimal performance through an implicit, activity-dependent regularization mechanism that biases the network toward a low-dimensional subspace, thereby reducing the computational overhead of explicit structural constraints.

Abstract

The ability to robustly maintain and update continuous variables is a hallmark of working memory. While classical continuous attractor networks suffer from severe fine-tuning fragility, standard artificial recurrent neural networks (RNNs) like GRUs and LSTMs typically fail to stably learn continuous manifolds, instead shattering the state space into discretized point attractors. To bridge this gap, we draw inspiration from divisive normalization, a canonical neural computation widely observed across cortical circuits, and propose the Recurrent Divisive Normalization Network (RDNN), a minimal and algebraically isolated model of dynamic division. Through dynamical systems analysis on canonical working memory tasks, we demonstrate that this biophysical constraint allows the network to converge to robust, high-fidelity slow manifolds. Furthermore, we analyze the gradient dynamics of divisive normalization during Backpropagation Through Time (BPTT), showing that it introduces an activity-dependent local gradient scaling. This scaling dampens parameter updates in highly active regimes, which empirically aligns with a significant self-compression of the network's effective rank, confining the recurrent dynamics to a tight, low-dimensional subspace while avoiding the optimization pathologies associated with explicit low-rank factorization. Finally, ablations demonstrate that while subtractive inhibition can maintain static memories, divisive normalization is mathematically essential to prevent manifold shattering under time-varying inputs. Our findings identify divisive normalization not merely as a biological artifact, but as a critical computational mechanism for learning high-fidelity continuous representations.

Related papers