Drift and Dependence: Layer-wise Information-Theoretic Bounds for Replay-Based Continual Learning
Tieliang Gong, Zhongbo Zhang, Wen Wen, Yong-Jin Liu
Xi'an Jiaotong University · Tsinghua University
cs.LG, stat.ML
Submitted: 2026-08-12
Updated: 2026-08-13
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: The paper "Drift and Dependence: Layer-wise Information-Theoretic Bounds for Replay-Based Continual Learning" develops a layer-wise information-theoretic framework to analyze the generalization
Terminology
Summary
The paper Drift and Dependence: Layer-wise Information-Theoretic Bounds for Replay-Based Continual Learning
develops a layer-wise information-theoretic framework to analyze the generalization behavior of replay-based continual learning (CL). The central problem addressed is that replay—mixing a small buffer of past examples into current training—is effective against catastrophic forgetting, but its generalization behavior is shaped by two coupled effects that existing analyses fold into a single hypothesis-level quantity: (1) finite memory replaces each past distribution with an empirical proxy, and (2) repeated reuse couples the buffer, the current data, and the final hypothesis through a shared optimization trajectory.
The paper's main result, Theorem IV.1 (Hierarchical Synergy–Drift Bound), decomposes the expected generalization gap into a replay-induced representation drift and an optimization-dependence term. The theorem states: "Let W be the output of a learner minimizing the empirical risk L̂(W). Assume the loss function l(W, Z) is σ-subgaussian for all Z ∈ Z and W ∈ W. For any split layer l ∈ 0,..., L, we have genW ≤ (T − 1)√(2σ2K(l)) + √(2σ2/Neff (S(l) + P(l) − R(l) + C(l)))". Here, Neff = 1/((T−1)/m + 1/n) is the effective sample size, K(l) is the replay-centroid drift term measuring the KL divergence between the replay centroid and the population distribution, and the optimization variance term is further decomposed into stability (S(l)), plasticity (P(l)), interaction (R(l)), and residual-coupling (C(l)) components.
The paper explains that The replay-centroid drift term captures the bias induced by finite memory: at layer l, the mismatch between the replay centroid QAl,Yi,W1:l and the corresponding population distribution PAl,Yi,W1:l.
A key feature is that this term generally does not vanish as training proceeds, even though the raw exemplars are drawn i.i.d. from the task.
The optimization term measures the optimization dependence under the actual training representations,
with the prefactor √(2σ2/Neff) set by the effective sample size. The paper notes that "When replay is effectively unlimited (m → ∞), the old-task contribution (T − 1)/m vanishes and the current-task dependence recovers the usual O(1/√n) rate. In contrast, when the current-task size grows under a fixed replay budget (n → ∞), Neff approaches the finite limit m/(T − 1), so the variance term does not vanish but settles at a finite floor of order √((T − 1)/m), which we call the finite-memory variance floor."
The paper provides two corollaries of the main theorem. Corollary IV.2 (Input-Layer Bound) shows that at the input layer l = 0, the bound simplifies to genW ≤ √(2σ2/Neff I(S; W)), where S is the original data pair. The paper notes that the replay centroid QX,Yi,W coincides with the population PX,Yi,W and the training sequence is block-wise i.i.d., therefore K(0) = 0 and C(0) = 0.
Corollary IV.3 (Output-Layer Bound) shows that at the output layer l = L, the bound simplifies to genW ≤ (T − 1)√(2σ2K(L)) + √(2σ2/Neff C(L)), where there is no trainable mapping downstream of the representation, hence the stability–plasticity–synergy tradeoff over Wl+1:L disappears and S(L) = P(L) = R(L) = 0.
The paper then presents a Wasserstein relaxation of the drift term in Theorem IV.4 (Hierarchical Wasserstein Bound), which states: "Assume the loss function l(·, y) is ρ0-Lipschitz and the activation functions ϕl are ρl-Lipschitz. Then, at time T, we have: genW ≤ min l∈ 0,...,L E[ρ̄l(W) · Σ i=1 T W1(P̂Al,YSi,W1:l, PAl,Yi,W1:l)], where ρ̄l(W):= ρ0(1 ∨ Π h=l+1 L ρh∥Wh∥op). This relaxation
admits a finite, geometrically interpretable bound even under support mismatch, at the cost of losing the explicit structural decomposition of the optimization term. The layer minimization reveals a depth-dependent trade-off:
At shallow layers, representations tend to be shared across tasks, keeping W1 small, but any residual mismatch is amplified by a long suffix network, inflating ρ̄l(W). At deeper layers, this amplification shrinks while representations become increasingly task-specific, enlarging the drift. The optimal layer l⋆ is where these two competing effects are balanced."
This leads to Corollary IV.5 (Generalization Funnel Layer), which defines any minimizer of the geometric upper-bound proxy
as a generalization funnel layer. The paper states: "When drift increases with depth while suffix sensitivity decreases, this product attains its minimum at an intermediate layer, providing a bound-level rationale for intermediate-layer stabilization strategies such as feature distillation or partial freezing."
The paper then operationalizes the optimization term for stochastic gradient Langevin dynamics (SGLD) in Section V. Proposition V.1 (Heritage) splits the conditional mutual information into a heritage term and an increment term: I(UT(l); ΘR(T)W1:l) ≤ I(UT(l); Θ0(T)W1:l) + I(UT(l); ΘR(T)Θ0(T), W1:l) =: HT(l) + ∆T(l).
Corollary V.2 (Cumulative information budget) shows that the heritage term admits the following cumulative-budget upper bound: HT(l) ≤ Σ t=1 T−1 ∆t(l).
The paper explains: "Corollary V.2 reveals that HT(l) can be upper bounded by the cumulative sum of past increments. Therefore, to control the cross-task heritage HT(l), it suffices to control the sequence of within-task increments ∆t(l) t<T."
Theorem V.3 (Task-Wise SGLD-Based Bound) provides the main SGLD result: For task t trained by SGLD for Rt steps, we have genW(t) ≤ (t − 1)√(2σ2Kt(l)) + √(2σ2/Neff,t (Ct(l) + Σ s=1 t Σ r=1 Rs 1/2 E[log det(I + (ηs,r2/τs,r2) Ms,r(l))])).
The paper explains that this "reduces the suffix information complexity to a cumulative trajectory budget of log-determinants. This budget is governed by the ratio of learning rate to injected noise ηs,r2/τs,r2 and the conditional second-moment matrix Ms,r(l)."
The paper then performs a second-order expansion in Lemma V.4 and Proposition V.5 to separate the per-step cost into two orthogonal forces: Optimization Instability, driven by gradient covariance; and Interaction Cost, driven by the alignment of old-task and new-task mean gradients in a sensitivity-aware metric.
Proposition V.5 (Instability–mean decomposition) states: Let As,r(l):= I + αs,r Vs,r(l) ≻ 0 with αs,r = ηs,r2/τs,r2. Then 1/2 log det(I + αs,r Ms,r(l)) = 1/2 log det As,r(l) + 1/2 log(1 + αs,r µs,r(l)⊤ As,r(l)−1 µs,r(l)).
This leads to Definition V.6 (Sensitivity metric), which defines Hs,r(l):= As,r(l)−1 ≻ 0, and the induced cosine cosH(u, v):= ⟨u, v⟩H / (∥u∥H ∥v∥H). The paper explains: Geometrically, Hs,r(l) induces a stability-aware Mahalanobis metric: it discounts directions of high curvature while prioritizing flat subspaces.
Corollary V.7 (Alignment of old/new gradients) gives the interaction term a concrete geometric reading: "whether old and new task gradients align determines whether the per-step budget is small or large. When the two gradients align within the stable subspaces identified by H (constructive synergy), they reinforce a common descent direction... When they conflict (interference), the updates fight rather than reinforce."
The paper then provides interpretable diagnostics. Corollary V.8 (Scalar diagnostic decomposition) yields "three monitorable scalar quantities. tr(Σs,r(old)) measures gradient inconsistency within the replay buffer under the current representation, with elevated values serving as an early indicator of catastrophic forgetting when replay no longer provides stable optimization anchors. tr(Σs,r(new)) indicates the intrinsic stochasticity of the new task gradients. ⟨µs,r(old), µs,r(new)⟩ directly quantifies cross-task alignment, distinguishing positive transfer from destructive interference." Corollary V.9 (Sensitivity-weighted local mixing) derives a local mixing coefficient λ⋆ that minimizes the H-metric energy of the mixed mean update.
The experiments validate the theory across controlled and benchmark settings. In the controlled test of the effective-sample-size term, the paper reports: The fitted slope is βb = −0.443 ± 0.027 (95% bootstrap confidence interval [−0.498, −0.391], K = 200 replications per point), so the interval lies just above −1/2.
The paper also reports: "Across the 16 cells the proxy and the measured gap are not merely rank-correlated (Spearman ρ = 0.94 at the cell level) but essentially linear: a straight-line fit of the gap against 1/√Neff gives R2 = 0.97 (Pearson r = 0.98)."
For the drift–sensitivity trade-off, the paper reports: "The drift forms a pronounced 'bathtub': Dl is large at both the input and the final representation and small across a broad interior, so the drift–amplification product is minimized strictly inside the network in every one of the 12 cells. The funnel location
moves systematically deeper as the network deepens: its absolute index increases with L, and the Spearman's ρ between funnel index and depth falls between 0.8 and 1.0 across all three method–dataset pairs."
For the SGLD optimization-branch diagnostics, the paper reports in Table I that the alignment term cosH has task-controlled partial correlations with pairwise forgetting of −0.948 (ER) and −0.794 (DER++) on Split-CIFAR-100, and −0.796 (ER) and −0.645 (DER++) on Split-TinyImageNet, with confidence intervals that exclude zero. The paper notes: The sign is exactly the one the theory predicts: when old- and new-task gradients align in the stability-aware metric H, forgetting is small; when they anti-align, forgetting is large.
The paper also honestly states a scope boundary: "The iCaRL columns are included deliberately as a boundary of the framework rather than omitted... cosH is no longer a reliable predictor for iCaRL: its CI includes zero on Split-CIFAR-100 and the sign even reverses on Split-TinyImageNet."
The paper concludes: "We developed a layer-wise information-theoretic theory of replay generalization in continual learning. The central message is that replay generalization is governed by two coupled mechanisms: a replay-induced representation drift created by finite memory, and an optimization-complexity term created by shared downstream parameters. The main Synergy–Drift theorem isolates these two mechanisms in a single decomposition. The Wasserstein result revisits the first branch when KL-based drift is ill-posed under support mismatch, while the SGLD result instantiates the second branch through trajectory-level gradient moments. Together, these results provide a common language for studying replay bias, task interaction, and layer-wise sensitivity in continual learning."
Improvements for AI systems
Based on the paper, here are specific improvements for AI systems:
1. Layer-Aware Replay Buffer Selection
-
Improvement: Instead of replaying raw exemplars uniformly, select and weight replay samples per layer based on the drift term K(l). The system can compute which layers have the highest centroid drift and prioritize replaying examples that minimize drift at those critical layers.
-
Capability: The AI system can dynamically allocate limited replay memory to minimize representation drift at the
funnel layer
(where drift×amplification is minimized), improving generalization with the same buffer size.
2. Stability-Aware Gradient Alignment for Replay
-
Improvement: Use the sensitivity metric H (from Definition V.6) to weight replay gradients during optimization. Instead of naively mixing old and new task gradients, the system computes the H-metric alignment cosH between old and new task gradients and adjusts the replay learning rate or gradient magnitude accordingly.
-
Capability: The system can automatically detect when replay gradients are anti-aligned with new-task gradients (interference) and reduce their influence, or amplify them when aligned (constructive synergy), reducing catastrophic forgetting without manual tuning.
3. Early-Warning Forgetting Detector
-
Improvement: Monitor the scalar diagnostic tr(Σ(old)) (gradient inconsistency within replay buffer) as a real-time indicator. When this value spikes, the system triggers additional replay or regularization.
-
Capability: The AI system can predict imminent catastrophic forgetting before accuracy drops, enabling proactive intervention rather than reactive recovery.
4. Adaptive Learning Rate / Noise Scheduling for SGLD
-
Improvement: Use the cumulative log-determinant budget from Theorem V.3 to schedule learning rate η and noise τ. The system adjusts η2/τ2 per task to keep the information budget below a threshold that guarantees the generalization bound.
-
Capability: The system can automatically tune its stochastic optimization to balance stability and plasticity, avoiding both overfitting to new tasks and forgetting old ones.
5. Depth-Dependent Feature Freezing
-
Improvement: Use the generalization funnel layer (Corollary IV.5) to decide which layers to freeze or distill during continual learning. The system computes the drift-amplification product at each layer and freezes layers below the funnel while allowing deeper layers to adapt.
-
Capability: The AI system can automatically determine the optimal layer to stabilize, reducing computation and improving generalization without manual architecture design.
6. Replay Budget Allocation Across Tasks
-
Improvement: Use the effective sample size Neff = 1/((T−1)/m + 1/n) to optimally split a fixed total replay budget m across tasks. The system allocates more replay to tasks with smaller current-task sizes n, since the finite-memory variance floor √((T−1)/m) dominates when n is large.
-
Capability: The system can maximize worst-case generalization across all tasks given a fixed memory constraint, rather than using equal replay per task.
7. Task-Interaction-Aware Optimization
-
Improvement: Use the interaction term R(l) (from the Synergy–Drift theorem) to detect when old and new task gradients conflict in the H-metric. When R(l) is large, the system switches to a more conservative optimizer (e.g., lower learning rate or higher noise) for that layer.
-
Capability: The AI system can automatically adapt its optimization strategy per layer and per task based on measured gradient interactions, improving stability without sacrificing plasticity.
8. Support-Mismatch Detection
-
Improvement: Use the Wasserstein relaxation (Theorem IV.4) to detect when the replay buffer and current data have non-overlapping support (e.g., distribution shift). The system can flag this condition and switch to a more robust replay strategy (e.g., generative replay) when the KL-based drift is ill-posed.
-
Capability: The system can recognize when its replay buffer is no longer representative and trigger fallback mechanisms, preventing silent degradation.
9. Hyperparameter-Free Replay Weighting
-
Improvement: Use the local mixing coefficient λ⋆ (from Corollary V.9) to automatically compute the optimal weight for mixing old and new gradients in each update step, rather than using a fixed replay ratio.
-
Capability: The AI system can self-tune the replay-to-new-data gradient ratio at each step, maximizing stability without requiring manual hyperparameter search.
10. Forgetting-Aware Model Selection
-
Improvement: Use the heritage term H(l) (from Proposition V.1) as a model selection criterion. The system can compare different architectures or initialization strategies by their predicted cross-task information transfer, choosing the one with the lowest heritage bound.
-
Capability: The AI system can automatically select the architecture and initialization that minimizes expected forgetting before training begins, saving computation and improving final performance.
Sources
- An Empirical Investigation of Catastrophic Forgetting in Gradient-Based Neural Networks
- Continual Learning and Catastrophic Forgetting
- Progressive Neural Networks
- Information-theoretic Online Memory Selection for Continual Learning
- Efficient Lifelong Learning with A-GEM
- Gradient Projection Memory for Continual Learning
- Understanding Forgetting in Continual Learning with Linear Regression
- Information-Theoretic Generalization Bounds of Replay-based Continual Learning
- Understanding the Generalization Ability of Deep Learning Algorithms: A Kernelized Renyi's Entropy Perspective
- Why and When Deep is Better than Shallow: Implementation-Agnostic State-Transition Model of Deep Learning
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks