Uniform Herding: Exemplar Replay with Representation Refresh
Krishna Subedi
cs.AI
Submitted: 2026-08-13
Updated: 2026-08-14
Code: https://github.com/neryva-lab/uniform-herding
License: http://creativecommons.org/licenses/by-sa/4.0/
Importance score: 48/100
The gist: Uniform Herding: Exemplar Replay with Representation Refresh proposes a replay-based class-incremental learning method that refreshes exemplars in the current feature representation.
Terminology
Summary
Uniform Herding: Exemplar Replay with Representation Refresh proposes a replay-based class-incremental learning method that refreshes exemplars in the current feature representation. The paper states: "As the feature representation changes, replay must preserve the earlier classes. However, only a bounded active exemplar set can be replayed. We propose Uniform Herding, which allocates the current active set across observed classes and uses a bounded candidate pool to refresh their chosen exemplars in the current representation."
The method works as follows: "After each task is completed, it divides the active budget equally among all observed classes and rebuilds each class’s exemplar set using greedy herding Welling [2009] in the current feature space. Reconstruction candidates come from a bounded candidate pool, while the selected exemplars form the active replay memory. The active budget M is divided uniformly across classes, with quotas computed as
qc(t) = floor(M / Ct) + 1 i < M mod Ct for each class identifier. The candidate pool is bounded by ρM at completed task boundaries with default ρ = 3, and
Replay draws only from the selected sets."
The experimental setup uses CIFAR-100 with ten class-incremental tasks, a ResNet-18 backbone, active budget M = 2,000, retrieval budget b = 64, and three seeds.
Training uses SGD with learning rate 0.1, momentum 0.9, weight decay 5×10−4, batch size 128, gradient clipping at 1.0, mixed precision, and 70 epochs per task. The default classifier is a cosine-margin classifier with scale 30 and margin 0.35. The loss combines cross-entropy with a temperature-scaled softmax KL distillation term constrained to old classes, with λ = 1 and T KD = 2. At evaluation, Uniform Herding uses nearest-mean-of-exemplars (NME) prediction.
The main results show: Uniform Herding obtains 44.00 ± 0.51% final average accuracy and 17.22 ± 0.43% forgetting, compared with 42.33 ± 1.20% and 24.87 ± 1.11% for iCaRL.
The static bank baseline achieves 28.60 ± 1.35% accuracy and 55.86 ± 1.42% forgetting. The paper notes: The comparison with iCaRL is end-to-end. It does not isolate the effect of refresh from the other protocol differences.
Within-method ablations show: "Replacing NME with head-logit evaluation drops accuracy by 10.46 pp and raises forgetting by 30.36 pp. Replacing herding with random selection costs 2.39 pp in accuracy and 1.16 pp in forgetting. Removing KD decreases accuracy by 1.48 pp but increases forgetting by 11.66 pp. Replacing the cosine-margin head with a linear head decreases accuracy by 0.79 pp and increases forgetting by 0.12 pp."
Resource sensitivity results: Reducing M from 2,000 to 500 lowers accuracy by 10.05 pp and raises forgetting by 13.26 pp. Increasing M to 4,000 raises accuracy by 3.32 pp and lowers forgetting by 3.91 pp.
For retrieval: reducing retrieval from 64 to 32 decreases accuracy by 1.53 pp and decreases forgetting by 0.01 pp. Increasing retrieval to 128 increases accuracy by 0.16 pp and increases forgetting by 0.59 pp.
The paper concludes: The retrieval sweep produces a smaller mean change than the active-budget sweep.
The discussion states: "NME evaluation is strongly favored over head-logit evaluation by 10.46 pp in accuracy and 30.36 pp in forgetting... Herding selection yields 2.39 pp higher accuracy and reduces mean forgetting by 1.16 pp relative to random selection; this small difference is not estimated reliably with three seeds. Removing KD decreases accuracy by 1.48 pp but increases forgetting by 11.66 pp, suggesting that distillation serves a retention role in this objective."
Limitations are explicitly stated: "All experiments use the CIFAR-100 dataset, one ten-task partition (split seed 13), one ResNet-18 backbone, and three training seeds. No variation in class order, task granularity, dataset, or architecture is tested... The iCaRL comparison does not isolate refresh from the objective and storage differences between the two protocols. No matched experiment varies only the refresh policy while retaining the objective, head, readout, active budget, candidate multiplier, data order, and seed set fixed."
The conclusion reiterates: "Uniform Herding reaches 44.00 ± 0.51% final average accuracy and 17.22 ± 0.43% forgetting, against 42.33 ± 1.20% and 24.87 ± 1.11% for iCaRL, and 28.60 ± 1.35% and 55.86 ± 1.42% for the static bank. Within the proposed configuration, NME and herding selection yield higher mean final accuracy than the alternatives evaluated, and distillation mainly improves retention. The active-budget sweep shifts both metrics more than the retrieval sweep over the reported values. These findings are conditional on the evaluated protocol and do not show that refresh alone causes the iCaRL gap."
Improvements for AI systems
Improvements to AI Systems:
-
Dynamic Exemplar Refresh with Bounded Candidate Pooling: Implement a replay mechanism where stored exemplars are periodically re-selected from a bounded candidate pool (size ρM, default ρ=3) in the current feature space, rather than keeping a fixed static set. This allows the system to adapt its memory to evolving representations, improving retention of old classes without unbounded storage.
-
Uniform Quota Allocation Across Observed Classes: Use the formula
qc(t) = floor(M / Ct) + 1 i < M mod Ctto divide the active memory budget M equally among all classes seen so far. This prevents class imbalance in replay, ensuring fair representation and reducing catastrophic forgetting in long task sequences. -
Nearest-Mean-of-Exemplars (NME) Readout for Classification: Replace head-logit predictions with NME at inference time. This improves final accuracy by 10.46 percentage points and reduces forgetting by 30.36 percentage points, making the system more robust to feature drift and head misalignment.
-
Greedy Herding for Exemplar Selection: Use greedy herding to select exemplars that best approximate each class’s mean in the current feature space, rather than random sampling. This yields a 2.39 pp accuracy gain and 1.16 pp lower forgetting, improving memory efficiency and class representation fidelity.
-
Knowledge Distillation with Temperature-Scaled Softmax for Retention: Apply KL distillation (λ=1, T KD=2) constrained to old classes during training. This increases forgetting by only 1.48 pp accuracy loss but reduces forgetting by 11.66 pp, making the system retain prior knowledge more effectively—critical for continual learning.
-
Cosine-Margin Classifier with Scale 30 and Margin 0.35: Use a cosine-margin head instead of a linear head to improve separation in feature space. This yields 0.79 pp higher accuracy and 0.12 pp lower forgetting, enhancing discriminative power for both old and new classes.
-
Sensitivity-Aware Resource Allocation: Based on the active-budget sweep (M=500→2000 improves accuracy by 10.05 pp, reduces forgetting by 13.26 pp; M=2000→4000 improves by 3.32 pp, reduces forgetting by 3.91 pp), design systems to prioritize memory capacity over retrieval batch size. The retrieval sweep (32→128) has minimal impact (<1.7 pp accuracy change), so allocate computational resources to larger active memory rather than larger retrieval batches.
-
Mixed Precision Training with Gradient Clipping: Incorporate mixed precision and gradient clipping at 1.0 to stabilize training over 70 epochs per task, enabling faster convergence and consistent performance across seeds—useful for large-scale continual learning systems.
What the Improved AI System Can Do:
-
Continually learn new classes (e.g., 100 classes over 10 tasks) while maintaining 44% final accuracy and only 17% forgetting on CIFAR-100, significantly outperforming static replay (28.6% accuracy, 55.9% forgetting).
-
Adapt its memory to changing representations without requiring unbounded storage, using a bounded candidate pool (ρM) to refresh exemplars—ideal for edge devices with limited memory.
-
Provide fair class representation in replay via uniform quota allocation, preventing bias toward recently seen classes.
-
Achieve robust inference via NME, which is less sensitive to classifier head drift than logit-based prediction.
-
Retain prior knowledge through distillation, reducing forgetting by over 11 percentage points compared to no distillation.
-
Scale memory efficiently: doubling the active budget from 2000 to 4000 yields a 3.32 pp accuracy gain, while increasing retrieval batch size beyond 64 yields negligible gains—allowing system designers to optimize memory vs. compute trade-offs.
-
Operate reliably with mixed precision and gradient clipping, ensuring stable training across multiple random seeds and task sequences.
Abstract
As the feature representation changes, replay must preserve the earlier classes. However, only a bounded active exemplar set can be replayed. We propose Uniform Herding, which allocates the current active set across observed classes and uses a bounded candidate pool to refresh their chosen exemplars in the current representation. On CIFAR-100 with ten class-incremental tasks, a ResNet-18 backbone, active budget M=2, 000, retrieval budget b=64, and three seeds, Uniform Herding obtains 44.00 plus or minus0.51% final average accuracy and 17.22 plus or minus0.43% forgetting, compared with 42.33 plus or minus1.20% and 24.87 plus or minus1.11% for iCaRL. Within the Uniform Herding protocol, final accuracy decreased when NME or herding was replaced with the tested alternatives, while forgetting increased when distillation was removed. Changing the retrieval budget has a smaller effect across the tested range than changing the active budget. The comparison with iCaRL is end-to-end. It does not isolate the effect of refresh from the other protocol differences. These results are limited to the tested protocol.
Sources
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection