Towards Realistic Class-Incremental Learning with Free-Flow Increments
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Free-Flow Class-Incremental Learning: Towards Robust CIL under Variable Class Arrivals".
Jane: The paper was written by Zhiming Xu, Baile Xu, Jian Zhao, Furao Shen and Suorong Yang from Nanjing University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Alright, welcome back to the show, everybody. Today we are digging into a fresh arXiv paper that's got a mouthful of a title: "Towards Realistic Class-Incremental Learning with Free-Flow Increments." Jane, I have to say, just reading that title got me excited, because it's poking at something that's been bugging me for a while.
Jane: Oh, absolutely, Tom. And for our listeners who might not be deep in the machine learning weeds, let me break down what we're even talking about here. Normally, when we train a model to recognize, say, a hundred different objects, we give it all the pictures at once. But in the real world, new objects show up over time. That's class-incremental learning — the model learns a few new things, then a few more, and it has to remember the old stuff without forgetting it.
Tom: Right, and the catch is, in most research papers, they test this by giving the model nice, equal-sized batches of new classes. Like, ten classes, then ten more, then ten more. It's clean, it's tidy, but it's not how reality works.
Jane: Exactly. And that's the whole point of this paper. The authors from Nanjing University — Zhiming Xu, Baile Xu, Jian Zhao, Furao Shen, and Suorong Yang — they're saying, look, in the real world, sometimes you get one new class, sometimes you get twenty-five. And that irregular, "free-flow" arrival of new classes breaks a lot of existing methods.
Tom: And they prove it, too. They ran a bunch of standard, well-known algorithms under this new free-flow setup, and the accuracy just tanks. We're talking about drops of several points, sometimes more than ten points, compared to the same methods under the old, equal-sized setup.
Jane: Yeah, and that's a big deal, Tom. Because it means we've been building and testing these systems under conditions that are almost too easy. It's like training a driver only on a straight, empty highway and then being surprised when they struggle in city traffic.
Tom: That's a great analogy, Jane. So the paper isn't just pointing out a problem, though. They're proposing a fix, which is what we're going to dig into next. But first, I want our listeners to sit with that idea — the problem is real, and it's messy.
Jane: And the fix is clever. It's all about making the learning process more stable when the number of new classes is bouncing around all over the place. Stick around, because we're about to get into the nitty-gritty of how they actually pulled that off.
Summary: Tom: So we've established that "Towards Realistic Class-Incremental Learning with Free-Flow Increments" is tackling a real problem. Jane, can you walk us through what the paper actually found when they started stress-testing these models?
Jane: Sure, Tom. So they took seven different, popular class-incremental learning methods — you've got your replay methods, your distillation methods, your dynamic expansion methods — and they ran them on standard benchmarks like CIFAR-one hundred and ImageNet. But instead of the usual equal-sized tasks, they shuffled the class counts per step. Sometimes a step would have two new classes, sometimes twenty.
Tom: And the results were pretty stark, right?
Jane: Oh, uniformly bad. Every single method lost accuracy compared to the equal-sized setup. But the interesting part was *how* they failed. They showed confusion matrices in the paper, and you could see the models developing this weird bias. They'd start over-predicting the early classes and ignoring the newest ones.
Tom: Yeah, I saw that. It's like the model gets so destabilized by the irregular flow that it just clings to whatever it learned first. The paper calls this an "exposure imbalance" — some classes get seen a lot, others barely get seen, and the gradient updates go haywire.
Jane: Right. And the deeper issue is that the standard loss functions, like cross-entropy, are computed as an average over all the samples in a mini-batch. So if you have a step with only two new classes, those two classes dominate the gradient. But if the next step has twenty classes, each one gets a much smaller slice of the learning signal.
Tom: So the model's learning is being driven by the *schedule* of arrivals, not by the actual content of the data. That's a fundamental flaw.
Jane: Exactly. And that's why they came up with their main fix, which they call the Class-Wise Mean objective. Instead of averaging the loss over all samples, they average the loss within each class first, and then average those class averages equally. That way, whether a step brings two classes or twenty, each class gets an equal vote in the update.
Tom: That's a really elegant idea, Jane. It's like making sure every student in the class gets the same amount of time to speak, regardless of how many students show up that day.
Jane: Precisely. And that single change already gave a nice boost to all the baselines they tested. But they didn't stop there. They also added some method-specific tweaks, which is what we're going to get into now.
Tom: I love it. So we've got the core fix, and now we're going to see how they tailored it to different families of algorithms. Don't go anywhere.
Improvements: Jane: Welcome back. We're still talking about "Towards Realistic Class-Incremental Learning with Free-Flow Increments." Tom, we covered the main loss function fix. But the paper goes further, right?
Tom: It does. And this is where it gets really interesting, because they don't just apply one universal patch. They look at different types of methods and address their specific weaknesses under free-flow conditions. Let me bring in our senior researcher, Lu, to help unpack this.
Lu: Thanks, Tom. So, for methods that use knowledge distillation — that's where an old model teaches a new model — they found that the teaching signal gets corrupted by the new classes. So they propose a simple but powerful change: only apply the distillation loss to the replayed old samples, not the new ones. It's called replay-only distillation.
Jane: So the old model only teaches on data it actually knows about, rather than trying to give advice on brand-new classes it has no clue about. That makes a lot of sense.
Lu: Exactly. And for methods that use contrastive learning or other auxiliary losses, they found that the scale of those losses changes depending on how many classes are in the current step. So they normalize those losses to keep the gradient magnitudes stable.
Tom: And then there's the weight alignment fix, which is for a specific family of methods. Meng, you're our engineer — can you explain what weight alignment is and why it breaks?
Meng: Sure. So after training a step, some methods, like the WA baseline, do a post-hoc calibration. They look at the average magnitude of the classifier weights for old classes versus new classes, and they scale the new ones to match. The idea is to correct for the bias where new classes have larger weight norms.
Jane: And that works fine when you have a decent number of new classes, but under free-flow?
Meng: Right. If you only have one or two new classes, the average weight norm for those classes is a really noisy estimate. So the calibration over-corrects, and you end up with a worse model. The paper's fix, DIWA, is a dynamic version that weakens the alignment when the class increment is small and strengthens it when the increment is large.
Tom: So it's like an adaptive volume knob for the calibration. And the results show that stacking all these fixes on top of the class-wise mean loss gives a big boost across the board.
Lu: It does. On CIFAR-one hundred for example, they show that a method like BiC, which drops from forty-four point six nine percent accuracy in the equal-split setting to seventeen point one five percent under free-flow, jumps back up to forty-four point two five percent with their framework. That's a massive recovery.
Jane: And they even tested it on an extreme case where the model first learns ninety classes and then gets one or two at a time. TagFex, a strong method, nearly collapsed to one percent accuracy, but with their framework it stayed stable.
Meng: And the best part is, all these additions are cheap. They measured the training time and it's basically a wash. No extra computational burden for a huge gain in robustness.
Tom: That's the kind of engineering we like to hear. So we've got a problem, a fix, and a validation. Let's wrap this up in our final segment.
Conclusion: Tom: Alright, we're at the finish line for our discussion on "Towards Realistic Class-Incremental Learning with Free-Flow Increments." Jane, what's the big takeaway for our listeners?
Jane: The big takeaway is that the standard way we test class-incremental learning is too clean. The paper shows that when you let the number of new classes vary freely, which is what happens in the real world, most existing methods break down. But they also show that this isn't a hopeless problem — with a few smart adjustments, you can make those methods much more resilient.
Tom: And I love that the fixes are model-agnostic. They work with replay methods, distillation methods, and expansion methods. It's not about inventing a brand new architecture; it's about fixing the fundamentals of how we compute the loss and calibrate the model.
Lu: I think the impact here goes beyond just benchmarks. This is a step toward making continual learning actually deployable. Think about a robot in a warehouse that needs to learn about a new product one week, and then twenty new products the next week. This framework helps it handle that irregular flow gracefully.
Meng: From an engineering standpoint, the fact that it's cheap is huge. You can retrofit these changes onto existing systems without a big rewrite. That lowers the barrier for adoption.
Lalam: And culturally, I see this as a move toward more human-like learning. We don't learn in perfectly balanced batches. We learn in bursts and trickles. Making eye that can handle that irregularity brings it closer to how we actually acquire knowledge, which could make eye assistants more adaptable in dynamic, real-world settings.
Jane: That's a beautiful way to put it, Lalam. So, we've said goodbye to this paper, but the ideas are definitely going to stick with us. Thanks to everyone for tuning in, and we'll be back soon with another exciting paper from the arXiv.
Tom: See you next time, folks!
Zhiming Xu, Baile Xu, Jian Zhao, Furao Shen, Suorong Yang
Nanjing University
cs.LG
Submitted: 2026-08-17
Updated: 2026-08-18
Comments: 15pages, 5figures, 3 tables
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 78/100
Key concepts
- Class-Incremental Learning (CIL)
- This is a machine learning process where a model learns new classes over time. Unlike standard training where all pictures are given at once, CIL requires the model to learn new things while successfully remembering previously learned concepts without forgetting them.
- Free-Flow Increments
- This refers to the real-world scenario where new classes arrive irregularly. Instead of receiving equal batches of ten classes repeatedly, the model might receive one class or twenty-five, making the learning process unstable and challenging for existing methods.
- Class-Wise Mean Objective
- This is a core fix proposed by the authors. Instead of averaging loss across all samples in a batch (which is skewed by irregular arrivals), this method averages the loss within each class first, then averaging those class averages equally, ensuring every new class gets an equal vote in the learning update.
Terminology
Summary
Summary
This paper introduces and formalizes a new, more realistic setting for Class-Incremental Learning (CIL) termed Free-Flow Class-Incremental Learning (FFCIL). The authors argue that standard CIL benchmarks are evaluated under idealized conditions characterized by balanced task partitions and data distribution,
where the model receives new categories in equal or near-equal portions across steps. However, they note that this design facilitates clean and comparable evaluation, but it imposes an artificial regularity that practical data streams rarely satisfy.
In real applications, whether a learning step introduces a single novel class or a massive influx of diverse categories, the model must integrate them dynamically while correctly differentiating all observed classes.
The FFCIL problem is formally defined by two conditions: (1) Free-flow: Each step t introduces a non-empty new class set C t with highly variable size, where C t 1 and C t-C t-1 is unbounded; (2) Non-repetition: Previously observed classes do not reappear, i.e., C t C s = for all t not equal to s. The authors note that no restriction is imposed between consecutive steps, enabling highly unbalanced updates (e.g., learning a single class on D t and tens of classes on D t+1).
The paper identifies that such irregular, highly variable increments introduce severe exposure imbalances and classifier bias, destabilizing optimization and dramatically amplifying catastrophic forgetting, leading to substantial performance degradation.
The authors analyze how free-flow increments perturb standard optimization dynamics, observing that "varying incoming class sizes induce erratic gradient magnitudes in instance aggregated losses and step-dependent variations in auxiliary objectives (e.g., knowledge distillation). Furthermore, the extreme heterogeneity across updates undermines the reliability of step-wise statistics, causing post-hoc weight alignment to become overly aggressive and skewed."
To address these challenges, the authors propose a model-agnostic framework with several key components:
-
Class-Wise Mean (CWM) Objective: The paper shows that standard cross-entropy loss, computed as an instance-level empirical risk, is equivalent to a weighted sum of class-conditional mean losses where the weight is the within-batch class frequency pi c = n c/B. In FFCIL, this frequency becomes
highly step-dependent, amplifying per-class contributions when few classes arrive and diluting them when many classes arrive.
The CWM objective replaces this with a uniform aggregation: L cwm CE = 1 over C batch sum c in C batch (1 over n c sum i in b c CE(p theta(x i), y i)). This ensureseach present class contributes equally regardless of its sample count.
-
Replay-Only Distillation: The paper demonstrates that vanilla knowledge distillation (KD) decomposes into old-class and new-class gradient terms, with weights proportional to their batch fractions. In FFCIL,
the fraction of new-class samples B new/B changes markedly with the number of arriving classes, so the relative contribution of L new fluctuates across steps and makes distillation gradients inconsistent.
The proposed solution applies distillationexclusively on replayed old-class samples to avoid interference from the unstable L new term,
formulated as L KD ro = -B old over B times 1 over C old sum c in C old 1 over n c sum i in I c i. -
Loss Scale Normalization: For methods with contrastive learning and knowledge transfer objectives (e.g., TagFex), the paper normalizes the InfoNCE loss by (N eff) where N eff is the effective number of valid negatives per anchor, and normalizes the knowledge transfer loss L kl by C t, since
its scale is sensitive to C t.
-
Dynamic Intervention Weight Alignment (DIWA): The paper notes that conventional Weight Alignment (WA) applies a uniform alignment strength regardless of increment size. However, "Small increments provide unreliable estimates of new-class weight statistics, making full alignment prone to over-calibration, whereas larger increments yield more stable statistics and thus benefit from stronger alignment.
DIWA introduces an intervention factor eta t = 1-(1-eta min) (-C t-1 over tau), which
increases the alignment strength as C t grows and weakens it when fewer classes are introduced." The final scaling factor is gamma t = (1- eta t) + eta t mu old over mu new.
The paper also applies CWM to auxiliary losses in dynamic-expansion methods like DER and MEMO, replacing the standard cross-entropy on auxiliary classifiers with a CWM-based form.
Experiments are conducted on CIFAR-100, VTAB, and ImageNet datasets, evaluating seven baselines: Replay, iCaRL, WA, BiC, DER, MEMO, and TagFex. The results show that CIL methods across different paradigms all suffer an accuracy drop under the FFCIL setting.
For example, on CIFAR-100, BiC drops from 44.69% (equal tasks) to 17.15% (FFCIL), and WA drops from 51.83% to 14.98%. The proposed framework consistently improves performance, e.g., BiC improves to 44.25% and WA to 49.43% under FFCIL with the framework. Confusion matrix analysis for BiC shows that under FFCIL, the predicted label distribution is clearly skewed toward earlier classes, while the most recently learned classes are markedly under-predicted,
whereas the proposed method improves the accuracy for most classes and reduces the prediction bias.
The paper also studies the impact of different step-size schedules (ascending, descending, fluctuating), finding that different step schedules have a substantial impact on the final accuracy,
with descending schedules causing clear performance drops and fluctuating schedules also degrading performance. The proposed method consistently improves performance across all schedules.
Under an extreme FFCIL schedule (90 classes initially, then 1-2 classes per step), both DER and TagFex suffer substantial drops, with TagFex exhibiting a near-collapse behavior, with accuracy degrading to around 1%.
The proposed method effectively mitigates this issue and maintains stable performance.
Ablation studies show that the CWM loss yields consistent accuracy improvements across all baselines, and adding other components further improves performance. Training-time analysis shows that CWM, DIWA, and replay-only distillation do not increase the training time; instead, they even lead to a slight reduction in runtime, while scale normalization introduces only a negligible time increase.
The paper concludes that FFCIL induces consistent performance degradation under standard training objectives, while the proposed strategies substantially improve robustness and accuracy,
and suggests future work may explore model architectures for FFCIL and specific FFCIL algorithms.
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems and what the improved system can do:
-
Current limitation: Standard cross-entropy loss weights each class by its sample frequency in the mini-batch, causing unstable gradients when class increments vary in size.
-
Improvement: Implement the CWM objective that computes loss per class, then averages uniformly across classes present in the batch. This removes the implicit frequency-based weighting.
-
What the improved system does: Maintains stable gradient magnitudes and update directions regardless of whether a step introduces 1 class or 25 classes, preventing optimization instability and reducing catastrophic forgetting.
-
Current limitation: Vanilla distillation (Eq. 7) includes new-class samples, whose contribution fluctuates wildly with class increment size, destabilizing the distillation gradient.
-
Improvement: Apply distillation loss exclusively on replayed old-class samples (Eq. 9), with a calibration factor (B old/B) to preserve overall distillation strength.
-
What the improved system does: Produces consistent distillation gradients across steps, preserving old-class knowledge reliably even when new-class influx varies dramatically.
-
Current limitation: Contrastive loss (InfoNCE) and knowledge transfer losses scale with the number of valid negatives (N eff) and new-class subspace dimension (C t), respectively, making them unreliable under free-flow increments.
-
Improvement: Normalize contrastive loss by log(N eff) (Eq. 14) and knowledge transfer loss by C t (Eq. 15).
-
What the improved system does: Maintains consistent loss magnitudes for auxiliary objectives regardless of step composition, preventing over- or under-weighting of these terms during training.
-
Current limitation: Standard Weight Alignment (WA) applies uniform calibration strength after every step, but small class increments produce unreliable statistics, leading to over-calibration.
-
Improvement: Introduce an intervention factor η t (Eq. 18) that modulates alignment strength based on the number of new classes C t, using a temperature parameter τ to control saturation.
-
What the improved system does: Applies strong calibration when many new classes arrive (reliable statistics) and weak calibration when few classes arrive (unreliable statistics), preventing over-correction and maintaining classifier balance.
-
Current limitation: Methods like DER and MEMO use instance-averaged cross-entropy on auxiliary classifiers, inheriting the same frequency-bias problem.
-
Improvement: Replace auxiliary loss with CWM form (Eq. 12) over step-relative labels.
-
What the improved system does: Stabilizes auxiliary classifier training under free-flow increments, ensuring balanced learning across the
other
category and newly introduced classes.
-
Handles arbitrary class arrival patterns: Learns effectively whether a step introduces 1 class, 5 classes, or 25 classes, without requiring predefined equal-size tasks.
-
Maintains accuracy under extreme schedules: Remains stable even when 90 classes are learned first, followed by 1–2 classes per step (where baseline methods collapse to 1% accuracy).
-
Reduces prediction bias: Produces balanced predictions across all learned classes, avoiding the recency bias (favoring recent classes) or the reverse bias (under-predicting recent classes) seen in standard methods under free-flow conditions.
-
Zero additional training overhead: All improvements are computationally free—CWM and DIWA actually reduce runtime slightly, while normalization adds negligible cost.
-
Model-agnostic compatibility: Works across replay-based, distillation-based, and dynamic-expansion methods (validated on Replay, iCaRL, BiC, WA, DER, MEMO, TagFex).
-
Consistent gains across datasets: Improves final accuracy by 1.5–14.5 percentage points on CIFAR-100, VTAB, and ImageNet under free-flow schedules, while reducing forgetting by 0.3–6.1 points.
Sources
- Latest Advancements Towards Catastrophic Forgetting under Data Scarcity: A Comprehensive Survey on Few-Shot Class Incremental Learning
- A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark
- A Model or 603 Exemplars: Towards Memory-Efficient Class-Incremental Learning
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks