Continual Distillation Learning for Rehearsal-Free Class-Incremental Learning via Decoupled Prompting

arXiv:2407.13911 · cs.CV, cs.LG · Submitted 2024-07-18 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Continual Distillation Learning for Rehearsal-Free Class-Incremental Learning via Decoupled Prompting".

Jane: The paper was written by the authors from University of Texas at Dallas.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the arXiv radio hour, everyone. I’m Tom, and I’ve got my co-host Jane here with me. We’re looking at a paper that’s got a mouthful of a title today: “Continual Distillation Learning: Knowledge Distillation in Prompt-based Continual Learning.” Jane, let’s be real — that title is a lot. Can you break it down for our listeners?

Jane: Absolutely, Tom. So there are two big ideas jammed into that title. The first is continual learning, which is when a model learns a bunch of tasks one after another, like learning to recognize dogs, then cats, then birds, without forgetting the earlier ones. The second is knowledge distillation, which is when you take a big, smart model and use it to teach a smaller, faster model. This paper is about putting those two together in a really specific way.

Tom: Right, and I love this because the setup is so practical. You’ve got these big vision models — like a giant ViT-Large — that perform great on continual learning tasks, but they’re slow and expensive to run. So the dream is to have a small student model learn from that big teacher, so you get the accuracy without the heavy compute.

Jane: And the authors — Qifan Zhang, Yunhui Guo, and Yu Xiang from UT Dallas — they noticed something interesting. When you just take existing distillation methods and drop them into this continual learning setting, they don’t work well. The student barely improves. So they had to invent a new way to make the transfer actually stick.

Tom: That’s the hook for me. It’s not just “here’s a new trick.” It’s “here’s a problem nobody saw coming, and here’s why the old tricks fail.” I mean, imagine trying to teach a student by having them copy your notes, but every time they start a new chapter, you throw away the old notebook. That’s basically what’s happening.

Jane: Exactly. And that’s what we’re going to dig into — why the old methods break, and what they built instead. Stick around, because the solution involves something called “KD prompts,” and it’s genuinely clever.

Tom: Before we jump into the technical weeds, I want to give a quick shout-out to the authors for framing this so clearly. They even have a project website. It’s rare to see a paper that’s this well-organized around a new problem statement.

Jane: It really is. And the fact that they tested it across three different continual learning methods — L2P, DualPrompt, and CODA-Prompt — means the solution isn’t just a one-off. It generalizes.

Tom: Alright, so we’ve got the title and the big picture. Next up, we’re going to talk about what they actually found when they tried the standard distillation approaches. That’s where the surprises start.

Summary: Tom: So we’re back, and we’re still on “Continual Distillation Learning: Knowledge Distillation in Prompt-based Continual Learning.” Jane, last segment we set the stage. Now let’s talk about what the paper actually found when they ran the experiments.

Jane: Right. So they set up a teacher-student scenario. The teacher is a big model, like ViT-Large, and the student is smaller, like ViT-Base or ViT-Small. Both are using prompt-based continual learning methods — these are models that keep the backbone frozen and only train little “prompt” vectors that get inserted into the network.

Tom: And the results were pretty disappointing for the standard methods. They tried logit distillation — that’s the classic Hinton approach where the student learns to match the teacher’s output probabilities. They also tried feature distillation, where the student tries to match the internal representations. And they tried some newer variants like DKD and ReviewKD.

Jane: And none of them gave a big boost. On ImageNet-R, for example, the student with no distillation got about sixty-seven point four percent accuracy. Adding the best standard method — DeiT’s distillation token — got it to about seventy point seven percent. So there’s some improvement, but it’s modest, and the forgetting rate stays high.

Tom: Forgetting rate is key here. That’s how much the model loses on old tasks after learning new ones. The standard methods weren’t just underperforming — they were still forgetting a lot. And the authors figured out why. They call it “distillation information forgetting.”

Jane: Let me explain that. In prompt-based continual learning, each task picks its own set of prompts from a shared pool. So task one might use prompts A, B, and C. Task two might use prompts D, E, and F. The problem is, when the student learns from the teacher on task one, the knowledge gets baked into prompts A, B, and C. But when task two comes along, those prompts aren’t used anymore. So the distilled knowledge is just… gone.

Tom: It’s like writing notes in a notebook, then throwing the notebook away and starting a fresh one for each chapter. The student never accumulates the teacher’s wisdom across tasks.

Jane: Exactly. And that’s the core insight of the paper. They didn’t just say “distillation doesn’t work.” They diagnosed why it doesn’t work in this specific setting.

Tom: And that diagnosis leads to their solution, which we’ll get into next. But I want to pause on the numbers because they’re actually pretty striking. When they used their own method — which they call KDP — the student with ViT-Small got to about seventy-one point nine percent on ImageNet-R. That’s better than any of the standard distillation methods, and it’s getting close to the teacher’s seventy-six point four percent.

Jane: And on CIFAR-one hundred the same pattern. KDP pushed the small student from eighty-two point two percent up to eighty-four point three percent, while the best standard method only got to eighty-three point eight percent. So it’s a consistent win.

Tom: So the summary is: standard distillation doesn’t transfer well in continual learning because of the prompt selection mechanism, and the authors proved it with experiments. Now we need to talk about how they fixed it.

Jane: And that fix is the heart of the paper. Let’s take a quick break, and when we come back, we’ll talk about KD prompts and why they’re such a clever workaround.

Improvements: Tom: We’re back on “Continual Distillation Learning: Knowledge Distillation in Prompt-based Continual Learning.” Jane, we left off with the problem — distillation information forgetting. Now let’s talk about the fix. What did they actually build?

Jane: So the authors introduced something called KD prompts. These are additional trainable prompts that are inserted into the student’s frozen ViT backbone, but they’re completely separate from the main prompt pool used for continual learning.

Tom: And the key difference is that these KD prompts are global — they’re shared across all tasks. They don’t get re-selected for each new task. So when the student learns from the teacher on task one, that knowledge accumulates in the KD prompts, and it’s still there for task two, task three, and so on.

Jane: That’s the elegant part. The original prompts are task-specific because that’s what helps with continual learning — each task gets its own space. But the KD prompts are task-agnostic. They’re like a permanent notebook that the student keeps adding to, rather than throwing away.

Tom: And they didn’t stop there. They also added a distillation token and a separate classifier, kind of like what DeiT does. So the student has its main classifier for the actual task, and a KD classifier that specifically learns from the teacher’s logits.

Jane: Right. And the loss function combines both — the student still learns the true labels for the current task, but it also gets a KL divergence loss that pushes its KD classifier to match the teacher’s output. They balance those with a parameter alpha, which they set to zero point five.

Tom: Let’s talk about the results because they’re impressive. On ImageNet-R with CODA-Prompt, the ViT-Small student jumped from sixty-seven point four percent to seventy-one point nine percent with KDP. That’s a four point five point gain. And the forgetting rate dropped from eight point five percent down to five point six percent.

Jane: And the same pattern held across all three continual learning methods they tested — L2P, DualPrompt, and CODA-Prompt. So this isn’t a hack that only works with one setup. It’s a general improvement.

Tom: They also ran ablations to make sure they weren’t just adding parameters for no reason. They compared adding the same number of trainable prompts into the main prompt pool versus using them as KD prompts. The KD prompts won by a clear margin.

Jane: That’s a really important control. It shows the improvement isn’t just from having more trainable parameters. It’s specifically from having prompts that are designed to be global and cross-task.

Tom: And they also tried the obvious alternative — unfreezing part of the ViT backbone to let it learn. That failed badly. The accuracy dropped to about thirty percent and the forgetting rate shot up to fifty-eight percent. So freezing the backbone and using prompts is definitely the right call.

Jane: So the improvement is real, it’s consistent, and it’s principled. The KD prompts solve the exact problem they identified.

Tom: Now I want to bring in some other voices. Lu, you’re the researcher — what do you think about this approach? Is there a bigger implication here?

Lu: I think the big implication is that this opens the door for continual learning on edge devices. You can train a large teacher once, then continually distill into a small student that can actually run on a phone or a robot. The student keeps learning new tasks without forgetting, and it stays fast.

Tom: That’s a great point. And Meng, from the engineering side — does this actually hold up in practice?

Meng: The main cost is that you have to train the teacher alongside the student, which doubles the training time. But for deployment, the student is what matters. And the paper shows the student gets close to the teacher’s accuracy while being much cheaper to run. That’s a trade-off most teams would take.

Tom: Alright, we’ve got the problem, the solution, and the results. Let’s wrap this up in the next segment.

Conclusion: Tom: And we’re at the finish line for “Continual Distillation Learning: Knowledge Distillation in Prompt-based Continual Learning.” Jane, give us the final summary.

Jane: Sure. The paper identifies a real problem — standard knowledge distillation doesn’t work well in prompt-based continual learning because the prompt selection mechanism causes distilled knowledge to be lost between tasks. They call this distillation information forgetting.

Tom: And their fix is KD prompts — global, task-agnostic prompts that sit outside the main prompt pool and accumulate knowledge across all tasks. Combined with a distillation token and classifier, this gives consistent improvements across three different continual learning methods and two datasets.

Jane: The numbers speak for themselves. On ImageNet-R, the small student went from sixty-seven point four percent to seventy-one point nine percent with KDP. On CIFAR-one hundred from eighty-two point two percent to eighty-four point three percent. And the forgetting rate dropped significantly in both cases.

Tom: And the ablation studies show it’s not just about adding parameters — it’s about where you put them. The KD prompts are specifically designed to be cross-task, and that makes all the difference.

Lu: I’d add that this has real-world implications for deploying continual learning on resource-constrained devices. You get the accuracy of a big model with the speed of a small one, and the student can keep learning new tasks over time.

Meng: And from an engineering standpoint, the trade-off is clear. You spend more time training, but you get a deployable model that’s much cheaper to run. That’s a win for most real-world applications.

Lalam: And I think there’s a cultural angle too. This kind of efficient continual learning could make adaptive AI systems more accessible — think assistive technologies that learn new user preferences over time without needing a huge server farm behind them.

Tom: That’s a beautiful way to end it. So, to the authors — Qifan Zhang, Yunhui Guo, and Yu Xiang — great work. You’ve given us a new problem, a clear diagnosis, and a clever solution.

Jane: And to our listeners, thanks for tuning in. We’ll be back with another paper soon. Until then, keep learning — and don’t forget what you learned yesterday.

Tom: See you next time, everyone.

University of Texas at Dallas

cs.CV, cs.LG

Submitted: 2024-07-18

Updated: 2026-09-05

Comments: Accepted at CoLLAs 2026

Project page: https://irvlutd.github.io/CDL

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 44/100

The gist: L2P, DualPrompt, and CODA-Prompt.

Key concepts

Continual Learning
This is when a model learns multiple tasks sequentially, such as recognizing dogs, cats, and birds, without forgetting the knowledge gained from earlier tasks.
Knowledge Distillation
This technique involves using a large, smart model (teacher) to teach a smaller, faster model (student) by having the student mimic the teacher's outputs or representations.
Distillation Information Forgetting
This is the problem where standard distillation methods fail in continual learning. Because prompts are task-specific, knowledge learned from one task is lost when moving to a new task because those specific prompts are no longer used.
KD Prompts
These are additional trainable prompts that are separate from the main prompt pool and remain global across all tasks. They accumulate teacher knowledge over time, solving the forgetting problem.

Terminology

Summary

Summary

The paper introduces a new research problem termed Continual Distillation Learning (CDL), which aims to use knowledge distillation (KD) to improve prompt-based continual learning (CL) models. The authors motivate this problem by observing that for prompt-based CL methods using ViTs, larger ViTs achieve better performance, and thus, it is appealing to use large models in prompt-based CL. However, larger models introduce more computation during inference. The CDL problem is defined as a teacher-student setup where, at each task in a continual learning sequence, a teacher model (with a larger ViT backbone) is first updated, and then a student model (with a smaller ViT backbone) is updated using the data of the new task and the teacher model. The authors explicitly limit the scope to prompt-based CL models because they experimentally verified that large models may not result in better performance for the traditional CNN-based CL models such as iCaRL or LWF, as these models update backbones and deeper and larger CNNs tend to overfit new tasks.

The paper's central finding is that existing KD methods are ineffective in the CDL setup. The authors state, When applying existing knowledge distillation methods, i.e., logit distillation [13, 40], feature distillation [28, 4], to the CDL problem, we found that the results are not satisfactory. They identify the root cause as a phenomenon they name distillation information forgetting. In prompt-based CL, each task independently reselects prompt components from a prompt pool via a key-query mechanism. Since the ViT backbone is frozen, the knowledge embedded within them from the teacher model cannot be transferred to the next task. Specifically, if task A uses prompt subset PA and task B uses a disjoint subset PB, then the prompts in PB has no information about the teacher model, forcing the student to learn from scratch for task B. The authors note that a naive solution of unfreezing part of the ViT backbone fails, as experiments show it leads to severe forgetting (average accuracy dropping to 29.98% and forgetting rate reaching 58.21%).

To address this, the authors propose a novel method named Knowledge Distillation based on Prompts (KDP). KDP introduces globally accessible prompts, specifically designed for knowledge distillation, into the frozen ViT backbone of the student model, which they call KD prompts. These prompts are independent of the key-query mechanism of the prompt pool, are not limited to specific tasks, and can facilitate cross-task distillation. The KD prompts are inserted into the ViT backbone using prefix-tuning, and a separate KD classifier (with a corresponding distillation token) is added to the student model to process the teacher's soft labels via logit distillation. The final training loss for the student model combines a student classification loss, a student knowledge distillation loss (KL divergence), and the student prompt pool loss, balanced by hyperparameters α and λ.

The authors conduct experiments on CIFAR-100 and ImageNet-R datasets in a class-incremental setting with 10 tasks, using three prompt-based CL methods: L2P, DualPrompt, and CODA-Prompt. They test two teacher-student pairs: ViT-Large to ViT-Base, and ViT-Base to ViT-Small. The results show that KDP consistently outperforms existing KD methods (KD, DKD, FitNets, ReviewKD, DeiT) across all settings. For example, with CODA-Prompt on ImageNet-R, KDP achieves 71.92% average accuracy (vs. 67.44% for the student without distillation) for ViT-Base to ViT-Small, and 78.62% (vs. 76.42% for the student without distillation) for ViT-Large to ViT-Base. On CIFAR-100, KDP achieves 84.31% and 87.13% for the respective pairs. The authors also demonstrate that KDP generalizes across all three prompt-based CL methods, with the combination of CODA-Prompt and KDP achieving state-of-the-art performance.

Ablation studies reveal that: (1) the length of KD prompts (Lp) has limited influence on performance, while increasing the number of layers where KD prompts are inserted (n) improves accuracy, so they set Lp=6 and n=12; (2) the best performance is achieved when both KD prompts and the KD classifier are used together; (3) unfreezing the ViT backbone leads to catastrophic forgetting; and (4) adding the same number of trainable prompts directly into the CL prompt pool is inferior to KDP, confirming that the improvement is due to the global, cross-task nature of KD prompts rather than merely increasing trainable parameters.

The paper concludes that KDP effectively enhances the distillation performance in comparison to existing KD methods in the CDL setup. The authors acknowledge a limitation: CDL models require training a teacher model and a student model jointly. The total training time and memory consumption are increased.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in an AI system, along with what the improved system can do:


  1. Add a Distillation Information Forgetting prevention mechanism
  • Implementation: Insert globally shared, task-agnostic KD prompts into the frozen ViT backbone of the student model, independent of the key-query prompt selection mechanism.

  • Benefit: Prevents loss of teacher knowledge when the model switches to a new task (which would otherwise re-select different CL prompts and discard distilled information).

  1. Add a dedicated distillation token and classifier
  • Implementation: Insert a learnable distillation token at the first ViT layer, connected to a separate KD classifier. Use KL divergence with temperature (τ=2) to align student logits with teacher logits.

  • Benefit: Separates continual learning (CL) loss from distillation loss, reducing interference and improving overall accuracy and forgetting rate.

  1. Modify the prefix-tuning mechanism to accept both CL and KD prompts
  • Implementation: Extend the prefix-tuning function to concatenate both CL prompts and KD prompts into the key/value of each MSA layer.

  • Benefit: Enables simultaneous task-specific learning and global knowledge transfer without unfreezing the backbone (which causes severe forgetting).

  1. Add a weighted loss function combining CL, KD, and prompt-pool losses
  • Implementation: Use LS = (1−α)·L classification + α·L KD + λ·L pool, with α=0.5 and λ=1.

  • Benefit: Balances task learning and teacher imitation, improving accuracy (e.g., +4.5% on ImageNet-R for ViT-Small) while reducing forgetting rate.

  1. Insert KD prompts across all 12 ViT blocks (not just the first few)
  • Implementation: Place KD prompts (length 6) in every block from layer 1 to 12.

  • Benefit: Maximizes cross-layer knowledge transfer, leading to higher accuracy (verified in ablation studies).

  • Continual learning without catastrophic forgetting:

The system can learn new tasks sequentially while retaining high accuracy on previously seen classes, even when the backbone is frozen.

  • Efficient deployment of large-model knowledge:

A small ViT-Small student can achieve performance close to a ViT-Base teacher (e.g., 71.92% vs. 76.42% on ImageNet-R, and 84.31% vs. 86.16% on CIFAR-100), with significantly lower inference cost.

  • Generalize across multiple prompt-based CL methods:

The improvements work with L2P, DualPrompt, and CODA-Prompt, making the system adaptable to different continual learning frameworks.

  • Reduce forgetting rate significantly:

For example, on ImageNet-R with CODA-Prompt, the forgetting rate drops from 8.52% (no distillation) to 5.61% with KDP, and from 6.52% to 2.08% with L2P.

  • Avoid the pitfalls of unfreezing the backbone:

Unlike naive approaches that unfreeze ViT layers (which cause accuracy to collapse to 30% and forgetting to 58%), the system maintains stable performance by keeping the backbone frozen and using KD prompts.

  • Support real-time, resource-constrained deployment:

The student model is compact and fast, suitable for mobile devices, robots, or edge servers, while still benefiting from the knowledge of a large teacher model.

These improvements directly address the core limitations identified in the paper and enable a practical, efficient, and robust continual learning system that can be deployed in real-world scenarios.

Abstract

Prompt-based continual learning has shown strong performance in rehearsal-free class-incremental learning by adapting learnable prompts while freezing a pre-trained Vision Transformer (ViT) backbone. However, the effect of backbone scale remains underexplored. We observe that larger ViT backbones consistently yield better continual learning performance, which motivates us to study how to transfer such capability from a larger model to a smaller one. In this paper, we introduce Continual Distillation Learning (CDL), a new setting for knowledge distillation in rehearsal-free prompt-based continual learning. We show that conventional distillation methods provide only limited gains in CDL, mainly because task-specific prompts are forced to encode both continual adaptation and distillation knowledge, while lacking a persistent mechanism for cross-task knowledge transfer. To address this problem, we propose Decoupled Continual Distillation Learning (D-CDL), which introduces persistent Knowledge-Distillation prompts (KD-prompts) and a dedicated KD branch to explicitly decouple distillation from task adaptation. The proposed KD-prompts are propagated across tasks as a global carrier of teacher knowledge, while the original prompts remain responsible for continual learning. D-CDL is simple, general, and can be integrated into various prompt-based continual learning frameworks. Extensive experiments on Split CIFAR-100 and Split ImageNet-R across four representative continual learning methods show that D-CDL consistently outperforms existing distillation baselines and substantially improves student performance under different teacher-student settings.

Sources

Related papers