Continual Distillation Learning for Rehearsal-Free Class-Incremental Learning via Decoupled Prompting
summary
The gist
L2P, DualPrompt, and CODA-Prompt.
In short
The episode discusses a paper on Continual Distillation Learning for Rehearsal-Free Class-Incremental Learning via Decoupled Prompting. The hosts analyze why standard distillation methods fail in this setting due to 'distillation information forgetting' and introduce 'KD prompts' as a solution. Experiments show KD prompts significantly improve accuracy and reduce forgetting rates across different continual learning methods.
Key concepts
- Continual Learning
- This is when a model learns multiple tasks sequentially, such as recognizing dogs, cats, and birds, without forgetting the knowledge gained from earlier tasks.
- Knowledge Distillation
- This technique involves using a large, smart model (teacher) to teach a smaller, faster model (student) by having the student mimic the teacher's outputs or representations.
- Distillation Information Forgetting
- This is the problem where standard distillation methods fail in continual learning. Because prompts are task-specific, knowledge learned from one task is lost when moving to a new task because those specific prompts are no longer used.
- KD Prompts
- These are additional trainable prompts that are separate from the main prompt pool and remain global across all tasks. They accumulate teacher knowledge over time, solving the forgetting problem.
Terminology used across episodes
This episode discusses
- Continual Distillation Learning for Rehearsal-Free Class-Incremental Learning via Decoupled Prompting · Paper Radio
- Distilling Knowledge via Knowledge Review
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- PathNet: Evolution Channels Gradient Descent in Super Neural Networks
- Lifelong Machine Learning with Deep Streaming Linear Discriminant Analysis
- Adam: A Method for Stochastic Optimization
- Learning without Forgetting
- Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing
- Gradient Episodic Memory for Continual Learning
- Toward Understanding Catastrophic Forgetting in Continual Learning
- FitNets: Hints for Thin Deep Nets
- Training data-efficient image transformers & distillation through attention
- Three scenarios for continual learning
- A Comprehensive Survey of Continual Learning: Theory, Method and Application
- ViTKD: Practical Guidelines for ViT feature knowledge distillation
- Lifelong Learning with Dynamically Expandable Networks
- Decoupled Knowledge Distillation
The paper
Continual Distillation Learning for Rehearsal-Free Class-Incremental Learning via Decoupled Prompting · Read on arXiv
University of Texas at Dallas
Prompt-based continual learning has shown strong performance in rehearsal-free class-incremental learning by adapting learnable prompts while freezing a pre-trained Vision Transformer (ViT) backbone. However, the effect of backbone scale remains underexplored. We observe that larger ViT backbones consistently yield better continual learning performance, which motivates us to study how to transfer such capability from a larger model to a smaller one. In this paper, we introduce Continual Distillation Learning (CDL), a new setting for knowledge distillation in rehearsal-free prompt-based continual learning. We show that conventional distillation methods provide only limited gains in CDL, mainly because task-specific prompts are forced to encode both continual adaptation and distillation knowledge, while lacking a persistent mechanism for cross-task knowledge transfer. To address this problem, we propose Decoupled Continual Distillation Learning (D-CDL), which introduces persistent Knowledge-Distillation prompts (KD-prompts) and a dedicated KD branch to explicitly decouple distillation from task adaptation. The proposed KD-prompts are propagated across tasks as a global carrier of teacher knowledge, while the original prompts remain responsible for continual learning. D-CDL is simple, general, and can be integrated into various prompt-based continual learning frameworks. Extensive experiments on Split CIFAR-100 and Split ImageNet-R across four representative continual learning methods show that D-CDL consistently outperforms existing distillation baselines and substantially improves student performance under different teacher-student settings.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Continual Distillation Learning for Rehearsal-Free Class-Incremental Learning via Decoupled Prompting".
Jane: The paper was written by the authors from University of Texas at Dallas.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the arXiv radio hour, everyone. I’m Tom, and I’ve got my co-host Jane here with me. We’re looking at a paper that’s got a mouthful of a title today: “Continual Distillation Learning: Knowledge Distillation in Prompt-based Continual Learning.” Jane, let’s be real — that title is a lot. Can you break it down for our listeners?
Jane: Absolutely, Tom. So there are two big ideas jammed into that title. The first is continual learning, which is when a model learns a bunch of tasks one after another, like learning to recognize dogs, then cats, then birds, without forgetting the earlier ones. The second is knowledge distillation, which is when you take a big, smart model and use it to teach a smaller, faster model. This paper is about putting those two together in a really specific way.
Tom: Right, and I love this because the setup is so practical. You’ve got these big vision models — like a giant ViT-Large — that perform great on continual learning tasks, but they’re slow and expensive to run. So the dream is to have a small student model learn from that big teacher, so you get the accuracy without the heavy compute.
Jane: And the authors — Qifan Zhang, Yunhui Guo, and Yu Xiang from UT Dallas — they noticed something interesting. When you just take existing distillation methods and drop them into this continual learning setting, they don’t work well. The student barely improves. So they had to invent a new way to make the transfer actually stick.
Tom: That’s the hook for me. It’s not just “here’s a new trick.” It’s “here’s a problem nobody saw coming, and here’s why the old tricks fail.” I mean, imagine trying to teach a student by having them copy your notes, but every time they start a new chapter, you throw away the old notebook. That’s basically what’s happening.
Jane: Exactly. And that’s what we’re going to dig into — why the old methods break, and what they built instead. Stick around, because the solution involves something called “KD prompts,” and it’s genuinely clever.
Tom: Before we jump into the technical weeds, I want to give a quick shout-out to the authors for framing this so clearly. They even have a project website. It’s rare to see a paper that’s this well-organized around a new problem statement.
Jane: It really is. And the fact that they tested it across three different continual learning methods — L2P, DualPrompt, and CODA-Prompt — means the solution isn’t just a one-off. It generalizes.
Tom: Alright, so we’ve got the title and the big picture. Next up, we’re going to talk about what they actually found when they tried the standard distillation approaches. That’s where the surprises start.
Summary: Tom: So we’re back, and we’re still on “Continual Distillation Learning: Knowledge Distillation in Prompt-based Continual Learning.” Jane, last segment we set the stage. Now let’s talk about what the paper actually found when they ran the experiments.
Jane: Right. So they set up a teacher-student scenario. The teacher is a big model, like ViT-Large, and the student is smaller, like ViT-Base or ViT-Small. Both are using prompt-based continual learning methods — these are models that keep the backbone frozen and only train little “prompt” vectors that get inserted into the network.
Tom: And the results were pretty disappointing for the standard methods. They tried logit distillation — that’s the classic Hinton approach where the student learns to match the teacher’s output probabilities. They also tried feature distillation, where the student tries to match the internal representations. And they tried some newer variants like DKD and ReviewKD.
Jane: And none of them gave a big boost. On ImageNet-R, for example, the student with no distillation got about sixty-seven point four percent accuracy. Adding the best standard method — DeiT’s distillation token — got it to about seventy point seven percent. So there’s some improvement, but it’s modest, and the forgetting rate stays high.
Tom: Forgetting rate is key here. That’s how much the model loses on old tasks after learning new ones. The standard methods weren’t just underperforming — they were still forgetting a lot. And the authors figured out why. They call it “distillation information forgetting.”
Jane: Let me explain that. In prompt-based continual learning, each task picks its own set of prompts from a shared pool. So task one might use prompts A, B, and C. Task two might use prompts D, E, and F. The problem is, when the student learns from the teacher on task one, the knowledge gets baked into prompts A, B, and C. But when task two comes along, those prompts aren’t used anymore. So the distilled knowledge is just… gone.
Tom: It’s like writing notes in a notebook, then throwing the notebook away and starting a fresh one for each chapter. The student never accumulates the teacher’s wisdom across tasks.
Jane: Exactly. And that’s the core insight of the paper. They didn’t just say “distillation doesn’t work.” They diagnosed why it doesn’t work in this specific setting.
Tom: And that diagnosis leads to their solution, which we’ll get into next. But I want to pause on the numbers because they’re actually pretty striking. When they used their own method — which they call KDP — the student with ViT-Small got to about seventy-one point nine percent on ImageNet-R. That’s better than any of the standard distillation methods, and it’s getting close to the teacher’s seventy-six point four percent.
Jane: And on CIFAR-one hundred the same pattern. KDP pushed the small student from eighty-two point two percent up to eighty-four point three percent, while the best standard method only got to eighty-three point eight percent. So it’s a consistent win.
Tom: So the summary is: standard distillation doesn’t transfer well in continual learning because of the prompt selection mechanism, and the authors proved it with experiments. Now we need to talk about how they fixed it.
Jane: And that fix is the heart of the paper. Let’s take a quick break, and when we come back, we’ll talk about KD prompts and why they’re such a clever workaround.
Improvements: Tom: We’re back on “Continual Distillation Learning: Knowledge Distillation in Prompt-based Continual Learning.” Jane, we left off with the problem — distillation information forgetting. Now let’s talk about the fix. What did they actually build?
Jane: So the authors introduced something called KD prompts. These are additional trainable prompts that are inserted into the student’s frozen ViT backbone, but they’re completely separate from the main prompt pool used for continual learning.
Tom: And the key difference is that these KD prompts are global — they’re shared across all tasks. They don’t get re-selected for each new task. So when the student learns from the teacher on task one, that knowledge accumulates in the KD prompts, and it’s still there for task two, task three, and so on.
Jane: That’s the elegant part. The original prompts are task-specific because that’s what helps with continual learning — each task gets its own space. But the KD prompts are task-agnostic. They’re like a permanent notebook that the student keeps adding to, rather than throwing away.
Tom: And they didn’t stop there. They also added a distillation token and a separate classifier, kind of like what DeiT does. So the student has its main classifier for the actual task, and a KD classifier that specifically learns from the teacher’s logits.
Jane: Right. And the loss function combines both — the student still learns the true labels for the current task, but it also gets a KL divergence loss that pushes its KD classifier to match the teacher’s output. They balance those with a parameter alpha, which they set to zero point five.
Tom: Let’s talk about the results because they’re impressive. On ImageNet-R with CODA-Prompt, the ViT-Small student jumped from sixty-seven point four percent to seventy-one point nine percent with KDP. That’s a four point five point gain. And the forgetting rate dropped from eight point five percent down to five point six percent.
Jane: And the same pattern held across all three continual learning methods they tested — L2P, DualPrompt, and CODA-Prompt. So this isn’t a hack that only works with one setup. It’s a general improvement.
Tom: They also ran ablations to make sure they weren’t just adding parameters for no reason. They compared adding the same number of trainable prompts into the main prompt pool versus using them as KD prompts. The KD prompts won by a clear margin.
Jane: That’s a really important control. It shows the improvement isn’t just from having more trainable parameters. It’s specifically from having prompts that are designed to be global and cross-task.
Tom: And they also tried the obvious alternative — unfreezing part of the ViT backbone to let it learn. That failed badly. The accuracy dropped to about thirty percent and the forgetting rate shot up to fifty-eight percent. So freezing the backbone and using prompts is definitely the right call.
Jane: So the improvement is real, it’s consistent, and it’s principled. The KD prompts solve the exact problem they identified.
Tom: Now I want to bring in some other voices. Lu, you’re the researcher — what do you think about this approach? Is there a bigger implication here?
Lu: I think the big implication is that this opens the door for continual learning on edge devices. You can train a large teacher once, then continually distill into a small student that can actually run on a phone or a robot. The student keeps learning new tasks without forgetting, and it stays fast.
Tom: That’s a great point. And Meng, from the engineering side — does this actually hold up in practice?
Meng: The main cost is that you have to train the teacher alongside the student, which doubles the training time. But for deployment, the student is what matters. And the paper shows the student gets close to the teacher’s accuracy while being much cheaper to run. That’s a trade-off most teams would take.
Tom: Alright, we’ve got the problem, the solution, and the results. Let’s wrap this up in the next segment.
Conclusion: Tom: And we’re at the finish line for “Continual Distillation Learning: Knowledge Distillation in Prompt-based Continual Learning.” Jane, give us the final summary.
Jane: Sure. The paper identifies a real problem — standard knowledge distillation doesn’t work well in prompt-based continual learning because the prompt selection mechanism causes distilled knowledge to be lost between tasks. They call this distillation information forgetting.
Tom: And their fix is KD prompts — global, task-agnostic prompts that sit outside the main prompt pool and accumulate knowledge across all tasks. Combined with a distillation token and classifier, this gives consistent improvements across three different continual learning methods and two datasets.
Jane: The numbers speak for themselves. On ImageNet-R, the small student went from sixty-seven point four percent to seventy-one point nine percent with KDP. On CIFAR-one hundred from eighty-two point two percent to eighty-four point three percent. And the forgetting rate dropped significantly in both cases.
Tom: And the ablation studies show it’s not just about adding parameters — it’s about where you put them. The KD prompts are specifically designed to be cross-task, and that makes all the difference.
Lu: I’d add that this has real-world implications for deploying continual learning on resource-constrained devices. You get the accuracy of a big model with the speed of a small one, and the student can keep learning new tasks over time.
Meng: And from an engineering standpoint, the trade-off is clear. You spend more time training, but you get a deployable model that’s much cheaper to run. That’s a win for most real-world applications.
Lalam: And I think there’s a cultural angle too. This kind of efficient continual learning could make adaptive AI systems more accessible — think assistive technologies that learn new user preferences over time without needing a huge server farm behind them.
Tom: That’s a beautiful way to end it. So, to the authors — Qifan Zhang, Yunhui Guo, and Yu Xiang — great work. You’ve given us a new problem, a clear diagnosis, and a clever solution.
Jane: And to our listeners, thanks for tuning in. We’ll be back with another paper soon. Until then, keep learning — and don’t forget what you learned yesterday.
Tom: See you next time, everyone.
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language