RL Forgets! Towards Continual Policy Optimization

summary

Video file (mp4)

The gist

Continual post-training for vision-language models (VLMs) has become central to adapting them to evolving tasks, but existing evidence suggests that reinforcement learning (RL) is not inherently

In short

Standard reinforcement learning methods fail to prevent catastrophic forgetting when vision-language models are continually post-trained. This research introduces MRCL, a new benchmark, and Continual Policy Optimization (CPO), a replay-free framework. CPO solves the problem by replacing an impossible constraint on old data with a parameter movement regularization technique to limit policy drift.

Key concepts

MRCL
A new benchmark for continual learning built from diverse multimodal datasets released in 2025 or later. It is designed to reduce contamination from pre-training data and provide a reliable testbed for evaluating how well models retain knowledge when continually adapting to new tasks.
Objective Mismatch
Standard RL uses KL regularization based on current task data to prevent over-optimization. However, forgetting occurs because the true problem is behavioral drift on old tasks. The mismatch means the current objective doesn't effectively constrain drift toward prior-task distributions.
Continual Policy Optimization (CPO)
A novel, replay-free framework that prevents forgetting without storing historical data. It relaxes the difficult constraint of penalizing old data by using sparse parameter movement regularization to limit policy drift based on an approximation of the prior-task KL divergence.
Prior-Task Behavioral KL Objective
The theoretical ideal is to constrain the new policy's behavior relative to previous tasks using a KL divergence calculated on old data. Since this is intractable, CPO approximates this constraint by using an empirical measure of parameter movement instead, ensuring the model respects prior task behaviors without needing old data.

Terminology used across episodes

This episode discusses

The paper

RL Forgets! Towards Continual Policy Optimization · Read on arXiv

School of Computer Science and Engineering, Southeast University

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "RL Forgets! Towards Continual Policy Optimization".

Jane: Continual post-training for vision-language models (VLMs) has become central to adapting them to evolving tasks, but existing evidence suggests that reinforcement learning (RL) is not inherently robust against catastrophic forgetting.

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So, this paper argues that just because we use reinforcement learning to adapt vision-language models doesn't mean they're safe from forgetting when learning new things continually. The authors set up a new benchmark called MRCL to test this assumption across diverse multimodal reasoning tasks.

Jane: They found that standard reinforcement learning methods actually suffer from severe catastrophic forgetting when applied in this continual post-training setting, which challenges the common idea that RL naturally prevents forgetting.

Lu: That finding is significant because it suggests the existing evidence we rely on might be missing important nuances because they are often drawn from outdated or very narrow datasets.

Meng: So, the core problem they identify is an objective mismatch between how current-task learning methods look and what causes old tasks to be forgotten.

Tom: Exactly! The paper points out that the KL regularization used in common policy optimization, like in GRPO, is only looking at data from the new task.

Jane: But catastrophic forgetting actually happens because of behavioral drift on prior-task distributions, which requires evaluating constraints on historical data instead.

Lu: That distinction between current-task and prior-task evaluation seems to be the central technical flaw they are addressing in their research.

Meng: It’s interesting how they framed it as an objective mismatch rather than just a simple forgetting problem, which gives us a better way to think about fixing it computationally.

Tom: And that mismatch is what motivates their new approach, which they call Continual Policy Optimization or CPO.

Jane: They propose CPO as a replay-free framework that tries to constrain prior-task behavioral drift without needing to store any old data at all.

Lu: The idea of relaxing the intractable historical KL constraint into a regularization term based on empirical parameter movement is a very creative way to tackle this without the massive memory overhead of storing old samples.

Meng: I appreciate the replay-free aspect; that drastically simplifies deployment for real-world applications where we can't keep infinite task data.

Tom: So, they showed that CPO works quite well, with Qwen3-VL-8B showing a reduction in forgetting by thirteen point seven percent and an improvement in pretrained capability of seven point zero percent.

Jane: Those results on Qwen3-VL-8B are pretty compelling evidence that this new optimization strategy is effective at preserving prior knowledge while still allowing for new learning.

Lu: It’s a strong indication that we can build more robust continual learning systems by focusing on the constraint mechanism rather than just relying on standard RL loss functions.

Meng: The scalability across different VLM scales mentioned in their experiments suggests this isn't just a small tweak for one model size; it has broader applicability.

Conclusion: Tom: So, wrapping up this discussion on "RL Forgets! Towards Continual Policy Optimization," we’ve seen how the authors identified a gap in current RL research regarding continual post-training for vision-language models and then proposed CPO as a solution that fixes the underlying objective mismatch.

Jane: Essentially, they showed that by focusing on parameter movement regularization instead of trying to perfectly constrain old data distributions, we can keep the model stable across many tasks without needing to keep massive archives of past experiences.

Lu: The implication here is that we can start designing continual learning frameworks based on understanding these underlying objective constraints rather than just applying standard RL loss functions blindly.

Meng: From a practical standpoint, this means our systems could be much more efficient in deployment because they don't need to burden themselves with storing every single task they’ve ever done.

Tom: It really shifts the focus toward how we structure the optimization process itself, rather than just feeding it more data.

Jane: I think the authors' work, specifically introducing MRCL and then developing CPO, gives us a much stronger set of tools to evaluate and improve these complex adaptation scenarios for AI models.

Lu: It opens up avenues for exploring how we can systematically design policy updates that respect the history of learned behaviors while aggressively pursuing new capabilities.

Meng: I think the real impact is in making large vision-language models more reliable and maintainable when they are put into production environments where they have to handle a stream of evolving tasks.

Tom: It’s clear this work provides a principled way to study continual post-training, showing that careful constraint design can actually lead to better preservation of old knowledge.

Jane: We should really keep an eye on how CPO scales and performs when we apply these concepts to even larger models or more complex multimodal scenarios beyond what they tested.

Lu: I’m excited to see the next steps where this framework is applied to truly open-ended, long-term learning challenges for these massive AI systems.

Meng: We'll be watching closely how the engineering team implements this; if it holds up under heavy load, it could really make a difference in how we deploy advanced vision models.

More episodes

← Home