RL Forgets! Towards Continual Policy Optimization

arXiv:2607.04364 · cs.LG · Submitted 2026-07-05 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "RL Forgets! Towards Continual Policy Optimization".

Jane: Continual post-training for vision-language models (VLMs) has become central to adapting them to evolving tasks, but existing evidence suggests that reinforcement learning (RL) is not inherently robust against catastrophic forgetting.

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So, this paper argues that just because we use reinforcement learning to adapt vision-language models doesn't mean they're safe from forgetting when learning new things continually. The authors set up a new benchmark called MRCL to test this assumption across diverse multimodal reasoning tasks.

Jane: They found that standard reinforcement learning methods actually suffer from severe catastrophic forgetting when applied in this continual post-training setting, which challenges the common idea that RL naturally prevents forgetting.

Lu: That finding is significant because it suggests the existing evidence we rely on might be missing important nuances because they are often drawn from outdated or very narrow datasets.

Meng: So, the core problem they identify is an objective mismatch between how current-task learning methods look and what causes old tasks to be forgotten.

Tom: Exactly! The paper points out that the KL regularization used in common policy optimization, like in GRPO, is only looking at data from the new task.

Jane: But catastrophic forgetting actually happens because of behavioral drift on prior-task distributions, which requires evaluating constraints on historical data instead.

Lu: That distinction between current-task and prior-task evaluation seems to be the central technical flaw they are addressing in their research.

Meng: It’s interesting how they framed it as an objective mismatch rather than just a simple forgetting problem, which gives us a better way to think about fixing it computationally.

Tom: And that mismatch is what motivates their new approach, which they call Continual Policy Optimization or CPO.

Jane: They propose CPO as a replay-free framework that tries to constrain prior-task behavioral drift without needing to store any old data at all.

Lu: The idea of relaxing the intractable historical KL constraint into a regularization term based on empirical parameter movement is a very creative way to tackle this without the massive memory overhead of storing old samples.

Meng: I appreciate the replay-free aspect; that drastically simplifies deployment for real-world applications where we can't keep infinite task data.

Tom: So, they showed that CPO works quite well, with Qwen3-VL-8B showing a reduction in forgetting by thirteen point seven percent and an improvement in pretrained capability of seven point zero percent.

Jane: Those results on Qwen3-VL-8B are pretty compelling evidence that this new optimization strategy is effective at preserving prior knowledge while still allowing for new learning.

Lu: It’s a strong indication that we can build more robust continual learning systems by focusing on the constraint mechanism rather than just relying on standard RL loss functions.

Meng: The scalability across different VLM scales mentioned in their experiments suggests this isn't just a small tweak for one model size; it has broader applicability.

Conclusion: Tom: So, wrapping up this discussion on "RL Forgets! Towards Continual Policy Optimization," we’ve seen how the authors identified a gap in current RL research regarding continual post-training for vision-language models and then proposed CPO as a solution that fixes the underlying objective mismatch.

Jane: Essentially, they showed that by focusing on parameter movement regularization instead of trying to perfectly constrain old data distributions, we can keep the model stable across many tasks without needing to keep massive archives of past experiences.

Lu: The implication here is that we can start designing continual learning frameworks based on understanding these underlying objective constraints rather than just applying standard RL loss functions blindly.

Meng: From a practical standpoint, this means our systems could be much more efficient in deployment because they don't need to burden themselves with storing every single task they’ve ever done.

Tom: It really shifts the focus toward how we structure the optimization process itself, rather than just feeding it more data.

Jane: I think the authors' work, specifically introducing MRCL and then developing CPO, gives us a much stronger set of tools to evaluate and improve these complex adaptation scenarios for AI models.

Lu: It opens up avenues for exploring how we can systematically design policy updates that respect the history of learned behaviors while aggressively pursuing new capabilities.

Meng: I think the real impact is in making large vision-language models more reliable and maintainable when they are put into production environments where they have to handle a stream of evolving tasks.

Tom: It’s clear this work provides a principled way to study continual post-training, showing that careful constraint design can actually lead to better preservation of old knowledge.

Jane: We should really keep an eye on how CPO scales and performs when we apply these concepts to even larger models or more complex multimodal scenarios beyond what they tested.

Lu: I’m excited to see the next steps where this framework is applied to truly open-ended, long-term learning challenges for these massive AI systems.

Meng: We'll be watching closely how the engineering team implements this; if it holds up under heavy load, it could really make a difference in how we deploy advanced vision models.

School of Computer Science and Engineering, Southeast University

cs.LG

Submitted: 2026-07-05

Updated: 2026-09-28

Code: https://github.com/MaolinLuo/CPO

Importance score: 88/100

The gist: Continual post-training for vision-language models (VLMs) has become central to adapting them to evolving tasks, but existing evidence suggests that reinforcement learning (RL) is not inherently

Key concepts

MRCL
A new benchmark for continual learning built from diverse multimodal datasets released in 2025 or later. It is designed to reduce contamination from pre-training data and provide a reliable testbed for evaluating how well models retain knowledge when continually adapting to new tasks.
Objective Mismatch
Standard RL uses KL regularization based on current task data to prevent over-optimization. However, forgetting occurs because the true problem is behavioral drift on old tasks. The mismatch means the current objective doesn't effectively constrain drift toward prior-task distributions.
Continual Policy Optimization (CPO)
A novel, replay-free framework that prevents forgetting without storing historical data. It relaxes the difficult constraint of penalizing old data by using sparse parameter movement regularization to limit policy drift based on an approximation of the prior-task KL divergence.
Prior-Task Behavioral KL Objective
The theoretical ideal is to constrain the new policy's behavior relative to previous tasks using a KL divergence calculated on old data. Since this is intractable, CPO approximates this constraint by using an empirical measure of parameter movement instead, ensuring the model respects prior task behaviors without needing old data.

Terminology

Summary

Continual post-training for vision-language models (VLMs) has become central to adapting them to evolving tasks, but existing evidence suggests that reinforcement learning (RL) is not inherently robust against catastrophic forgetting. This work introduces a new benchmark and a novel replay-free framework, Continual Policy Optimization (CPO), to demonstrate that standard RL methods still suffer from severe forgetting during continual post-training by addressing the objective mismatch between current-task KL regularization and prior-task behavioral drift.

The gist

Continual RL still suffers from catastrophic forgetting under challenging evaluation settings when applied to vision-language models.

Problem Identification and Benchmark Introduction

The authors revisit the assumption that reinforcement learning is inherently robust to forgetting by introducing MRCL, a Multimodal Reasoning Continual Learning benchmark built from diverse multimodal datasets released in 2025 or later. This benchmark is designed to reduce pre-training data contamination and provide a reliable testbed for evaluating forgetting in continual post-training, addressing blind spots in existing evaluations where benchmarks are often outdated or tasks are too narrow. Experiments on MRCL reveal that continual RL still suffers from catastrophic forgetting, challenging the view that RL naturally solves continual learning.

Failure Analysis: Objective Mismatch

The research traces the failure of standard RL objectives to prevent forgetting to an objective mismatch. Specifically, it argues that the KL regularization used in common policy optimization methods (like GRPO) is evaluated on current-task data, i.e., Ex∼Tnew [KL (πθ∥πold)], which is mainly designed to prevent over-optimization. However, catastrophic forgetting is caused by behavioral drift on prior-task distributions, which requires evaluating the constraint on historical data, i.e., Ex∼Told [KL (πθ∥πold)]. The authors show that current-task KL can only serve as a forgetting proxy when the new-task distribution is closed to the prior-task distribution.

Proposed Solution: Continual Policy Optimization (CPO)

To constrain prior-task behavioral drift without accessing historical data, the paper proposes Continual Policy Optimization (CPO), a replay-free framework grounded in a prior-task behavioral KL objective. CPO relaxes the intractable historical KL constraint into sparse parameter-movement regularization, thereby limiting policy drift without storing old data. This is achieved by approximating the oracle prior-task KL divergence using an empirical parameter movement measure.

Implementation Details and Results

CPO maintains a cumulative protected set, initialized as S(0) = ∅, and optimizes the current policy with RL loss L RL i(θ) plus a masked L1 movement regularization term: L i(θ) = L RL i(θ) + λ m(i-1) 0 m(i-1) ⊙ (θ − ˜θ). The protected set S(i-1) is updated by adding the top p% moving coordinates after training task Ti. Extensive experiments across multiple model scales show that CPO consistently reduces forgetting while preserving, and in some cases improving, pretrained model capabilities. On Qwen3-VL-8B, CPO reduces forgetting by 13.7% and improves pretrained capability by 7.0%.

Key Contributions

The key contributions of this work are:

  1. Introducing MRCL, a new continual learning benchmark built from diverse multimodal datasets released in 2025 or later, which reduces the risk of pre-training data contamination.

  2. Proposing CPO, a replay-free continual RL framework grounded in the prior-task behavioral KL objective, which theoretically relaxes the intractable historical KL constraint into a parameter-movement surrogate to preserve prior-task behavior without storing old data.

  3. Demonstrating that CPO substantially reduces forgetting and better preserves pretrained capabilities compared to post-training baselines across multiple VLM scales.

Evaluation Metrics

The evaluation uses three aggregated accuracy metrics: Mean Finetune Accuracy (MFT), Mean Final Accuracy (MFN), and Mean Task Accuracy (MTA). These metrics are used to evaluate the continual learning process, showing that CPO improves final retention while preserving strong task-level adaptation, as evidenced by its superior performance in Table 10 across different model scales. The paper also introduces Cumulative mean Accuracy Improvement (CAI) and Cumulative mean Accuracy Delta (CAD) to quantify the dynamics of plasticity and stability visualized in Figure 3.

Conclusion

The work concludes that CPO provides a principled, replay-free optimization framework for studying continual post-training in modern vision-language models by constraining policy drift on protected parameters rather than penalizing current-task behavior. This approach is shown to be scalable to large VLMs due to its negligible computational and memory overhead. The combination of MRCL and CPO provides both a stronger benchmark and a principled framework for studying continual post-training in modern vision-language models.

Improvements for AI systems

Here are specific improvements to AI systems derived from the research presented in this paper, along with what these improved systems can achieve:


) Continual Policy Optimization (CPO) Framework Integration:

The core improvement is replacing standard reinforcement learning (RL) policy optimization methods (like GRPO or GSPO) with the proposed CPO framework. This framework addresses catastrophic forgetting by grounding the constraint not on current-task data, but on a computationally feasible surrogate for prior-task behavioral KL divergence derived from parameter movement.

  1. Improvement: Implement a replay-free continual RL mechanism that uses Parameter Movement as a noisy Fisher information proxy to constrain policy drift without storing historical data or replaying old samples.

  2. System Capability: AI models can undergo sequential task adaptation (continual post-training) on massive, dynamic datasets (e.g., evolving medical images, changing navigation environments) while maintaining high performance on previously learned tasks. This is crucial for real-world deployment where models must continuously learn new procedures or adapt to novel sensor data without forgetting fundamental skills.

) Multimodal Reasoning Benchmark and Evaluation:

The introduction of the Multimodal Reasoning Continual Learning benchmark (MRCL), built from diverse, recent multimodal datasets (2025+), provides a rigorous testing ground that moves beyond outdated benchmarks.

  1. Improvement: Establish and utilize the MRCL benchmark to systematically evaluate the plasticity-stability trade-off in Vision-Language Models (VLMs) across domains like medical understanding, complex geometry, navigation, and financial chart analysis.

  2. System Capability: AI developers can move past homogeneous datasets that mask forgetting issues. By training on MRCL, models become demonstrably more robust to out-of-domain knowledge decay across a wide spectrum of reasoning tasks—from identifying anatomical structures (MedBookVQA) to solving complex spatial puzzles (Puzzle).

) Objective Mismatch Diagnosis and Constraint Derivation:

The research provides a deep theoretical insight into why standard RL fails in continual learning: the objective mismatch between evaluating KL divergence on current-task data versus behavioral drift on prior-task distributions.

  1. Improvement: Develop diagnostic tools to identify specific failure modes in existing RL agents during adaptation, such as distinguishing between over-optimization (current-task KL) and behavioral drift (prior-task KL).

  2. System Capability: Researchers can precisely engineer the loss functions of continual learning algorithms. Instead of blindly applying a general KL penalty, they can tailor the regularization to target the specific type of forgetting occurring in a given model architecture or data stream, leading to more efficient and effective adaptation strategies.

) Scalability for Large Models:

The CPO framework specifically addresses the computational bottleneck associated with traditional importance-based methods like Exact Fisher Information Matrix (FIM) estimation by using diagonal relaxation and empirical parameter movement proxies.

  1. Improvement: Create lightweight, scalable RL training pipelines suitable for very large VLMs (e.g., Qwen3-VL-8B) that require only standard parameter updates and a small set of historical movement statistics, rather than expensive replaying or dense matrix computations.

  2. System Capability: Continual learning becomes feasible for state-of-the-art foundation models that are prohibitively large for methods requiring full Fisher estimation. This allows for the development of truly massive, continuously learning AI systems used in high-stakes environments where computational resources are limited but data streams are constant.

) Precision in Policy Update Selection:

The paper demonstrates that using masked L1 regularization (as proposed in CPO) leads to sparse parameter movement, whereas dense L2 regularization can lead to cognitive collapse by allowing too much drift.

  1. Improvement: Implement adaptive, sparsity-inducing regularizers (like Masked L1) during the RL phase of continual learning.

  2. System Capability: The resulting AI model exhibits superior stability and retention of prior knowledge, ensuring that adaptation is precise—preserving critical learned features while only allowing necessary parameter adjustments for new tasks, thus preventing catastrophic forgetting during complex sequential learning regimes.

Sources

Related papers