Denser not equal to Better: Limits of On-Policy Self-Distillation for Continual Post-Training
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Denser not equal to Better".
Tom: Continual post-training enables foundation models to acquire new knowledge while preserving existing capabilities, and this work investigates whether on-policy self-distillation (SDPO) can reliably serve as a stabilizer for continual learning.
Jane: First, who's behind it and why it matters.
Title and authors: Tom: So, to recap, we're looking at the paper "Denser not equal to Better: Limits of On-Policy Self-Distillation for Continual Post-Training" and the authors are essentially challenging the optimistic idea that on-policy self-distillation is a straightforward way to acquire new skills while preserving what the model already knows.
Jane: Exactly, Tom, they lay out how this approach can accelerate in specific areas like in-domain specialization when things are stable, but then they show it struggles significantly when we try to move into out-of-distribution scenarios.
Lu: The authors make a very clear point about the ingredients involved: the data source and the objective function determine which continuation distribution gets reinforced at each step, which is a big concept in how we think about model updates.
Meng: I wonder what this means for practical deployment; if we rely too much on dense token supervision, are we setting ourselves up for instability when encountering novel inputs?
Lalam: For Lalam, the implication is that simply feeding the model more dense signals isn't enough; we need a method that respects the structure of both the old and new knowledge simultaneously.
The paper's summary: Tom: Moving into what they actually show, this paper summarizes how SDPO performs relative to other methods like GRPO across various tasks, and it highlights that SDPO can exhibit stronger forgetting than sequence-level on-policy reinforcement learning methods like GRPO when things aren't perfectly set up.
Jane: It’s interesting because the summary points out that the effectiveness of SDPO is highly conditional on the quality and stability of those teacher signals, which is a major caveat they bring up.
Lu: They introduce a specific trade-off: increasing supervision density strengthens the local learning signal but also boosts sensitivity, domain mismatch, and accumulated artifacts in a way that isn't always beneficial.
Meng: That trade-off sounds like it could be a headache for real-world systems; we want to learn fast, but if we introduce too much noise or sensitivity, the model just becomes brittle.
Lalam: From Lalam’s view, this confirms that the signals themselves have to be high quality and predictable for any form of distillation to work effectively in a continual setting.
The paper's improvements: Tom: Now let's talk about what the paper suggests we should improve or change based on their findings; they propose strategies like "StableSDPO" which is a periodic refresh-and-freeze approach to manage that instability caused by teacher signal volatility.
Jane: That strategy is key because it shows that freshness alone doesn't guarantee good performance; you need the teacher signals to be delivered in a stable way for that frequency of updates to actually help.
Lu: They also look at how supervision structure matters, showing that more supervision only helps when it’s reliable, and they point out how Chain-of-Thought tokens can actually harm long or artifact-prone rationales in tasks like MATH and SCIENCE because the intermediate tokens aren't always tied to final correctness.
Meng: That suggests we shouldn't just blindly add more token targets; we need a smart system to weigh those targets based on their reliability for the specific task at hand, which is a tough thing to build reliably.
Lalam: Lalam thinks this means future AI culture should focus less on brute-force data input and more on building intelligent filters that assess the trustworthiness of every piece of distilled knowledge before integrating it.
Conclusion: Tom: Wrapping things up for "Denser not equal to Better: Limits of On-Policy Self-Distillation for Continual Post-Training," the paper concludes that SDPO is a powerful but fragile training signal that can increase drift and collapse, exhibiting weaker retention than GRPO in some continual post-training settings.
Jane: So, the main takeaway is that token supervision can be a rapid specialization signal, but it carries a real risk of introducing drift and interference if not handled carefully.
Lu: The authors stress that the future success of on-policy self-distillation depends entirely on safeguarding the signal through smart token weighting or teacher controls to make it more selective and stable.
Meng: For practical engineering, this means we need to focus our efforts on those stability strategies—like periodic freezing—instead of just chasing higher supervision densities across the board.
Lalam: Lalam agrees; the lesson for future AI systems is that token supervision can be a rapid specialization signal, but it's a potentially dangerous signal if we don't build in safeguards to keep it stable and artifact-aware.
Meng Wang, Haohan Zhao, Wenzhuo Liu, Lu Yang, Geng Liu, Haiyang Guo, Guo-Sen Xie, Gaofeng Meng
Centre for Artificial Intelligence and Robotics, HKISI, CAS 2 Institute of Automation, CAS University of Chinese Academy of Sciences 4Nanjing University of Science and Technology
cs.LG, cs.CL
Submitted: 2026-07-02
Updated: 2026-10-02
Code: https://github.com/Moenupa/SDPO-CL
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 92/100
The gist: Continual post-training enables foundation models to acquire new knowledge while preserving existing capabilities, and this work investigates whether on-policy self-distillation (SDPO) can reliably
Key concepts
- On-policy Self-Distillation (SDPO)
- A method where a model learns by distilling knowledge from its own previous outputs, using teacher signals to guide updates. The paper tests if this process can reliably preserve old skills while learning new ones during continual post-training.
- Teacher Signal Stability
- The quality and consistency of the signals provided by the teacher model are crucial for SDPO success. If teacher signals are unstable or noisy, they introduce artifacts and domain mismatches into the student model, leading to poor performance rather than improvement.
- Drift and Collapse
- These describe negative changes in the model during training. Drift refers to parameter shifts where the model moves farther from its original state. Collapse occurs when the model amplifies artifacts, such as repeating specific tokens, due to a confirmation bias loop with the teacher.
- Sequence-level On-Policy RL (e.g., GRPO)
- These methods use rewards based on entire sequences of actions or tokens to update the model. The paper contrasts SDPO with these methods, showing that sequence-level rewards provide a more conservative update bias, resulting in better retention and less forgetting.
Terminology
Summary
Continual post-training enables foundation models to acquire new knowledge while preserving existing capabilities, and this work investigates whether on-policy self-distillation (SDPO) can reliably serve as a stabilizer for continual learning. The central finding is that while SDPO can accelerate in-domain specialization when teacher signals are stable and well-aligned, it struggles to generalize to out-of-distribution scenarios, exhibiting stronger forgetting than sequence-level on-policy reinforcement learning methods like GRPO.
The Core Problem
The paper explores the optimistic view that on-policy learning can mitigate forgetting in continual post-training, specifically focusing on self-distillation policy optimization (SDPO). While some work suggests SDPO may be a practical recipe for acquiring new skills while preserving prior capabilities, the authors argue that this view conflates two distinct ingredients: where the data come from and which objective is used to update the model. Specifically, when every token becomes a training target, teacher drift, noisy rationales, formatting conventions, and domain-specific artifacts can be reinforced as repeatedly as useful behavior.
This leads to the conclusion that on-policy self-distillation may not automatically inherit the conservative update bias of sequence-level on-policy RFT.
Key Findings on Signal Quality
The effectiveness of SDPO is highly conditional upon the quality and stability of its teacher signals. The authors investigate how supervision density and teacher updates affect performance, revealing a crucial trade-off: supervision density introduces a trade-off: it strengthens the local learning signal, but also increases sensitivity, domain mismatch, and accumulated artifacts.
They demonstrate that freshness alone does not predict performance,
as larger Exponential Moving Average (EMA) update rates can cause instability. To address this, they introduce the StableSDPO
strategy—a periodic refresh-and-freeze approach—which preserves teacher freshness at refresh points while removing step-to-step volatility. This decomposition shows that freshness is useful only when it is delivered through a stable teacher signal.
Impact of Supervision Density and Structure
The paper analyzes the effect of token supervision density, specifically by comparing standard SDPO with variants that incorporate Chain-of-Thought (CoT) tokens. They find that more supervision is useful only when reliable.
CoT distillation benefits short and structured traces, such as those in TOOLUSE training, but can actively harm long or artifact-prone rationales in tasks like MATH and SCIENCE. This contrast is attributed to the reliability of intermediate tokens; for instance, in long-form MATH tasks, CoT tokens are weakly tied to final correctness,
leading to diluted credit assignment and encouragement of imitation of reasoning artifacts.
Drift, Collapse, and Forgetting Mechanisms
SDPO exhibits significant parameter and response drift compared to GRPO. The analysis using Singular Value Decomposition (SVD) metrics—such as Normalized Spectral Shift (NSS), Principal-angle rotation (Max PA), Principal-mask overlap (PM overlap), and Principal-mask update sparsity—reveals that SDPO variants move farther from the base model than GRPO, indicating faster adaptation comes with larger drift.
This drift is diagnosed through three primary failure modes:
-
Parameter and response drift, where SDPO produces
notably larger spectral shifts and more principal-parameter modifications than GRPO,
which hurts continual stability. -
Collapse as artifact amplification, exemplified by SDPO5% collapsing with visible generation artifacts like repeatedly repeated tokens (e.g., “@boxed”). This is attributed to a
confirmation bias loop: the student emits the artifact, the teacher assigns it token-level support, and the next student update makes the artifact more likely.
-
Excess KL drift, where SDPO does not inherit the KL-minimal distribution induced by sequence-level rewards (GRPO). The theorem proves that
SDPO’s on-policy sampling controls where supervision is applied, but the distribution selected at each token is the teacher projection, not the KL-minimal sequence-reward projection,
whichpredict[s] additional forgetting.
Conclusion and Implications
The paper concludes that SDPO offers a powerful yet fragile training signal.
While it can substantially improve performance on the current training domain, it increases the risk of drift, interference, and even collapse. Across single-domain and staged continual post-training settings, SDPO exhibits weaker retention than GRPO,
with forgetting concentrated on neighboring-yet-misaligned tasks. The main lesson is that token supervision can be a rapid specialization signal, but also a potentially dangerous signal.
Therefore, the future of on-policy self-distillation for continual learning depends not on making the signal denser, but on "safeguarding it through token weighting, teacher controls, or data balancing strategies to be more selective, stable, and artifact-aware.
Improvements for AI systems
Thought: I need to act as an excellent, fastidious, and diligent AI researcher who is reading the provided scientific paper (DENSER ̸= BETTER: LIMITS OF ON-POLICY SELFDISTILLATION FOR CONTINUAL POST-TRAINING
). My response must focus solely on the specific improvements this research suggests for AI systems and what these improved systems can achieve. I must be very specific.
The core findings of the paper relate to on-policy self-distillation (SDPO) in continual post-training, contrasting it with sequence-level reward optimization (GRPO). The main insights are:
-
SDPO is a strong specialization method but fragile for continual learning; its benefits depend on teacher stability and token reliability.
-
Dense supervision amplifies both useful signals and harmful artifacts (drift, formatting errors) because token-level targets can reinforce repeated patterns (confirmation bias).
-
Parameter drift is significantly larger under SDPO compared to GRPO, indicating that faster adaptation comes with greater internal instability.
-
Token-level supervision is conditional: it helps short, structured traces (like TOOLUSE), but harms long, artifact-prone rationales (like MATH/SCIENCE) due to the dilution of credit assignment for incorrect intermediate steps.
-
The theoretical analysis shows that SDPO's teacher projection induces
excess KL drift,
meaning it does not inherently inherit the KL-minimal bias of sequence-level RL methods (GRPO).
I will structure my answer around these specific, actionable improvements and the resulting capabilities of the enhanced systems.
Abstract
Continual post-training enables foundation models to acquire new knowledge while preserving existing capabilities. Recent work suggests that on-policy learning can mitigate forgetting, with self-distillation as a particularly attractive approach. We revisit this optimistic claim through self-distillation policy optimization (SDPO). Our experiments show that SDPO accelerates in-domain specialization when teacher signals are stable and well aligned, but struggles to generalize out of distribution. In continual post-training, SDPO exhibits greater forgetting and can even collapse, whereas GRPO, the more established on-policy reinforcement learning method, adapts more conservatively and better preserves prior capabilities. Further analyses link these failures to increased drift in parameter and response space, and to amplification of high-frequency artifacts through a self-reinforcing teacher-student loop. Thus, on-policy data alone is insufficient for continual learning. Self-distillation is effective when teacher targets are stable and token-level supervision is reliable, but should not be treated as a default stabilizer for continual post-training. Our code is available at https://github.com/Moenupa/SDPO-CL.
Sources
- Constitutional AI: Harmlessness from AI Feedback
- Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Reinforcement Learning via Self-Distillation
- Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?
- Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-Training
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
- Let's Verify Step by Step
- ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning
- Olmo 3
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- RL's Razor: Why Online Reinforcement Learning Forgets Less
- Self-Distillation Enables Continual Learning
- Learning by Distilling Context
- Qwen3 Technical Report
- Solving math word problems with process- and outcome-based feedback
- On-Policy Context Distillation for Language Models
- Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks