Denser not equal to Better: Limits of On-Policy Self-Distillation for Continual Post-Training
summary
The gist
Continual post-training enables foundation models to acquire new knowledge while preserving existing capabilities, and this work investigates whether on-policy self-distillation (SDPO) can reliably
In short
The paper investigates whether on-policy self-distillation (SDPO) can stabilize continual learning by using teacher signals to guide model updates. While SDPO can speed up in-domain specialization when teacher signals are stable, it struggles with out-of-distribution data and exhibits stronger forgetting than sequence-level methods like GRPO. The key takeaway is that dense supervision alone does not guarantee better results.
Key concepts
- On-policy Self-Distillation (SDPO)
- A method where a model learns by distilling knowledge from its own previous outputs, using teacher signals to guide updates. The paper tests if this process can reliably preserve old skills while learning new ones during continual post-training.
- Teacher Signal Stability
- The quality and consistency of the signals provided by the teacher model are crucial for SDPO success. If teacher signals are unstable or noisy, they introduce artifacts and domain mismatches into the student model, leading to poor performance rather than improvement.
- Drift and Collapse
- These describe negative changes in the model during training. Drift refers to parameter shifts where the model moves farther from its original state. Collapse occurs when the model amplifies artifacts, such as repeating specific tokens, due to a confirmation bias loop with the teacher.
- Sequence-level On-Policy RL (e.g., GRPO)
- These methods use rewards based on entire sequences of actions or tokens to update the model. The paper contrasts SDPO with these methods, showing that sequence-level rewards provide a more conservative update bias, resulting in better retention and less forgetting.
Terminology used across episodes
This episode discusses
- Denser not equal to Better: Limits of On-Policy Self-Distillation for Continual Post-Training · Paper Radio
- Constitutional AI: Harmlessness from AI Feedback
- Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Reinforcement Learning via Self-Distillation
- Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?
- Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-Training
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
- Let's Verify Step by Step
- ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning
- Olmo 3
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- RL's Razor: Why Online Reinforcement Learning Forgets Less
- Self-Distillation Enables Continual Learning
- Learning by Distilling Context
- Qwen3 Technical Report
- Solving math word problems with process- and outcome-based feedback
- On-Policy Context Distillation for Language Models
- Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
The paper
Denser not equal to Better: Limits of On-Policy Self-Distillation for Continual Post-Training · Read on arXiv
Meng Wang, Haohan Zhao, Wenzhuo Liu, Lu Yang, Geng Liu, Haiyang Guo, Guo-Sen Xie, Gaofeng Meng
Centre for Artificial Intelligence and Robotics, HKISI, CAS 2 Institute of Automation, CAS University of Chinese Academy of Sciences 4Nanjing University of Science and Technology
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Denser not equal to Better".
Tom: Continual post-training enables foundation models to acquire new knowledge while preserving existing capabilities, and this work investigates whether on-policy self-distillation (SDPO) can reliably serve as a stabilizer for continual learning.
Jane: First, who's behind it and why it matters.
Title and authors: Tom: So, to recap, we're looking at the paper "Denser not equal to Better: Limits of On-Policy Self-Distillation for Continual Post-Training" and the authors are essentially challenging the optimistic idea that on-policy self-distillation is a straightforward way to acquire new skills while preserving what the model already knows.
Jane: Exactly, Tom, they lay out how this approach can accelerate in specific areas like in-domain specialization when things are stable, but then they show it struggles significantly when we try to move into out-of-distribution scenarios.
Lu: The authors make a very clear point about the ingredients involved: the data source and the objective function determine which continuation distribution gets reinforced at each step, which is a big concept in how we think about model updates.
Meng: I wonder what this means for practical deployment; if we rely too much on dense token supervision, are we setting ourselves up for instability when encountering novel inputs?
Lalam: For Lalam, the implication is that simply feeding the model more dense signals isn't enough; we need a method that respects the structure of both the old and new knowledge simultaneously.
The paper's summary: Tom: Moving into what they actually show, this paper summarizes how SDPO performs relative to other methods like GRPO across various tasks, and it highlights that SDPO can exhibit stronger forgetting than sequence-level on-policy reinforcement learning methods like GRPO when things aren't perfectly set up.
Jane: It’s interesting because the summary points out that the effectiveness of SDPO is highly conditional on the quality and stability of those teacher signals, which is a major caveat they bring up.
Lu: They introduce a specific trade-off: increasing supervision density strengthens the local learning signal but also boosts sensitivity, domain mismatch, and accumulated artifacts in a way that isn't always beneficial.
Meng: That trade-off sounds like it could be a headache for real-world systems; we want to learn fast, but if we introduce too much noise or sensitivity, the model just becomes brittle.
Lalam: From Lalam’s view, this confirms that the signals themselves have to be high quality and predictable for any form of distillation to work effectively in a continual setting.
The paper's improvements: Tom: Now let's talk about what the paper suggests we should improve or change based on their findings; they propose strategies like "StableSDPO" which is a periodic refresh-and-freeze approach to manage that instability caused by teacher signal volatility.
Jane: That strategy is key because it shows that freshness alone doesn't guarantee good performance; you need the teacher signals to be delivered in a stable way for that frequency of updates to actually help.
Lu: They also look at how supervision structure matters, showing that more supervision only helps when it’s reliable, and they point out how Chain-of-Thought tokens can actually harm long or artifact-prone rationales in tasks like MATH and SCIENCE because the intermediate tokens aren't always tied to final correctness.
Meng: That suggests we shouldn't just blindly add more token targets; we need a smart system to weigh those targets based on their reliability for the specific task at hand, which is a tough thing to build reliably.
Lalam: Lalam thinks this means future AI culture should focus less on brute-force data input and more on building intelligent filters that assess the trustworthiness of every piece of distilled knowledge before integrating it.
Conclusion: Tom: Wrapping things up for "Denser not equal to Better: Limits of On-Policy Self-Distillation for Continual Post-Training," the paper concludes that SDPO is a powerful but fragile training signal that can increase drift and collapse, exhibiting weaker retention than GRPO in some continual post-training settings.
Jane: So, the main takeaway is that token supervision can be a rapid specialization signal, but it carries a real risk of introducing drift and interference if not handled carefully.
Lu: The authors stress that the future success of on-policy self-distillation depends entirely on safeguarding the signal through smart token weighting or teacher controls to make it more selective and stable.
Meng: For practical engineering, this means we need to focus our efforts on those stability strategies—like periodic freezing—instead of just chasing higher supervision densities across the board.
Lalam: Lalam agrees; the lesson for future AI systems is that token supervision can be a rapid specialization signal, but it's a potentially dangerous signal if we don't build in safeguards to keep it stable and artifact-aware.
More episodes
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck