Influence-Directed Distillation: Solving the Diversity Bottleneck in Sampled-Token On-Policy Distillation

summary

Video file (mp4)

The gist

The paper introduces Influence-Directed Distillation, a novel framework designed to overcome the critical "diversity bottleneck" inherent in traditional sampled-token on-policy distillation methods

In short

The discussion of Influence-Directed Distillation addresses a failure in Sampled-Token On-Policy Distillation where models lose diversity despite improving accuracy. The researchers propose using a diagnostic tool, First-Order Local Entropy Influence, to identify problematic updates. They then apply Divergence-Adaptive Shrinkage to selectively correct these entropy drains, resulting in highly efficient models that maintain both accuracy and broad solution variety.

Key concepts

Sampled-Token On-Policy Distillation (OPD) Failure
This is a problem where a model gets better at finding a single correct answer (pass@one increases), but fails to inherit the variety of the teacher model, causing performance on multiple answers (pass@k) to plateau. This failure stems from specific entropy contraction in the model's decision-making process.
First-Order Local Entropy Influence (IH(y))
This diagnostic tool provides a single signal that measures whether an update is shrinking or expanding the overall uncertainty of the model. It helps identify if a model is becoming too certain, which limits its adaptability and potential for novel ideas.
Divergence-Adaptive Shrinkage
Instead of applying a blanket penalty to all entropy-reducing updates, this intervention strategically punishes only those updates causing the most trouble. It uses an adaptive weight based on how much the student and teacher already agree, achieving targeted correction without uniform penalties.

Terminology used across episodes

This episode discusses

The paper

Influence-Directed Distillation: Solving the Diversity Bottleneck in Sampled-Token On-Policy Distillation · Read on arXiv

Run Yang, Runpeng Dai, Jie Sun, Jielei Zhang, Fan Zhou, Hongtu Zhu, Peiyi Li, Longwen Gao

BiliBili Inc. · University of North Carolina at Chapel Hill · University of Science and Technology of China · Shanghai University of Finance and Economics

Sampled-token on-policy distillation (OPD) efficiently transfers capabilities from teacher to student using student-generated tokens, requiring teacher probabilities only for sampled tokens. Yet it frequently suffers from diversity distillation failure: the student's pass@1 improves while its pass@ k plateaus, failing to inherit the teacher's diversity. To explain this, we introduce First-Order Local Entropy Influence, a signed first-order proxy that decouples each update's entropy effect into the teacher--student log-probability gap and the student's local probability structure, and empirically links entropy contraction to negative-influence positions. Motivated by this, we propose Influence-Directed Adaptive On-Policy Distillation (IDA-OPD): rather than relying on costly full-vocabulary Forward-KL objectives, it preserves entropy-expanding updates while replacing entropy-contracting ones with divergence-adaptive advantage shrinkage, using only the teacher's sampled-token log-probability. Experiments on reasoning-oriented distillation show IDA-OPD consistently improves pass@ k, inheriting the teacher's diversity through distillation, matches the strongest teacher-informed methods at strictly lower cost, and broadly maintains vanilla OPD's pass@1, all without full-vocabulary teacher information.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Influence-Directed Distillation: Solving the Diversity Bottleneck in Sampled-Token On-Policy Distillation".

Jane: The paper was written by Run Yang, Runpeng Dai, Jie Sun, Jielei Zhang, Fan Zhou et al. from BiliBili Inc. and University of North Carolina at Chapel Hill and University of Science and Technology of China and Shanghai University of Finance and Economics.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Summary: Tom: So, having established the problem with sampled-token OPD, the authors of Influence-Directed Distillation explain *why* it fails by looking at entropy. They found that even though a model might be getting better at finding a correct answer—which is reflected in pass@one improving—it’s failing to inherit the teacher’s variety, which is measured by pass@k plateauing.

Jane: That disparity between passing just one question versus passing multiple questions is what the researchers are trying to solve. The core of their summary is that this diversity failure comes down to a specific type of entropy contraction happening in the model's decision-making process.

Lu: The theory here is subtle but powerful: they found that you can’t just look at the difference between what your student model thinks and what the teacher model thinks; you need to look at how that difference changes the local probability structure of each token. That’s where the math gets really elegant.

Meng: I appreciate them providing a clear diagnostic tool, which is called First-Order Local Entropy Influence, IH(y). It gives us a single signal to understand if an update is shrinking or expanding the overall uncertainty of the model.

Lalam: This concept of entropy as a measure of "uncertainty" is very helpful because it tells us whether the model is getting too certain, which in our world translates to less adaptability and less potential for novel ideas.

Tom: That makes sense; it’s a way to measure if the model is collapsing its own possible solutions. But since this tool, IH(y), is so powerful, Jane wonders how they use it to actually fix the problem without just applying a blanket penalty across all of those entropy-contracting updates.

Improvements: Jane: The next segment is where things get very smart; they propose a specific intervention called Divergence-Adaptive Shrinkage. Instead of just punishing every update that reduces entropy, they only punish the ones that are actually causing the most trouble.

Tom: It’s not just any low-divergence updates that cause problems, though; the authors show through Figure two that these massive populations of highly aligned tokens are where the biggest aggregate entropy drain happens. So, instead of a uniform penalty, they use a divergence-adaptive weight, w y.

Lu: I think this is the most clever part of their design because it’s not just shrinking everything that goes down; they are strategically deciding which ones to shrink based on how much the student and teacher already agree. If you agree, you shrink the correction; if you disagree, you preserve it.

Meng: From an engineering standpoint, this adaptive weight is a massive efficiency win because we’re still only using that single sampled-token log-probability; we aren't bringing back all those expensive full-vocabulary logits that other methods require.

Lalam: The implication for us is that we can achieve greater diversity by making very precise, targeted adjustments instead of applying blunt force, which will lead to more robust and less biased reasoning in the future.

Tom: That precision is the key, so let's look at how this all translates into performance in the final segment.

Conclusion: Jane: We’ve seen that Influence-Directed Distillation uses a clever diagnostic tool to find trouble and then use a highly selective intervention to solve the problem of diversity failure in sampled-token OPD. The results, as shown in Table one are quite impressive for every benchmark they tested.

Tom: They consistently outperform other methods that operate under the same efficiency constraints, while also achieving remarkable pass@k scores on AIME and HMMT problems. It’s a huge leap forward in making these models more capable of complex reasoning.

Lu: The fact that they are matching or even exceeding the teacher's performance on specific benchmarks is a testament to the power of this targeted approach, suggesting that we haven't fully understood how to transfer diversity yet.

Meng: I think the practical impact is huge; by maintaining pass@one accuracy while boosting pass@k, we are building LLMs that can be both highly reliable and genuinely creative. The computational cost is minimal compared to other methods, which is a massive win for deployment.

Lalam: For us at the frontier of AI, this means we can move past simply optimizing for the single most probable answer and start designing models that respect the full spectrum of valid human-level solutions.

Tom: It’s a truly sophisticated piece of work by all those authors, BiliBili, UNC, and their colleagues in China. We'll wrap up our discussion with a final thought from each of us before heading off to the next paper.

Lu: I just hope this opens the door for even more complex interactions between systems that can benefit from this kind of directed learning.

Meng: I'm excited to see how this technology scales into real-world applications, given its low overhead and high performance gains.

Lalam: I believe the future is a model that is not just smart, but diverse, and Influence-Directed Distillation helps achieve that balance.

Tom: A fantastic discussion on Influence-Directed Distillation: Solving the Diversity Bottleneck in Sampled-Token On-Policy Distillation; thanks everyone for sharing their insights.

Conclusion: Tom: So, we’re wrapping up our discussion on Influence-Directed Distillation: Solving the Diversity Bottleneck in Sampled-Token On-Policy Distillation, which is a real game-changer for how we train large language models.

Jane: It successfully demonstrated that by identifying and gently curbing the specific updates that drain the model's entropy—the ones that cause it to become too certain—we can unlock much more diversity without sacrificing accuracy.

Tom: And I think the results are just incredibly strong, consistently beating previous methods while staying highly efficient, which is exactly what we want to see.

Lu: This isn't just about better math scores though; it's about a massive jump in reasoning ability across different domains. Imagine this level of diversity applied to complex scientific hypothesis generation or creative problem-solving in AI systems.

Meng: From an engineering standpoint, the fact that IDA-OPD keeps the computational cost low while achieving these results is truly impressive. We can deploy this on more scale and maintain performance without needing massive compute budgets for full teacher distributions.

Lalam: The cultural impact of having an AI that isn't collapsing its potential pathways is profound. It fosters a system that encourages exploration, leading to better, more nuanced solutions to societal problems rather than just the most statistically probable ones.

Tom: I agree with Lalam; the ability to retain that high-entropy potential is a huge step toward robust reasoning.

Jane: That’s right, we're moving past simply optimizing for pass@one and starting to truly optimize for how much a model can explore all valid solutions.

Lu: It opens up entire new avenues of inquiry in AI design, allowing us to push the boundaries of what's possible with existing hardware.

Meng: And since it’ efficient, we can actually run these massive distillation runs at a higher frequency and scale than before.

Lalam: Ultimately, this means our AI models are becoming more comprehensive partners in discovery rather than just sophisticated calculators.

Tom: It’s clear that Influence-Directed Distillation is a major milestone for LLM training, and I think we've seen the impact across all facets of it today.

Jane: Before heading into the next paper, let's take a quick breath and appreciate this was a fantastic discussion on Influence-Directed Distillation: Solving the Diversity Bottleneck in Sampled-Token On-Policy Distillation.

More episodes

← Home