SelFusion: Self-distillation for Diffusion Language Models

arXiv:2608.22898 · cs.CL · Submitted 2026-08-24 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "SelFusion: Self-distillation for Diffusion Language Models".

Jane: Diffusion language models (DLMs) offer faster inference capabilities compared to autoregressive (AR) large language models (LLMs), making them suitable for real-time applications.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, focusing on the title and authors of "SelFusion: Self-distillation for Diffusion Language Models," what's the main idea they're pushing? It seems like a clever name combining self-distillation with diffusion language models.

Jane: The paper is essentially proposing a new self-distillation framework specifically designed to boost the generation quality of Diffusion Language Models by using their own internal noising process as the source for learning.

Lu: What's compelling about the authors is that they are directly addressing the limitations found when applying conventional knowledge distillation methods to these models, which usually result in little or even negative performance changes.

Meng: I see why they focused on self-distillation; it avoids the whole bottleneck of needing a separate, high-quality teacher model that we'd have to train beforehand.

Lalam: It really shifts the focus from external dependencies to internal model capabilities, which is a very smart direction for advancing AI research.

The paper's summary: Tom: So, summarizing what the paper actually proposes with "SelFusion," it seems they introduce two distinct forward passes within the same DLM: an easy mode and a hard mode.

Jane: That's right, Tom. The key concept is that the easy mode uses a lower masking ratio compared to the hard mode, which theoretically should give it access to more context and thus generate more accurate predictions.

Lu: But they also admit that this easy mode isn't always more accurate than the hard mode and can actually be overconfident when it gets things wrong, so they have to introduce a way to handle that.

Meng: That’s a fair caveat; if the easy mode is overly confident in its errors, then we don't want that knowledge transferring poorly to the student model during distillation.

Lalam: So, they tackle the issue of potential overconfidence head-on by setting up a bidirectional distillation strategy that checks which mode is more correct based on token-level accuracy.

The paper's improvements: Tom: Moving onto the improvements they suggest, SelFusion introduces a dynamic way to pick which way to distill knowledge based on whether both modes were correct or incorrect for a given token.

Jane: It defines three specific rules for this bidirectional distillation strategy, essentially telling the system which mode should be the teacher target depending on the prediction outcomes.

Lu: The mechanism is quite intricate, especially how it dynamically determines the optimal direction based on that token-level correctness check, which seems like a sophisticated way to manage uncertainty.

Meng: From an engineering viewpoint, this dynamic switching adds complexity to the training loop, but if it leads to better knowledge transfer efficiency across different noise levels, that's worth the effort.

Lalam: This dynamic selection process is what makes SelFusion distinct; it moves beyond a fixed one-way distillation and lets the model adapt its learning strategy based on real-time token performance during training.

Conclusion: Tom: To wrap up, we've seen that "SelFusion: Self-distillation for Diffusion Language Models" proposes using two modes—easy and hard—and a dynamic, bidirectional approach to distillation to get better knowledge transfer without an external teacher.

Jane: In essence, the paper shows how leveraging the inherent noising process within a DLM can lead to substantial performance gains over existing methods by intelligently selecting the best teacher signal for each token.

Lu: The results they show, where it even surpasses some teacher models and achieves high scores on benchmarks like Self-inst and Vicuna, suggest that this internal mechanism is genuinely effective at bridging the gap between DLMs and larger models.

Meng: From a practical standpoint, the finding that it eliminates the need for separately trained teacher training really cuts down on computational overhead compared to traditional KD setups.

Lalam: This work opens up a path toward creating self-distilled DLMs that are faster for inference and more accurate in reasoning, which could significantly improve how we deploy high-speed AI applications in our daily lives.

Hyeong Soo Lim, Jin Young Kim, Eun Seo Seo Min Ho Jang Ji Won Yoon

Department of Artificial Intelligence, Chung-Ang University

cs.CL

Submitted: 2026-08-24

Updated: 2026-08-24

Comments: Published as a main conference paper at ACL 2026

Journal ref: Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026), pages 22077-22089, 2026

DOI: 10.18653/v1/2026.acl-long.1008

Code: https://github.com/scai-research/SelFusion_official

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 76/100

The gist: Diffusion language models (DLMs) offer faster inference capabilities compared to autoregressive (AR) large language models (LLMs), making them suitable for real-time applications.

Key concepts

Diffusion Language Models (DLMs)
These models are proposed because they offer faster inference capabilities compared to autoregressive large language models. They are considered suitable for real-time applications.
Self-distillation
This is a new framework that uses the internal noising process of a DLM itself as the source for learning. It aims to boost generation quality by leveraging the model's own workings instead of needing a separate teacher model.
Easy Mode and Hard Mode
SelFusion introduces two distinct forward passes within the same DLM. The easy mode uses a lower masking ratio, theoretically providing more context for accurate predictions, while the hard mode has a higher masking ratio.

Terminology

Summary

Diffusion language models (DLMs) offer faster inference capabilities compared to autoregressive (AR) large language models (LLMs), making them suitable for real-time applications. However, their generation quality often lags behind AR models, and conventional knowledge distillation (KD) methods applied to DLMs yield only marginal gains or even degrade performance. This paper proposes SelFusion, a novel self-distillation framework that enables effective KD for DLMs without requiring an external teacher model by leveraging the inherent noising process within the DLM itself. The proposed method dynamically determines the optimal distillation direction based on token-level correctness, leading to substantial performance improvements over existing KD baselines and even surpassing teacher models in many configurations.

SelFusion Framework Overview

SelFusion is a self-distillation framework designed to improve DLM generation quality by utilizing the model's noising process to create two distinct forward modes: the 'easy mode' and the 'hard mode'. The key idea is that the input to the easy mode has a lower masking ratio than that of the hard mode, which is expected to yield more accurate predictions. The paper notes that the easy mode assigns approximately 10% higher probability to the correct token than the hard mode throughout training, suggesting it serves as a more accurate teacher.

Bidirectional Distillation Strategy

To overcome potential issues where the easy mode might be overconfident on incorrect tokens, SelFusion introduces bidirectional KD between these two modes. The distillation direction is dynamically determined based on token-level correctness, defined by the following rule:

  1. Both correct: When both modes predict correctly, the mode assigning higher probability to the predicted token serves as the distillation target.

  2. Both wrong: When both modes predict incorrectly, the mode assigning higher probability to the ground-truth token serves as the distillation target.

  3. One correct: When only one mode predicts correctly, that mode serves as the distillation target.

Addressing Overconfidence with RMSNorm

The easy mode's access to more context during generation can lead to overconfident predictions. To mitigate this, SelFusion applies RMSNorm-based logit calibration by inserting it between the final hidden state and the LM head of the easy mode. This selectively suppresses overconfidence: "When both modes were correct, RMSNorm moderately reduced the top-1 probability from approximately around 90% to about 60%. More importantly, when both modes were incorrect, RMSNorm substantially suppressed the overconfident probability from around 40% to about 10%, approximately 75% reduction."

Objective Function and Training

The total training objective for SelFusion jointly optimizes three components: the hard mode diffusion loss, the easy mode diffusion loss, and the token-wise bidirectional distillation loss. The losses are defined as:

  1. Diffusion Losses for Easy and Hard Mode: These are computed over masked positions, reweighted by masking probabilities (Eqs. 3 and 4).

  2. Bidirectional KD Loss: This measures the discrepancy between the two modes' temperature-scaled distributions using Kullback-Leibler (KL) divergence (Eq. 6).

The final objective is defined as: LSelFusion = L(h)diff + L(e)diff + Lbkd (Eq. 7). This combined loss allows both modes to update simultaneously in a single backward pass.

Experimental Results and Performance

Experimental results on instruction-following tasks consistently show that SelFusion substantially outperforms other KD methods with external LLM and DLM teachers. Notably, the proposed method even surpasses the teacher models, including both DLMs and LLMs. For instance, SelFusion achieved scores of 12.87 on Self-inst and 17.08 on Vicuna, exceeding the respective LLM teacher scores of 11.04 and 14.95. Furthermore, the method demonstrates strong generalization to larger diffusion step settings (32 steps), achieving the best average performance across benchmarks in this setting. The training cost analysis confirms that SelFusion eliminates teacher training and thus reduces total training computation compared to KD methods that rely on a separately trained teacher.

Analysis of Mismatch and Generalization

The paper identifies the primary challenge as the distribution mismatch between AR models and the NAR student, specifically concerning top-k logit scale mismatch where LLMs show peaked distributions dominated by the top-1 token, while DLMs exhibit flatter distributions. SelFusion addresses this by dynamically selecting targets based on token correctness. The analysis of masking strategies shows that even when both modes are incorrect, the easy mode consistently assigned higher probability to the correct token than the hard mode, supporting its role as a student-friendly teacher. The framework's effectiveness is further validated across different tasks, with SelFusion consistently outperforming strong baselines in both instruction-following and summarization benchmarks like SAMSum.

Improvements for AI systems

As a fastidious and diligent researcher, I have analyzed SelFusion: Self-distillation for Diffusion Language Models. The proposed framework addresses the critical performance gap between efficient Diffusion Language Models (DLMs) and high-quality autoregressive (AR) models by introducing a novel self-distillation mechanism.

Here are the specific improvements that can be made to AI systems using SelFusion, along with what these improved systems can achieve:


I. Enhanced Generation Quality in Low-Latency LLM/DLM Architectures

The core improvement is the ability to achieve near state-of-the-art generation quality in models optimized for parallel decoding (like DLMs) without relying on slow, external teacher models (like massive AR LLMs).

  1. The system can perform high-quality instruction following and text generation using a student model (e.g., a 472M or 8B parameter DLM backbone) trained entirely via self-distillation from its own outputs across different noise levels within the diffusion process.

  2. This allows for the deployment of faster, real-time inference models that maintain high semantic coherence and factual accuracy comparable to much larger, slower AR models (LLMs).

II. Superior Knowledge Transfer in DLM Architectures (DLM-to-DLM Distillation)

The framework specifically solves the problem where DLMs cannot effectively learn from other DLMs due to inherent generation quality limitations of those teachers.

  1. The system can effectively transfer knowledge between different diffusion models (e.g., distilling a 8-step DLM into a 16-step DLM) by leveraging the bidirectional distillation process, which dynamically selects the most reliable source for each token prediction based on real-time confidence and correctness checks.

  2. This enables efficient fine-tuning and model compression within the diffusion domain without needing to generate slow, high-quality sequence targets from an external AR teacher.

III. Robustness Against Overconfidence in Diffusion Models

The integration of RMSNorm-based logit calibration prevents a common failure mode in distillation: overconfident errors.

  1. The improved system will exhibit significantly reduced catastrophic forgetting or poor generalization when distilling knowledge, as the calibration selectively suppresses high-probability predictions that are incorrect, preventing the student from learning spurious correlations from flawed teacher outputs.

  2. This results in a more stable and reliable transfer of knowledge across all training steps and masking ratios (e.g., 8-step vs. 16-step configurations).

IV. Adaptive Distillation Strategy

The bidirectional KD mechanism allows the model to dynamically switch distillation directions based on token-level performance, which is superior to fixed, one-way distillation methods.

  1. The system can adapt its learning strategy mid-training: if the easy mode (lower masking) produces a highly confident but wrong prediction, it immediately switches its focus to the hard mode prediction for that specific token, ensuring the student learns from the most reliable signal available at that moment.

  2. This adaptive behavior leads to substantially higher accuracy gains (up to 16% relative improvement over baselines) compared to naive logit-level or sequence-level KD methods.

In summary, SelFusion enables the creation of a class of Self-Distilled DLMs that are faster for inference, more accurate in reasoning, and require no external teacher dependency, effectively closing the generation quality gap between efficient diffusion models and large autoregressive models.

Abstract

Diffusion language models (DLMs) alleviate the inherent latency bottleneck of autoregressive (AR) large language models (LLMs), but their degraded generation quality limits practical applicability. Although knowledge distillation (KD) can be a promising direction for improving performance, we empirically find that naively applying conventional KD yields only marginal gains, or even degrades generation quality. Based on these observations, we propose a novel self-distillation framework for DLMs, namely SelFusion. To enable effective KD without an external teacher model, SelFusion performs two forward passes with different masking levels, defining the hard mode with a larger masking probability and the easy mode with a smaller masking probability. However, the easy mode is not always more accurate than the hard mode and can be overconfident on incorrect tokens. Thus, we introduce bidirectional KD between the two modes, which can dynamically determine the distillation direction based on token-level correctness. Experimental results on instruction-following tasks show that the proposed self-distillation substantially outperforms other KD methods with external LLM and DLM teachers. In many configurations, the student trained with SelFusion even surpasses the performance of the LLM teacher, providing a practical path toward improving DLM generation quality. Source code can be found at https://github.com/scai-research/SelFusion official

Sources

Related papers