Looped Diffusion Transformer

arXiv:2609.40305 · cs.CV, cs.LG · Submitted 2026-09-30 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Looped Diffusion Transformer".

Jane: Looped Diffusion Transformer explores an alternative way to scale computation in text-to-image models by repeatedly running shared Transformer blocks within each denoising step,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: Let's start with the title and the folks behind it; "Looped Diffusion Transformer" is quite descriptive of what they’re doing, suggesting a focus on looping computation within diffusion models.

Jane: The authors are a team from SenseTime Research and LeapLab at Tsinghua University, which tells us we’re looking at some top-tier research coming out of the AI community.

Lu: I find the combination of diffusion methods with looped computation really intriguing because it touches on how sequential generation processes can be structured for better quality.

Meng: I'm curious if their specific architecture choice, using MMDiT as a base, is what makes this looping strategy viable compared to other architectures.

Lalam: The paper introduces Looped Diffusion Transformer Yong Xien Chng and Tianyi Chen and their team, which shows the collaborative nature of this kind of deep research in the AI space.

The paper's summary: Tom: Now that we know the basics, let's look at what they actually achieved in terms of summarizing their main contribution with "Looped Diffusion Transformer."

Jane: The core idea is that instead of just running one set of blocks, they repeat shared Transformer blocks inside each denoising step, which increases computational depth without changing the number of parameters.

Lu: This looped computation enables iterative refinement of internal representations, and the authors suggest this iterative process can support latent visual reasoning through progressive correction across those loops.

Meng: So instead of just one pass to get an image, you’re essentially running a mini-denoising process inside each main denoising step, which is a big structural change for how we think about these models.

Lalam: It means the model gets multiple chances to correct its internal state before moving on to the next major denoising stage, which should lead to more robust outputs in theory.

The paper's improvements: Tom: The paper points out that naive looping doesn't actually improve image quality consistently, because they found issues with weak supervision and attention updates eroding local information.

Jane: They address this by proposing Looped-DiT, which incorporates deep supervision across intermediate loops combined with self-modulating attention to stabilize those featurization issues.

Lu: The deep supervision part involves decoding each intermediate loop output through the shared post-loop blocks to supervise all predictions against the same clean-image target, which is a smart way to guide learning.

Meng: And then there’s the self-modulating attention, with things like Gated Attention or Exclusive Self Attention designed to regulate attention updates based on current hidden states.

Lalam: That self-modulating part is crucial because it stops the repeated application of those shared blocks from causing redundant updates that degrade local information across the loops.

Conclusion: Tom: So, to wrap up this discussion on "Looped Diffusion Transformer," the main implication is that looping computation offers a way to scale visual generation by increasing depth without needing more parameters.

Jane: The paper shows that combining deep supervision and self-modulating attention tackles the problems of saturation and spatial information loss seen in naive looping, leading to better image quality.

Lu: What excites me is how this iterative refinement suggests a path toward latent visual reasoning, where the model corrects errors progressively across the loops without needing explicit reasoning tokens.

Meng: Practically speaking, if we can achieve six point five times larger performance with lower compute by using this looping method, that really changes our inference budget planning for large models.

Lalam: For me, the fact that we can get superior robustness and efficiency through these mechanisms means this paper could significantly improve the general capabilities of text-to-image systems across various cultural contexts.

Yong Xien Chng, Tianyi Chen, Wenwen Tong, Haiwen Diao, Zhongang Cai, Lei Yang, Ziwei Liu

SenseTime Research 2LeapLab, Tsinghua University 3Nanyang Technological University

cs.CV, cs.LG

Submitted: 2026-09-30

Updated: 2026-09-30

Comments: 21 pages, 9 figures

Code: https://github.com/OpenSenseNova/Looped-DiT

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 85/100

The gist: Looped Diffusion Transformer explores an alternative way to scale computation in text-to-image models by repeatedly running shared Transformer blocks within each denoising step, effectively

Key concepts

Looped Diffusion Transformer (Looped-DiT)
This architecture uses a specific sequence of stages—pre-loop A, looped stage B, and post-loop C—where the parameters of the middle stage B are reused across multiple iterations (the loop depth). This allows computation to be deepened by repeating this shared block structure many times.
Deep Supervision
Instead of only checking the final output after all loops, this technique applies a loss function at every intermediate loop depth. By supervising predictions at each stage, the model is encouraged to maintain high quality and correct errors progressively throughout the entire generation process.
Self-Modulating Attention (SMA)
This mechanism regulates how attention updates occur across different loops to prevent redundant information. It uses methods like Gated Attention or Exclusive Self Attention (XSA) to explicitly control the magnitude or direction of attention contributions, ensuring each loop iteration makes meaningful, non-redundant changes.
Latent Visual Reasoning
The authors observe that deeper looping stages cause the model to progressively correct earlier mistakes. This behavior suggests that the iterative process mimics a form of latent visual reasoning, where the model resolves complex visual constraints by refining its internal representations through successive error correction.

Terminology

Summary

Looped Diffusion Transformer explores an alternative way to scale computation in text-to-image models by repeatedly running shared Transformer blocks within each denoising step, effectively increasing computational depth without increasing parameter count. This approach enables iterative refinement of internal representations, which the authors find can support latent visual reasoning through progressive correction of errors across loops.

The gist

Looped computation offers a promising way to scale visual generation models by repeatedly applying shared Transformer blocks within each denoising step to increase computational depth without increasing parameter count.

Core Methodology and Architecture

The Looped Diffusion Transformer (Looped-DiT) is built on MiniT2I, a pixel-space denoiser based on MMDiT. The architecture partitions the 17 MMDiT blocks into three sequential groups: a pre-loop stage A, a looped stage B, and a post-loop stage C. Given an input image and text condition, the forward computation is defined as:

h(0) = A(hinput)

h(r) = B(h(r−1)), r = 1,..., N

xˆ0 = C(h(N))

Here, N denotes the loop depth. The critical aspect is that the parameters of the middle block group (B) are shared across iterations, meaning increasing N increases effective computational depth without increasing parameter count. This allows loop depth to be varied independently of the number of denoising steps.

Addressing Naive Looping Limitations

Initial experiments showed that naive looping fails to consistently improve image quality. The authors trace this problem to weak supervision across intermediate loops and unregulated attention updates that progressively erode local information. Specifically, they found that performance can saturate and eventually decline beyond the training loop depth. This degradation coincided with a steady loss of linearly decodable spatial information, as measured by a ridge-regression probe where the R2 value dropped from 0.865 after the 1st loop to 0.562 after 8 loops.

Proposed Solutions: Deep Supervision and Self-Modulating Attention

To overcome these challenges, Looped-DiT incorporates two key mechanisms:

  1. Deep Supervision: This applies the flow-matching objective to predictions at every loop depth rather than relying solely on the final state. For each loop depth n = 1,..., N, the hidden state is decoded through the shared post-loop stage C to create a prediction xˆ(n)0. The overall training objective combines supervision across loop depths: L = X Σ n=1 to N wnln, where wn controls the supervision strength at loop n.

  2. Self-Modulating Attention (SMA): This regulates attention updates across loops based on current hidden states to prevent redundancy. Two realizations are studied:

(a) Gated Attention:

The modulation is a token-dependent scalar gate: Ggate i,h = σ w⊤ g,hui + bg,h. The output is then modulated as z gate i,h = Ggate i,h oi,h. This explicitly controls the magnitude of each head contribution.

(b) Exclusive Self Attention (XSA):

This provides a parameter-free realization where the modulation Gxsa i,h projects the attention output onto the subspace orthogonal to the token’s own value direction: Gxsa i,h = I − vˆi,hvˆ⊤ i,h. This eliminates the self-value term from the update.

Key Findings and Performance Gains

The introduction of Looped-DiT consistently outperforms non-looped baselines under matched-parameter and matched-compute settings. Notably:

  1. A 260M-parameter Looped-DiT model can surpass a model 6.5× larger across multiple text-to-image benchmarks while requiring 4.9× lower inference compute.

  2. Loop depth acts as a complementary axis for inference scaling, where allocating additional computation to increasing loop depth yields greater gains than allocating it to more denoising steps.

  3. Successive loops exhibit behaviors suggestive of latent visual reasoning, with deeper loops progressively correcting earlier errors and resolving interdependent constraints without explicit textual reasoning traces.

Ablation Study Results

The study systematically ablated components to confirm the necessity of the proposed mechanisms:

(a) Looping alone:

Looping alone improves over the baseline model.

(b) Deep Supervision:

Both weighting schemes (Exponential and Final + Mean) outperform final-loop-only supervision, with Final + Mean performing better.

(c) Self-Modulating Attention:

Both Gated Attention and XSA improve over unmodulated attention, with XSA providing larger gains. Combining Final + Mean supervision with XSA yields the highest average score.

Improvements for AI systems

As a fastidious researcher, I have analyzed the Looped Diffusion Transformer (Looped-DiT) paper. The core innovation is demonstrating that repeatedly running shared Transformer blocks within each denoising step—looped computation—can scale visual generation models efficiently by increasing computational depth without increasing parameter count or sequence length.

Here are the specific improvements and capabilities you can gain by implementing Looped-DiT in your AI systems:


)

)

Improve text-to-image (T2I) generation efficiency and quality through iterative refinement. Looped-DiT allows models to perform latent reasoning by iteratively correcting errors within the hidden representations during the denoising process, rather than relying solely on a single pass.

Specific Capabilities:

  1. textbfIncreased Generation Quality via Error Correction (Latent Reasoning):

  2. The system can resolve complex, interdependent visual constraints that are difficult for standard models to handle in one step. As shown in Figure 5, successive loops allow the model to:

  3. Correct rendering errors introduced in earlier steps;

  4. Reorganize objects to satisfy spatial constraints (e.g., ensuring two toy elephants are correctly positioned relative to four white balloons);

  5. Remove extraneous or incorrect elements introduced by a previous loop iteration, leading to visually more coherent and accurate final images compared to non-looped baselines.

  6. textbfSuperior Inference Efficiency at Matched Compute Budgets (Scalability):

  7. The system can achieve performance comparable to significantly larger models (e.g., InternVL-U) while requiring substantially lower inference compute (up to 6.5× fewer FLOPs, as per Table 1).

  8. It offers a more effective use of inference budget than simply adding more denoising steps; allocating compute to increasing loop depth yields higher performance gains under fixed budgets (as shown in Figure 6).

  9. textbf Tunable Inference Strategy (Adaptive Looping):

  10. The system can implement Adaptive Looping, where a lightweight gating network decides whether to continue looping after each step based on predicted future loss reduction (Equation 7). This allows the model to dynamically adjust its computational effort per image, stopping early when prompt requirements are met, thereby optimizing the trade-off between quality and speed.

  11. textbf Enhanced Robustness via Self-Modulation:

  12. The integration of Self-Modulating Attention (specifically Exclusive Self Attention or Gated Attention) prevents the repeated application of shared blocks from causing redundant updates that erode local information. This ensures that attention updates are regulated based on current hidden states, preserving critical token information and preventing overwriting of correct representations across loops.

  13. textbf Enhanced Training Stability via Deep Supervision:

  14. During training, Deep Supervision applies the flow-matching objective to predictions at every intermediate loop depth (1 through 3). This provides a direct learning signal to earlier hidden states, preventing the optimization path from becoming excessively long and ensuring that all intermediate representations are learned effectively against the same clean target image.

In summary, implementing Looped-DiT enables AI systems to be simultaneously more robust (better reasoning), more efficient (lower compute/latency), and more controllable (adaptive looping) in text-to-image generation tasks.

Abstract

Improving text-to-image models has traditionally relied on increasing model size or the number of denoising steps. In this work, we explore an alternative way to scale computation by repeatedly running shared Transformer blocks within each denoising step, effectively increasing computational depth while keeping the parameter count fixed. This looped computation enables iterative refinement of internal representations without explicit reasoning tokens. However, naive looping fails to consistently improve image quality. We trace this problem to weak supervision across intermediate loops and unregulated attention updates that progressively erode local information. To overcome these challenges, we propose Looped Diffusion Transformer (Looped-DiT), which combines deep supervision across intermediate loops with self-modulating attention to stabilize looped feature updates. Under matched-parameter and matched-compute settings, Looped-DiT consistently outperforms non-looped baselines. Notably, a 260M-parameter looped model can surpass a model 6.5x larger across multiple text-to-image benchmarks while requiring 4.9x lower inference compute. Beyond this performance gain, we find that looped computation can offer a more effective form of iterative computation for diffusion models, with increasing loop depth yielding larger gains than adding more denoising steps under a fixed inference budget. Furthermore, deeper loops can progressively correct mistakes made in earlier loops, exhibiting behaviors suggestive of latent reasoning. Together, these results show that looped computation offers a promising way to scale visual generation models.

Sources

Related papers