Block Parallelism For Efficient Distributed Long-Context Diffusion Language Model Training
cs.LG
Submitted: 2026-09-16
Updated: 2026-09-16
Comments: 25 pages, 6 figures
Code: https://github.com/ScalingIntelligence/Turbo-dLLM
License: http://creativecommons.org/licenses/by/4.0/
The gist: Block diffusion language models (BDLMs) combine autoregressive dependencies across blocks with parallel denoising within blocks, but long-context training is constrained by distributed attention
Terminology
Abstract
Block diffusion language models (BDLMs) combine autoregressive dependencies across blocks with parallel denoising within blocks, but long-context training is constrained by distributed attention communication and activation memory. Conventional context parallelism (CP) shards the combined clean-plus-corrupted sequence by position, communicating shared clean K/V together with block-specific corrupted K/V and their gradients. We observe that the BDLM objective separates over target blocks. We introduce block parallelism (BP), a new distributed parallelism dimension that assigns each corrupted-block computation to one rank. To scale BP to long contexts, we introduce context-sharded block parallelism (CSBP), which also shards the shared clean sequence across those ranks. CSBP keeps corrupted K/V and gradients local, avoids replicated clean prefixes, and preserves BDLM training semantics. On 16 H200 GPUs at 256K context, CSBP improves throughput over the best baseline by 1.18-1.45x for supervised fine-tuning and 1.27-1.33x for conversion of autoregressive models to BDLMs, while matching or reducing peak HBM. Full-model speedup reaches 1.61x at 512K. On eight H100 GPUs, CSBP accelerates DFlash2 speculative-decoder training by 2.48x at 512K and 7.59x at 1M. In matched 12-hour DiffusionGemma 26B-A4B SFT runs, CSBP achieves higher pass rates at every trained checkpoint on SWE-bench Verified and Terminal-Bench Lite. Code: https://github.com/ScalingIntelligence/Turbo-dLLM
Sources
- Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models
- Striped Attention: Faster Ring Attention for Causal Transformers
- DFlash: Block Diffusion for Flash Speculative Decoding
- Training Deep Nets with Sublinear Memory Cost
- PaLM: Scaling Language Modeling with Pathways
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
- DiffusionGemma Technical Report
- The Llama 3 Herd of Models
- Nemotron-Labs-Diffusion: A Tri-Mode Language Model Unifying Autoregressive, Diffusion, and Self-Speculation Decoding
- LoongTrain: Efficient Training of Long-Sequence LLMs with Head-Context Parallelism
- DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models
- Reducing Activation Recomputation in Large Transformer Models
- GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
- SpecForge: A Flexible and Efficient Open-Source Training Framework for Speculative Decoding
- Ring Attention with Blockwise Transformers for Near-Infinite Context
- Online normalizer calculation for softmax
- LiteCoder-Terminal: Scaling Long-Horizon Terminal Environments for Learning Language Agents
- ZeRO: Memory Optimizations Toward Training Trillion Parameter Models
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
- Attention Is All You Need
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks