Training Variable Long Sequences with Data-Centric Parallel
Geng Zhang, Xuanlei Zhao, Kai Wang, Yang You
National University of Singapore
cs.AI, cs.LG
Submitted: 2026-07-14
Updated: 2026-08-11
License: http://creativecommons.org/licenses/by/4.0/
Terminology
Summary
Summary
This paper introduces Data-Centric Parallel (DCP), a framework designed to address the computational challenges of training deep learning models on variable long sequences. The authors argue that existing methods force a difficult trade-off between efficiency and ease-of-use. Simple approaches, such as bucket parallel, use static configurations that cause workload imbalance and low efficiency, while complex methods, such as compiler-based parallel approaches, introduce significant complexity and require substantial code changes for new models. To break this trade-off, DCP's core principle is to let the data itself drive the runtime
by dynamically adjusting direct runtime settings—including parallel size, gradient accumulation, and recomputation—based on each batch's sequence length.
The paper identifies three primary challenges in training variable long sequences. First, there is large variance in data length in real-world datasets, such as Panda-70M, which contains 70 million videos. A naive solution of grouping data of similar lengths harms model quality by disrupting independent and identically distributed sampling. Second, communication cost for sequence parallel is a major issue. Large sequence parallel (SP) sizes are necessary for the longest sequences to avoid out-of-memory errors, but this large SP becomes a major bottleneck for shorter sequences, which often dominate datasets, because the fixed communication cost remains high relative to the reduced computation. Third, workload imbalance arises in data parallelism because workers assigned batches of long sequences take significantly longer, forcing workers with short sequences to sit idle waiting for synchronization. The paper also notes a trade-off between balance and communication: statically reducing batch sizes for long sequences to balance workload leads to under-utilization of hardware and a massive increase in total communication costs due to more training steps.
DCP comprises two strategies: DCP-inter and DCP-intra. DCP-inter optimizes load balance by using gradient accumulation instead of reducing batch size. It first determines the best parallel settings for each sequence length based on profiling, then exploits gradient accumulation to balance execution time across data batches in each iteration, maintaining high throughput for slow batches and filling idle time of fast batches by running multiple batches. DCP-intra further exploits dynamic recomputation for better speed. It is motivated by the analysis that for short sequences, gradient checkpointing introduces considerable unnecessary computational overhead. DCP-intra strategically deactivates gradient checkpointing for shorter sequences and manages the increase in memory consumption by dynamically adjusting sequence parallel and batch size with ignorable cost. The paper formalizes the problem as two optimization targets: maximizing total throughput (sum of batch size times sequence length divided by execution time) and minimizing workload imbalance across batches.
The methods rely on a dual-layer profiling process. Based on the observation that transformer models often consist of multiple repeating layers, the profiling process only runs two layers for each batch size and sequence parallel size. The first layer is used for warm-up, and the forward execution time, backward execution time, and memory overhead for the second layer are recorded to estimate these metrics for the whole model. This profiling is fast and enables the runtime to select optimal configurations.
Empirical results demonstrate that DCP achieves up to a 2.88× speedup on 32 H200 GPUs. The evaluation was conducted on two typical transformer architectures: Transformer-1D (5B parameters) and Transformer-2D (1.2B parameters), across three synthesized datasets (short-sequence-dominated, balanced, and long-sequence-dominated). Key findings include: DCP-inter consistently outperforms the baseline (bucket parallel), delivering speedups of up to 2.70× and 1.68× for the two models; DCP-intra provides additional performance gains over DCP-inter, achieving a total speedup of up to 2.88×; improvements are most pronounced on datasets dominated by shorter sequences. The methods also significantly reduce workload imbalance, maintaining a consistently low imbalance ratio across all conditions. In scaling experiments, DCP demonstrates substantial scalability improvements, achieving near-linear scaling on both models, while the baseline exhibits significant throughput decay as it scales.
The paper also includes an ablation study showing that DCP-intra achieves 20-25% speedup over DCP-inter for sequence lengths less than 200k, directly corresponding to the overhead of gradient checkpointing that the method eliminates. Designed for generalization, DCP can be integrated into any model with at most 10 lines of code change. The paper notes two limitations: the method is restricted to Transformer-based models, and it is designed for a single model architecture, not systems of multiple distinct networks. Future work could enhance the method by developing predictive models for proactive selection of dynamic runtime parameters, exploring finer-grained intra-batch parallelism, and broadening the scope to other architectures.
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in an AI training system:
Current limitation: Fixed sequence-parallel (SP) size across all batches causes severe communication overhead for short sequences and OOM errors for long ones.
Improvement: Implement a runtime scheduler that profiles each incoming batch's sequence length and dynamically selects the optimal SP size (1, 2, 4, 8, 16, or 32) before execution. This eliminates the 40-60% communication overhead currently wasted on short sequences that don't need large SP.
Resulting capability: The system can train on datasets with sequence lengths ranging from 1K to 1M tokens without manual configuration, achieving up to 2.88× throughput improvement on 32 GPUs.
Current limitation: Bucket parallel reduces batch size for long sequences to balance workload, which underutilizes GPU memory and increases total communication steps.
Current limitation: Full activation checkpointing is applied uniformly, wasting 20-25% computation on short sequences where memory is not constrained.
Current limitation: Full-model profiling takes minutes and must be repeated for each configuration change.
Current limitation: No real-time monitoring of GPU utilization variance across workers.
-
Train on any sequence length distribution (short-dominated, balanced, or long-dominated) with minimal manual tuning
-
Achieve near-linear scaling across GPU clusters while maintaining >95% GPU utilization
-
Support both 1D and 2D attention architectures (standard transformers and video/protein models)
-
Adapt to changing data distributions mid-training without restarting or recompiling
-
Integrate with existing frameworks (DeepSpeed, PyTorch FSDP) with only 10 lines of code changes
The system is particularly effective for video generation, protein structure prediction, and long-context language models where sequence lengths vary dramatically within a single training run.
Abstract
Training deep learning models on variable long sequences poses significant computational challenges. Existing methods force a difficult trade-off between efficiency and ease-of-use. Simple approaches use static configurations that cause workload imbalance low efficiency, while complex methods introduces significant complexity and code change for new models. To break this trade-off, we introduce Data-Centric Parallel (DCP). Its core principle is to let the data itself drive the runtime. It achieves this by dynamically adjusting direct runtime settings (e.g., parallel size, gradient accumulation, recomputation) based on each batch's sequence length. Empirical results demonstrate that our method achieves up to a 2.88 times speedup on 32 H200 GPUs. Designed for generalization, it can be integrated into any model with 10 lines of code. We anticipate this simple yet effective approach will serve as a robust baseline and facilitate future advancements in distributed training for variable long sequences.
Sources
- Qwen Technical Report
- Training Deep Nets with Sublinear Memory Cost
- The Llama 3 Herd of Models
- USP: A Unified Sequence Parallelism Approach for Long Context Generative AI
- HunyuanVideo: A Systematic Framework For Large Video Generative Models
- CAME: Confidence-guided Adaptive Memory Efficient Optimization
- Movie Gen: A Cast of Media Foundation Models
- LLaMA: Open and Efficient Foundation Language Models
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
- Open-Sora: Democratizing Efficient Video Production for All
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection