Concertina: Data-Centric Adaptive Pipeline Parallelism for Efficient Heterogeneous Long-Context LLM Training
cs.DC, cs.AI
Submitted: 2025-09-25
Updated: 2026-09-14
Code: https://github.com/wsjdsg/InfiniPipe-code
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- ByteScale: Efficient Scaling of LLM Training with a 2048K Context Length on More Than 12,000 GPUs
- Optimal checkpointing for heterogeneous chains: how to train deep neural networks with limited memory
- DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models
- Reducing Activation Recomputation in Large Transformer Models
- Efficient Sequence Packing without Cross-contamination: Accelerating Large Language Models without Impacting Performance
- MiniMax-01: Scaling Foundation Models with Lightning Attention
- DISTFLASHATTN: Distributed Memory-efficient Attention for Long-context LLMs Training
- DeepSeek-V3 Technical Report
- Ring Attention with Blockwise Transformers for Near-Infinite Context
- GPT-4 Technical Report
- Striped Attention: Faster Ring Attention for Causal Transformers
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- The Llama 3 Herd of Models
- Zero Bubble Pipeline Parallelism
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
- Seq1F1B: Efficient Sequence-Level Pipeline Parallelism for Large Language Model Training
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- WLB-LLM: Workload-Balanced 4D Parallelism for Large Language Model Training
- Qwen3 Technical Report
Related papers
- iScheduler: Reinforcement Learning-Driven Continual Optimization for Large-Scale Resource Investment Problems
- SAMM: Sharded Automated Market Maker
- InferScale: GPU-Native KV Injection for Personalized LLM Serving
- Vigil: Accountable Liveness against Selective Silence
- Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAG
- Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing