OctoPipe: Reducing Pipeline Bubbles for Heterogeneous Models via Co-Optimizing Partitioning, Placement, and Scheduling
cs.DC, cs.AI
Submitted: 2025-09-28
Updated: 2026-09-02
Code: https://github.com/NVIDIA/Megatron-LM
Terminology
Sources
- The Llama 3 Herd of Models
- Kimi K2: Open Agentic Intelligence
- gpt-oss-120b & gpt-oss-20b Model Card
- Qwen3 Technical Report
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model
- MiniMax-01: Scaling Foundation Models with Lightning Attention
- Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
- Horovod: fast and easy distributed deep learning in TensorFlow
- DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models
- Training Deep Nets with Sublinear Memory Cost
- SPPO:Efficient Long-sequence LLM Training via Adaptive Sequence Pipeline Parallel Offloading
Related papers
- iScheduler: Reinforcement Learning-Driven Continual Optimization for Large-Scale Resource Investment Problems
- SAMM: Sharded Automated Market Maker
- InferScale: GPU-Native KV Injection for Personalized LLM Serving
- Vigil: Accountable Liveness against Selective Silence
- Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAG
- Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing