NoLoCo: No-all-reduce Low Communication Training Method for Large Models
cs.LG
Submitted: 2025-06-12
Updated: 2026-09-08
License: http://creativecommons.org/licenses/by/4.0/
The gist: Training large language models is generally done on clusters containing thousands of accelerators, communicating over a high-bandwidth interconnect.
Terminology
Abstract
Training large language models is generally done on clusters containing thousands of accelerators, communicating over a high-bandwidth interconnect. Scaling up these clusters is expensive and can become impractical, imposing limits on the size of models that can be trained. Several recent studies have proposed training methods that are less communication intensive, avoiding the need for compute clusters with extremely high interconnect speeds. These low communication training methods still employ a global synchronization step for model parameters, which can be too costly with a high number of participants, as the communication cost scales quadratically with group size. In this work, we propose a novel optimization method, NoLoCo, that does not explicitly synchronize all model parameters during training and does not require any collective communication. NoLoCo implicitly synchronizes model weights via a novel variant of the Nesterov momentum optimizer by partially averaging model weights within randomly selected subgroups. We provide both a theoretical convergence analysis of our optimizer and empirical results from language model training. Our method requires significantly less communication than fully sharded data parallel training and DiLoCo, a widely used low-communication baseline. Moreover, our method avoids global blocking communication, thereby reducing accelerator idle time. Our experiments show that NoLoCo is more communication-efficient than DiLoCo, improving final perplexity by up to 4% and converging up to 4 times faster in wall-clock time across a range of worker counts, model sizes, and communication bandwidths.
Sources
- Boosting Asynchronous Decentralized Learning with Model Fragmentation
- A Survey on Mixture of Experts in Large Language Models
- Communication-Efficient Language Model Training Scales Reliably and Robustly: Scaling Laws for DiLoCo
- Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus
- DiLoCo: Distributed Low-Communication Training of Language Models
- DiPaCo: Distributed Path Composition
- Efficient Training of Large Language Models on Distributed Infrastructures: A Survey
- Accelerating Large Language Model Training with 4D Parallelism and Memory Consumption Estimator
- The Llama 3 Herd of Models
- INTELLECT-1 Technical Report
- Eager Updates For Overlapped Communication and Computation in DiLoCo
- Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models
- Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- Ring Attention with Blockwise Transformers for Near-Infinite Context
- DeMo: Decoupled Momentum Optimization
- AudioPaLM: A Large Language Model That Can Speak and Listen
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
- Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters
- Seq1F1B: Efficient Sequence-Level Pipeline Parallelism for Large Language Model Training
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks