Distributed Training using an Intelligent Network
cs.LG, cs.DC, cs.NI
Submitted: 2026-08-26
Updated: 2026-08-26
License: http://creativecommons.org/licenses/by/4.0/
The gist: Distributed training across a wide area network (WAN) is challenging, as continuous parameter exchange by islands of compute is constrained by limited bandwidth, high latency, and uneven topology.
Terminology
Abstract
Distributed training across a wide area network (WAN) is challenging, as continuous parameter exchange by islands of compute is constrained by limited bandwidth, high latency, and uneven topology. We propose making the network an active participant in training. On the systems side, such networks should leverage (i) multicast technology to replicate outbound traffic and (ii) in-line FPGAs to aggregate inbound traffic, to ease egress and ingress bottlenecks. These technologies are used for training across workers within a data center, but this paper extends them to the WAN. On the algorithms side, we develop an optimization framework that produces rich synchronization schedules (namely, rotating cliques of islands) around the underlying network topology and these technologies, to maximize information exchange. Finally, we illustrate this on a nine-city topology modeled on the DoubleZero network, a live programmable WAN equipped with both technologies, and show how the optimal schedules shift with the network's capabilities. Together, these can narrow the gap to the gold standard of colocated training.
Sources
- DiLoCo: Distributed Low-Communication Training of Language Models
- Streaming DiLoCo with overlapping communication: Towards a Distributed Free Lunch
- Gemini: A Family of Highly Capable Multimodal Models
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
- In-Network Collective Operations: Game Changer or Challenge for AI Workloads?
- INTELLECT-1 Technical Report
- Prime Collective Communications Library -- Technical Report
- Factored Gossip DiLoCo: Reducing Blocking Communication in DiLoCo
- Asynchronous Local-SGD Training for Language Modeling
- Bandwidth-Aware Network Topology Optimization for Decentralized Learning
- MATCHA: Speeding Up Decentralized SGD via Matching Decomposition Sampling
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks