Accelerating Sharded Data Parallelism at Scale with Federated Learning
cs.DC, cs.AI, cs.PF
Submitted: 2026-09-17
Updated: 2026-09-17
Journal ref: Euro-Par 2026: Parallel Processing - 32nd European Conference on Parallel and Distributed Processing, Pisa, Italy, August 24-28, 2026, Proceedings, Part II
DOI: 10.1007/978-3-032-35251-4_30
Code: https://github.com/NVIDIA/nccl-tests
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- DiLoCo: Distributed Low-Communication Training of Language Models
- Low-Precision Training of Large Language Models: Methods, Challenges, and Opportunities
- Faster On-Device Training Using New Federated Momentum Algorithm
- The Llama 3 Herd of Models
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
- OPT: Open Pre-trained Transformer Language Models
Related papers
- iScheduler: Reinforcement Learning-Driven Continual Optimization for Large-Scale Resource Investment Problems
- SAMM: Sharded Automated Market Maker
- InferScale: GPU-Native KV Injection for Personalized LLM Serving
- Vigil: Accountable Liveness against Selective Silence
- Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAG
- Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing