The Life of a Token: from Words to Bits on the Wire
cs.DC, cs.LG, cs.NI, cs.PF
Submitted: 2026-09-17
Updated: 2026-09-17
Code: https://github.com/NVIDIA/nccl
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Scaling Laws for Neural Language Models
- Scaling Intelligence: Designing Data Centers for Next-Gen Language Models
- DeepSeek-V3 Technical Report
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Qwen2.5-Coder Technical Report
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- PaLM 2 Technical Report
- GPT-4 Technical Report
- Communication-Efficient Large-Scale Distributed Deep Learning: A Comprehensive Survey
- A Survey of Large Language Models
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
- PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel
- PipeDream: Fast and Efficient Pipeline Parallel DNN Training
- GSPMD: General and Scalable Parallelization for ML Computation Graphs
- BLOOM: A 176B-Parameter Open-Access Multilingual Language Model
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
- The Llama 3 Herd of Models
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
Related papers
- iScheduler: Reinforcement Learning-Driven Continual Optimization for Large-Scale Resource Investment Problems
- SAMM: Sharded Automated Market Maker
- InferScale: GPU-Native KV Injection for Personalized LLM Serving
- Vigil: Accountable Liveness against Selective Silence
- Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAG
- Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing