Reforge: Low-Latency Distributed GNN Serving with Selective Embedding Recomputation
cs.DC, cs.LG
Submitted: 2025-01-15
Updated: 2026-09-21
Comments: Extended version of the IPDPS'26 paper (https://doi.org/10.1109/IPDPS65963.2026.00071)
Code: https://github.com/NVIDIA/nccl
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- GPT-4 Technical Report
- How Attentive are Graph Attention Networks?
- Representation Learning on Graphs: Methods and Applications
- Training Compute-Optimal Large Language Models
- Adam: A Method for Stochastic Optimization
- DeeperGCN: All You Need to Train Deeper GCNs
- Quiver: Supporting GPUs for Low-Latency, High-Throughput GNN Serving with Workload Awareness
- Deep Graph Library: A Graph-Centric, Highly-Performant Package for Graph Neural Networks
Related papers
- iScheduler: Reinforcement Learning-Driven Continual Optimization for Large-Scale Resource Investment Problems
- SAMM: Sharded Automated Market Maker
- InferScale: GPU-Native KV Injection for Personalized LLM Serving
- Vigil: Accountable Liveness against Selective Silence
- Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAG
- Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing