EMA: Elastic and Performance Transparent Memory Across GPUs
cs.DC, cs.AI, cs.LG
Submitted: 2026-09-22
Updated: 2026-09-22
Code: https://github.com/NVIDIA/FasterTransformer
Terminology
Sources
- Training Deep Nets with Sublinear Memory Cost
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
- Harvest: Opportunistic Peer-to-Peer GPU Caching for LLM Inference
- M'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity
- Optimal checkpointing for heterogeneous chains: how to train deep neural networks with limited memory
- Evaluating Modern GPU Interconnect: PCIe, NVLink, NV-SLI, NVSwitch and GPUDirect
- Mixed Precision Training
- Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs
- TensorFlow-Serving: Flexible, High-Performance ML Serving
- LLaMA: Open and Efficient Foundation Language Models
- LightSeq: A High Performance Inference Library for Transformers
- ProTrain: Efficient LLM Training via Memory-Aware Techniques
- OPT: Open Pre-trained Transformer Language Models
Related papers
- iScheduler: Reinforcement Learning-Driven Continual Optimization for Large-Scale Resource Investment Problems
- SAMM: Sharded Automated Market Maker
- InferScale: GPU-Native KV Injection for Personalized LLM Serving
- Vigil: Accountable Liveness against Selective Silence
- Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAG
- Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing