LazyTrain: Limited-resource Allocation toward Zero-waste Yield Optimization in Large Language Model Training
Xiaojun Wu, Cehao Yang, Honghao Liu, Xueyuan Lin, Xuhui Jiang, Chengjin Xu, Jia Li, Jian Guo
IDEA Research · The Hong Kong University of Science and Technology (Guangzhou) · DataArcTech Ltd.
cs.CL
Submitted: 2026-08-12
Updated: 2026-08-13
Comments: 18 pages, 8 figures
Code: https://github.com/DataArcTech/LazyTrain
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: LazyTrain is an optimization-guided scheduler for limited-resource large language model (LLM) training.
Terminology
Summary
LazyTrain is an optimization-guided scheduler for limited-resource large language model (LLM) training. It is built as an optimization layer over a MegaTrain-style layer-streaming executor. The core idea is to formulate limited-resource training as a constrained scheduling problem, where the bottleneck is not just GPU memory capacity but also the competition for GPU compute windows, CPU-GPU traffic, mandatory parameter and gradient movement, and possible NVMe activation spill.
LazyTrain replaces fixed activation scheduling with a mixed-integer programming (MIP) model. The decision variables select checkpoint paths, activation homes (GPU HBM, CPU DRAM, or local NVMe), recomputation blocks, and communication assignments. The objective minimizes recomputation plus the incremental communication exposure induced by activation placement, rather than maximizing the amount of offload. The solved policy is loaded once before training and then executed by the runtime.
The method is formalized as a mixed-integer linear program (MILP). For a transformer with L layers and layer-boundary activations, a directed edge (s, e) represents one backward recomputation block. Binary variables select recomputation edges and materialized boundaries, while other binary variables assign each materialized boundary to GPU, CPU, or NVMe. Constraints enforce a legal path from input to output, unique placement, tier capacities, and PCIe/NVMe bandwidth windows. The objective minimizes recomputation cost plus activation-induced exposed communication, with small tie-breaking penalties.
LazyTrain also includes a coupled Hybrid 8-bit operator: 8-bit optimizer states reduce CPU-resident optimizer memory, while fast gradient clipping counteracts the additional CPU-side update overhead. The two mechanisms are deliberately coupled and ablated as one formal component.
Experiments were conducted on single-GPU H800 80GB and RTX 3090 24GB accelerators, from 3B to 27B models. The primary matched comparison uses Qwen3.6-27B on MetaMathQA with a 70/30 split, sequence length 1024, batch size 72, and one epoch. In this primary run, LazyTrain improves sustained throughput from 176.90 to 219.95 TFLOPS, a 1.24× increase, while using the same model, data, sequence length, and batch size. It reaches 1361 tokens/s, peaks at 68.84 GB of GPU memory, and obtains 95.42% exact-match accuracy on the full evaluation split.
The component ablation on the 27B model shows that removing MILP-based scheduling drops performance to 193.17 TFLOPS and 1195 tokens/s, reductions of 12.2% in both metrics, establishing MILP scheduling as the central contributor. Removing the Hybrid 8-bit operator yields 219.29 TFLOPS and 1357 tokens/s, only 0.3% below the complete system.
On RTX 3090, at every model scale, LazyTrain increases the maximum feasible batch size by one over MegaTrain and yields higher token throughput. For example, at 3B, LazyTrain reaches batch size 8 vs. 7 for MegaTrain; at 27B, batch size 2 vs. 1.
The final Qwen3.6-27B schedule uses 52 checkpoints: 30 on GPU, 21 on CPU, 1 on NVMe, with 13 recomputed boundaries. The solver reports 0.0 ms exposed PCIe/NVMe communication. A separate case study with a tighter 8 GB CPU activation budget shows 30 GPU, 11 CPU, 11 NVMe checkpoints and 13 recomputed boundaries.
The paper concludes that LazyTrain's mixed-integer scheduler is its primary method innovation, jointly optimizing activation checkpointing, tier placement, recomputation, and CPU-GPU-NVMe communication overlap. It improves sustained TFLOPS over MegaTrain by approximately 1.24× across H800 models, and on RTX 3090 achieves higher throughput and a one-unit larger feasible batch size. The limitations include single-GPU-only experiments, the evaluation split also being used for periodic loss monitoring, solver outputs not being separately instrumented for per-step stall time, and the scheduler being offline without runtime bandwidth adaptation.
Improvements for AI systems
Improvements to AI systems:
-
Optimization-guided resource scheduling: Replace heuristic or static memory management in LLM training with a mixed-integer linear programming (MILP) scheduler that jointly optimizes activation checkpointing, tier placement (GPU/CPU/NVMe), recomputation blocks, and communication overlap. This yields a 1.24× sustained throughput increase (176.90 → 219.95 TFLOPS) on a 27B model without changing model, data, or batch size.
-
Communication-exposure-aware objective: Instead of maximizing offload, minimize the sum of recomputation cost and incremental communication exposure induced by activation placement. This eliminates exposed PCIe/NVMe stalls (0.0 ms in the primary schedule), enabling higher effective compute utilization.
-
Hybrid 8-bit optimizer states with fast gradient clipping: Couple 8-bit CPU-resident optimizer states with accelerated CPU-side gradient clipping to reduce memory pressure without sacrificing update speed. This adds only 0.3% overhead while freeing GPU memory for larger activations or batch sizes.
-
Batch-size expansion capability: Use the scheduler to increase maximum feasible batch size by one unit at every model scale (e.g., 3B: batch 8 vs. 7; 27B: batch 2 vs. 1) on limited-memory GPUs like RTX 3090 24GB, enabling better gradient stability and data efficiency.
-
Offline optimal policy execution: Pre-solve the MILP once before training and execute the fixed policy at runtime, avoiding per-step scheduling overhead and enabling deterministic, reproducible training runs.
What the improved AI system can do:
-
Train larger models (up to 27B) on a single 24GB or 80GB GPU with throughput approaching that of multi-GPU setups, by eliminating memory-induced stalls.
-
Achieve higher token throughput (1361 tokens/s on 27B) while maintaining model accuracy (95.42% exact-match on MetaMathQA).
-
Run with larger batch sizes than previously possible on the same hardware, improving convergence and reducing training steps.
-
Operate with a flexible memory hierarchy (GPU HBM, CPU DRAM, NVMe) without manual tuning, automatically selecting optimal checkpoint paths and recomputation strategies.
-
Provide predictable training performance with zero exposed communication stalls, making it suitable for time-constrained or cost-sensitive training jobs.
Abstract
Training large language models on limited hardware is increasingly a scheduling problem across GPU compute, host memory, PCIe transfer, and storage bandwidth. Existing offloading systems reduce GPU residency, and MegaTrain shows that a CPU-master layer-streaming executor can train large models on a single GPU, but fixed checkpointing and placement heuristics still leave communication exposed on the critical path. We propose LazyTrain, an optimization layer over a layer-streaming executor. LazyTrain formulates checkpoint selection, activation placement, recomputation, and CPU-GPU-NVMe communication overlap as a mixed-integer scheduling problem, then executes the solved policy during training. It further couples 8-bit optimizer states with fast gradient clipping as a single Hybrid 8-bit operator: state compression reduces optimizer-state memory, while fast clipping counteracts the additional CPU-side update overhead. Across H800 experiments from Qwen2.5-3B to Qwen3.6-27B, LazyTrain improves sustained TFLOPS over matched baselines runs by approximately 1.24 times; RTX 3090 experiments likewise increase the maximum feasible batch size by one at each model scale. In the primary Qwen3.6-27B H800 MetaMathQA run, LazyTrain reaches 219.95 TFLOPS and 1361 tokens/s at batch size 72, peaks at 68.84,GB of GPU memory, and obtains 95.42% exact-match accuracy on the full evaluation split. The source code is available at https://github.com/DataArcTech/LazyTrain.
Sources
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
- Efficient Training of Large Language Models on Distributed Infrastructures: A Survey
- Kimi K2.5: Visual Agentic Intelligence
- ZeRO++: Extremely Efficient Collective Communication for Giant Model Training
- MegaTrain: Full Precision Training of 100B+ Parameter Large Language Models on a Single GPU
- Adam: A Method for Stochastic Optimization
- Cost-Efficient LLM Training with Lifetime-Aware Tensor Offloading via GPUDirect Storage
- GLM-5: from Vibe Coding to Agentic Engineering
- PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering