Gradient Prediction with Control Variates in the Cheap-Forward Regime
cs.LG, stat.ML
Submitted: 2025-11-07
Updated: 2026-09-02
License: http://creativecommons.org/licenses/by/4.0/
The gist: We study whether otherwise-idle inference resources could reduce the scarce-GPU cost of training.
Terminology
Abstract
We study whether otherwise-idle inference resources could reduce the scarce-GPU cost of training. Our analysis uses a simulated compute ledger in which fleet work is billed at a fraction of a scarce-GPU forward; all experiments run on a regular GPU. Our algorithm predicts gradients with a reduced-precision, inference-style reverse-mode program and combines many predictions with a few exact gradients through a control variate, so approximation error becomes variance rather than bias. On a 124M-parameter language model and selected short training windows, the method can lower simulated ledger cost relative to the tested baselines when fleet work is sufficiently cheap. Experiments spanning 10M-774M parameters show both transfers and failures. We do not test inference-only hardware, end-to-end distributed latency, or a full optimizer-by-batch-size baseline sweep.
Sources
- On the Convergence of SGD with Biased Gradients
- Deep Equals Shallow for ReLU Networks in Kernel Regimes
- Gradient Descent Happens in a Tiny Subspace
- Decoupled Weight Decay Regularization
- Characterizing the Spectrum of the NTK via a Power Series Expansion
- Eigenvalues of the Hessian in Deep Learning: Singularity and Beyond
- Low Rank Gradients and Where to Find Them
- Llama 2: Open Foundation and Fine-Tuned Chat Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks