Compute-Optimal Pretrain--Fine-tune in Ridge Gradient Descent
stat.ML, cs.LG
Submitted: 2026-09-14
Updated: 2026-09-14
License: http://creativecommons.org/licenses/by/4.0/
The gist: Pretraining followed by fine-tuning introduces a compute-allocation problem: under a fixed training budget, compute spent improving the upstream objective reduces the compute available for downstream
Terminology
Abstract
Pretraining followed by fine-tuning introduces a compute-allocation problem: under a fixed training budget, compute spent improving the upstream objective reduces the compute available for downstream adaptation. Despite its practical importance, this trade-off is not yet well understood theoretically, even in simple models. In this paper, we cast this allocation as a compute-split problem under a two-stage pretrain--fine-tune procedure with fixed total optimisation budget, using regularised least squares trained by gradient descent as a tractable setting. We characterise the optimal split under data-dependent evaluation geometries induced by the fine-tuning problem. Our results show that the allocation depends on how pretraining directions affect fine-tuning predictions and how fine-tuning shifts are seen through downstream data geometry. In particular, the relevant quantities are determined by prediction-relevant spectral components of the pretraining and fine-tuning empirical covariances. Technically, the analysis relies on a basis-invariant, eigenspace-level spectral decomposition, together with perturbative control of the non-commuting pretraining and fine-tuning dynamics.
Sources
- Delta Tuning: A Comprehensive Study of Parameter Efficient Methods for Pre-trained Language Models
- On the Provable Advantage of Unsupervised Pretraining
- Scaling Laws for Transfer
- Training Compute-Optimal Large Language Models
- Editing Models with Task Arithmetic
- Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
- LaMDA: Language Models for Dialog Applications
- Risk Comparisons in Linear Regression: Implicit Regularization Dominates Explicit Regularization
- When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning Method
Related papers
- Behavior of prediction performance metrics with rare events
- Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs
- A Posterior-Dynamics Framework for Imaging Inverse Problems with Pretrained Diffusion Priors
- One Permutation Is All You Need: Fast, Deterministic Feature Importance and Model Stress-Testing
- Online Conformal Prediction for Non-Exchangeable Panel Data
- Deep Time-Series Forecasting in 10 Years: A Survey