Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model
Mohammed Sabry, Sean Augenstein, Keith Rush, Lucio Dery
Google · Dublin City University
cs.CL, cs.AI
Submitted: 2026-08-13
Updated: 2026-08-14
Comments: Accepted at the Workshop on Methods and Opportunities at Small Scale (MOSS), COLM 2026
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 51/100
The gist: Mixture of Training (MoT) is a scaffolded modular pre-training procedure that partitions a target Transformer into contiguous layer blocks, trains each block inside a frozen pretrained aligner
Terminology
Summary
Mixture of Training (MoT) is a scaffolded modular pre-training procedure that partitions a target Transformer into contiguous layer blocks, trains each block inside a frozen pretrained aligner scaffold, and then recomposes the trained blocks with an optional short end-to-end adaptation pass. On a 1.3B-parameter Gemma-style model trained on C4, MoT provides a small-scale proof of mechanism: independently trained depth slices can be recomposed into a usable language model, and a quality-parity schedule reaches the same reported perplexity as the monolithic baseline. This parity setting processes more aggregate tokens and has a shorter idealized layer-equivalent critical path after aligner preparation; its effective compute advantage depends on reusing the aligner across runs. We therefore present MoT not as a general replacement for monolithic pre-training, but as a small-scale framework for studying whether scaffolded sub-runs can act as reusable training units.
The method works as follows: the target Transformer is written as a composition of K contiguous layer blocks, F = fK ◦ fK−1 ◦ · · · ◦ f1, where each fi is a target submodel. MoT trains these submodels independently but not in isolation, introducing a pretrained aligner A = aK ◦ · · · ◦ a1, sliced into the same number of blocks and shape-compatible with the target. For each target block fi, MoT constructs a scaffolded network Si = aK ◦ · · · ◦ ai+1 ◦ fi ◦ ai−1 ◦ · · · ◦ a1, where only fi is trainable and all aligner slices are frozen. Each scaffold is optimized with the standard next-token prediction loss. The aligner is required to be shape-compatible with the target model: it uses the same global width, attention-head dimensionality, feed-forward width, token embedding space, and output head. Each scaffold reuses the aligner’s token embeddings and output head; only fi is updated. The recomposed model retains these shared components but replaces the aligner’s Transformer slices with the trained target blocks.
MoT has three stages. Stage 0 prepares or selects the aligner and partitions both aligner and target into corresponding blocks. Stage 1 trains all scaffolded networks Si in parallel, updating only the target block in each scaffold. Stage 2 discards the aligner, recomposes F̂ = fK ◦ · · · ◦ f1, and optionally applies a short end-to-end adaptation pass. The recomposed model before Stage 2 adaptation is called the cold-composed model; its quality directly measures how well scaffolded training aligned the independently trained blocks.
The experiments evaluate MoT on the English portion of C4. The target model is a 12-layer, 1.3B-parameter decoder-only Transformer following the Gemma-1-2B width configuration: 256k token vocabulary, 2048-dimensional embeddings, RoPE, multi-query attention, and 16384-dimensional feed-forward layers. The monolithic baseline is trained end-to-end for 128k updates, processing 33.6B tokens and reaching perplexity 15.0 at 268.4 EFLOPs. Unless otherwise noted, MoT uses K = 2 target blocks, a 4-layer aligner, disjoint data streams for the two submodels, and evaluation is always performed on the recomposed target model rather than on individual scaffolds.
Table 1 summarizes the main quality–compute trade-off across three MoT schedules. Cold composition trains the two 6-layer submodels for 50k updates each and evaluates the recomposed model before any end-to-end adaptation, reaching PPL 19.3 at 128.2 train EFLOPs (157.9 fully charged). MoT + 15k adaptation adds a short Stage-2 pass, reaching PPL 15.9 at 159.7 train EFLOPs (189.4 fully charged). MoT quality parity reinvests part of the saved compute by extending submodel training to 75k updates and using a 30k adaptation pass, reaching PPL 15.0 at 255.3 train EFLOPs (285.0 fully charged). The critical-path estimates are 4.2×, 2.8×, and 1.7× respectively, assuming concurrent Stage 1 jobs after aligner preparation.
If the aligner is charged fully to a single run, the quality-parity schedule costs 255.3 + 29.7 = 285.0 EFLOPs, above the 268.4 EFLOP monolithic baseline. Quality parity is therefore not compute-saving in the fully charged single-run setting. More generally, if the same 29.7 EFLOP aligner is reused across R independent, shape-compatible target-model training runs, the effective quality-parity cost per run is 255.3 + 29.7/R. The effective cost falls below the monolithic baseline for R ≥ 3, so the quality-parity result is interpreted as evidence for an amortized-reuse regime rather than an unconditional efficiency gain. The 50k + 15k schedule remains below the full 128k-step baseline budget even when the aligner is fully charged, reaching PPL 15.9 at 128.2 + 31.5 + 29.7 = 189.4 EFLOPs.
The aligner is critical in this setting. Without it, cold-composition quality drops sharply: no aligner with K = 2 and shared data reaches PPL 38.9, versus 20.3 with the aligner. Disjoint streams improve PPL under the tested 4-layer aligner (19.3 vs 20.3) but worsen it without an aligner (50.4 vs 38.9). Increasing K from 2 to 4 lowers compute but worsens cold-composition PPL (24.8 vs 19.3 with aligner and disjoint data), exposing a quality–efficiency trade-off. These ablations provide behavioural evidence for the role of the aligner, but they do not by themselves fully diagnose the internal mechanism. The large degradation without an aligner is consistent with an interface-mismatch failure: independently trained depth slices do not necessarily produce hidden states that the next slice can use.
The experiments support three conclusions. First, independently trained submodels can be recomposed into a coherent language model when trained inside a shared aligner scaffold; without the scaffold, cold composition degrades sharply. Second, the remaining mismatch is mostly recoverable rather than catastrophic: a short adaptation pass closes most of the gap, and a longer quality-parity schedule reaches the same reported perplexity as the monolithic baseline. Third, MoT changes the engineering shape of pre-training. Instead of one coupled run, it creates smaller jobs that can be scheduled, restarted, and ablated independently, and potentially reused across compatible training efforts, making it especially suitable for small-scale training research.
The current study is intentionally small-scale. It uses one model family, one dataset, and a limited set of schedules, and reports perplexity rather than downstream reasoning, factuality, calibration, or robustness benchmarks. The study also does not yet include all compute-matched monolithic controls, such as baselines matched for fully charged MoT EFLOPs, aggregate token exposure, or estimated critical-path budget, nor does it report measured wall-clock time under equal hardware resources. Consequently, the current evidence demonstrates modular composability in this setting and identifies an amortized-reuse scheduling regime, rather than a uniform compute or equal-hardware speed advantage.
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems:
-
Modular Pre-training with Reusable Aligners: I can implement a training framework where a large Transformer is split into independently trainable depth blocks, each trained inside a frozen, shape-compatible aligner scaffold. This allows parallel, restartable, and independently ablatable sub-jobs, reducing the critical path of training from a single monolithic run to concurrent smaller runs (e.g., 4.2× shorter idealized critical path in the cold-composition setting). The improved system can train models with better scheduling flexibility and fault tolerance, as any sub-block can be retrained without restarting the whole model.
-
Amortized Compute Reuse Across Runs: By reusing a single pretrained aligner across multiple target-model training runs, I can reduce the effective compute per run below monolithic baselines. For the quality-parity schedule, the effective cost drops below the 268.4 EFLOP monolithic baseline when the aligner is reused for R ≥ 3 runs (cost = 255.3 + 29.7/R EFLOPs per run). The improved system can train multiple shape-compatible models (e.g., different datasets, tasks, or random seeds) at lower marginal cost, enabling efficient multi-task or multi-domain model families.
-
Cold-Composition Quality Diagnostics: I can use the cold-composed model (before any end-to-end adaptation) as a diagnostic tool to measure how well independently trained blocks align. This allows me to detect interface mismatches between depth slices early, without full fine-tuning. The improved system can automatically identify which block boundaries cause the largest hidden-state mismatches (e.g., by comparing PPL of cold-composed vs. adapted models) and then target additional training or architectural adjustments specifically to those interfaces, improving overall composability.
-
Short Adaptation Pass for Rapid Recovery: I can implement a two-stage training pipeline where after modular pre-training, a short end-to-end adaptation pass (e.g., 15k updates) recovers most of the quality gap (from PPL 19.3 to 15.9, close to the 15.0 baseline). The improved system can quickly salvage partially trained or recomposed models, reducing the need for full retraining when sub-blocks are swapped or updated.
-
Data Stream Isolation for Parallel Training: I can use disjoint data streams for each sub-block during scaffolded training, which improves cold-composition quality (PPL 19.3 vs. 20.3 with shared data) under the aligner. The improved system can train each depth slice on independent data subsets without cross-contamination, enabling more efficient parallel data loading and reducing memory bandwidth bottlenecks during multi-GPU training.
-
Aligner-Based Interface Standardization: I can use the frozen aligner as a fixed interface that maps hidden states between blocks, ensuring that each independently trained block produces outputs compatible with the next block. This allows me to train blocks with different hyperparameters (e.g., learning rates, batch sizes) without coordination, as long as they are shape-compatible. The improved system can support heterogeneous training configurations per block, optimizing each depth slice for its specific role (e.g., lower layers for syntax, higher layers for semantics) while maintaining global coherence.
-
Quality-Efficiency Trade-off Control: I can adjust the number of blocks K (e.g., from 2 to 4) to trade off compute savings against cold-composition quality (PPL worsens from 19.3 to 24.8 with K=4). The improved system can dynamically select K based on available compute and quality requirements, allowing users to choose between faster but slightly lower-quality pre-training or slower but higher-quality monolithic training.
-
Compute-Aware Scheduling: I can estimate the critical path and fully charged EFLOPs for different MoT schedules, enabling me to schedule training jobs on distributed hardware to minimize wall-clock time. For example, the cold-composition schedule has a 4.2× shorter critical path, so I can prioritize it for time-sensitive applications, while using quality-parity schedules for compute-constrained but quality-critical scenarios. The improved system can automatically select the optimal schedule based on hardware availability and quality targets.
Abstract
We ask whether language-model pre-training can be decomposed into smaller, independently trainable jobs that can later be recomposed into a coherent larger model. We introduce Mixture of Training (MoT), a scaffolded modular pre-training procedure that partitions a target Transformer into contiguous layer blocks, trains each block inside a frozen pretrained aligner scaffold, and then recomposes the trained blocks with an optional short end-to-end adaptation pass. On a 1.3B-parameter Gemma-style model trained on C4, MoT provides a small-scale proof of mechanism: independently trained depth slices can be recomposed into a usable language model, and a quality-parity schedule reaches the same reported perplexity as the monolithic baseline. This parity setting processes more aggregate tokens and has a shorter idealized layer-equivalent critical path after aligner preparation; its effective compute advantage depends on reusing the aligner across runs. We therefore present MoT not as a general replacement for monolithic pre-training, but as a small-scale framework for studying whether scaffolded sub-runs can act as reusable training units.
Sources
- Revisiting Model Stitching to Compare Neural Representations
- Model Stitching: Looking For Functional Similarity Between Representations
- Training Compute-Optimal Large Language Models
- L4: Diagnosing Large-scale LLM Training Failures via Automated Log Analysis
- m2mKD: Module-to-Module Knowledge Distillation for Modular Transformers
- DiffusionBlocks: Block-wise Neural Network Training via Diffusion Interpretation
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering