Rethinking Normalization Placement for LLMs: Post-Norm under Curriculum Depth Growing

arXiv:2608.13156 · cs.AI · Submitted 2026-08-13 · Read on arXiv

Sheng Ren, Yadong Wang, Naiqiang Tan, Jiangang Kong, Jun Fang, Rui Liu, Jun Wang, Kai Chen, Lipeng Liang, Xiang Chen

Nanjing University of Aeronautics and Astronautics · Didichuxing Co. Ltd

cs.AI

Submitted: 2026-08-13

Updated: 2026-08-14

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 95/100

Terminology

Summary

Affiliation: Nanjing University of Aeronautics and Astronautics; Didichuxing Co. Ltd


The paper investigates whether the standard preference for pre-norm normalization placement in Transformer language models persists when model depth is introduced through a curriculum rather than through conventional joint training. The authors state: "We ask whether this preference persists when depth is introduced through a curriculum. In curriculum depth growth, each appended block receives the boundary representation produced by a trained prefix, making normalization placement relevant to forward conditioning."

The central hypothesis is that "normalization placement interacts with the way depth is introduced. Under joint training, normalization primarily supports optimization through a complete stack. Under curriculum depth growth, it also shapes the forward distribution passed from the trained stack to each newly appended block."

The authors use a block-stack student with K=3 blocks and L=3 decoder layers per block, giving S=9 trainable layers total, distilled from a Qwen3-8B teacher with 36 decoder layers. The student copies the teacher embedding E and language-modeling head Wlm. The embedding remains frozen, while the language-modeling head and decoder layers are trained. All variants share the same attention and MLP sublayers—rotary position embeddings, SwiGLU activations, grouped-query attention, and scaled dot-product attention—so the comparison isolates the placement of RMSNorm relative to each residual update.

Pre-norm applies RMSNorm immediately before each sub-layer but does not explicitly normalize the residual stream after each residual update:

  • h′ = Attention(RMSNorm(h)) + h

  • hout = MLP(RMSNorm(h′)) + h′

  • A final RMSNorm is applied before the language-modeling head.

Post-norm applies RMSNorm after each residual addition:

  • u = Attention(h) + h, h′ = RMSNorm(u)

  • v = MLP(h′) + h′, hout = RMSNorm(v)

  • No additional RMSNorm after the final block; a single entry RMSNorm is applied before the first block.

Both variants contain two RMSNorm modules per decoder layer and exactly one model-boundary RMSNorm, giving them the same number of trainable normalization parameters.

Joint training: All K blocks are active from the first optimization step and trained end-to-end.

Curriculum growth: Blocks are appended sequentially. Each phase activates one additional block, carries over previously trained weights, and introduces a newly initialized block. At each phase transition, the optimizer state is reset and the learning-rate schedule is restarted with warmup.

Each grow phase runs for 50,000 steps, processing approximately 3.11B non-padding tokens per phase. The three grow phases consume the same data shard, so grow sees 3.11B unique tokens repeated across phases (9.34B total). Joint training runs 50,000 steps over the same 3.11B-token shard. The primary comparison matches unique-token exposure, while compute is examined separately through active-layer-token controls.

At each phase, the active student is trained with: Lp = ατ2KL(qτ teacher qτ p) + γCE(y, z p), where τ is the distillation temperature, z p denotes the logits of the active student, and y denotes the next-token labels.

Under joint training, pre-norm and post-norm reach nearly identical validation losses: CE values of 2.7603 and 2.7607 respectively, a difference of only 0.0004. Under curriculum growth, the ranking changes: post-norm reduces validation CE from 2.7658 to 2.7330 compared with pre-norm. This yields ∆joint = +0.0004 and ∆grow = −0.0328, where ∆ = Post − Pre. The grow gap is roughly eighty times larger than the joint gap. The resulting interaction, ∆grow − ∆joint = −0.0332, quantifies the observed shift toward post-norm under curriculum growth.

The per-phase results show where the ranking changes during curriculum growth. Pre-norm is slightly better in the shallow Phase 0, where the student contains only one block (CE 2.9802 vs. 2.9846). After additional blocks are appended, post-norm becomes better in both Phase 1 and Phase 2. This crossover matches the expected shift from optimizing a shallow block to conditioning the boundary distribution passed to appended blocks.

The authors test whether post-norm is already stronger when training a single block in isolation. Phase 0 results show pre-norm is marginally ahead on CE, KD loss, and PPL, with all gaps below 0.01 in CE/KD. This rules out post-norm forming a better shallow block as the source of its grow advantage; instead, the advantage appears only after additional blocks are appended.

The authors test whether additional student-decoder active compute explains the grow advantage. Post-joint improves from 2.7607 to 2.7518 when matched to grow on student active-layer tokens, and to 2.7428 when matched to grow on total training tokens. Both remain worse than post-grow (2.7330), despite the additional compute. This suggests that the post-norm grow result is not reducible to additional student active-layer compute.

The freeze variant follows the grow protocol exactly, except that at the start of each phase, previously completed blocks are frozen and only the newly appended block receives updates. Freeze is uniformly worse than retrain by approximately 0.144 CE for both placements. The placement gap, however, is essentially unchanged: ∆grow = −0.0328 versus ∆freeze = −0.0332, a difference of only −0.0004. This shows that retraining earlier blocks does not account for the placement gap.

Post-norm keeps block-boundary RMS controlled, whereas pre-norm drifts as depth grows: after block 0, the per-token RMS standard deviation rises from 5.81 to 19.98 under pre-norm but remains approximately 0.07–0.19 under post-norm. On the final-phase diagnostic batch, pre-norm shows a heavy-tailed distribution (kurtosis > 2400) versus a far less heavy-tailed post-norm distribution.

The extreme kurtosis suggests that pre-norm's drift is driven by a small subset of tokens rather than a uniform shift. The analysis localizes pre-norm's drift to structural tokens: document-boundary RMS grows from 119 to 425 (37× regular-token RMS at 150k), while regular-token RMS grows only from 7.1 to 11.5. Post-norm keeps both below 2.5 (2.45 vs. 2.40), so the drift is concentrated at structural-token massive activations rather than a uniform global shift.

On a fixed diagnostic batch, the authors bypass each entire block and measure ∆CE and angular distance. Under pre-norm growth, bypassing the final block changes CE by only 0.02, and its angular distance is 0.031, indicating a near-identity mapping. Removing any post-grow block instead raises CE by at least 2.38, with angular distances above 0.34. Neither joint model has a near-identity final block. Post-joint places low removal sensitivity on block 0 (∆CE = 0.55) but high sensitivity on block 2 (4.00), showing that uneven allocation alone does not imply the pre-grow failure pattern.

The placement gap emerges after blocks are appended. Transition-aligned loss reveals that post-norm incurs larger immediate increases at both transitions, with much slower recovery after the first and only a small delay after the second, whereas pre-norm changes more smoothly. The smooth pre-norm transition is consistent with a smaller initial perturbation, yet its final block is nearly identity-mapped; post-norm incurs a larger adjustment but develops nontrivial block changes.

The paper concludes: "Pre-norm and post-norm are nearly tied under joint training, whereas post-norm is better under curriculum depth growing. Single-block, compute-matched, freeze, and boundary diagnostics associate this crossover with controlled boundary scales, ruling out shallow-block quality, active-layer compute, or retraining alone. The results therefore motivate treating normalization placement and depth curriculum as coupled rather than independent design choices."

The authors recast curriculum depth growth as block-wise initialization, shifting the normalization question from full-depth optimization to inherited boundary conditioning. They demonstrate that normalization placement interacts with the training curriculum, with post-norm performing better under sequential depth growth in the evaluated block-stack setting. The analysis through single-block controls, block-removal analysis, and block-boundary activation diagnostics provides converging evidence for forward conditioning as the contributing mechanism.

Improvements for AI systems

Improvements to AI Systems Based on This Paper

  1. Curriculum-Aware Normalization Selection
  • Improvement: Implement an automated normalization-placement selector that switches from pre-norm to post-norm when training uses depth-growing curricula (e.g., progressive layer stacking, block-wise distillation).

  • What the improved system can do: Automatically choose post-norm for staged model growth, reducing validation loss by 0.03 CE (as shown) and avoiding the pre-norm boundary-scale drift that degrades performance in deeper phases.

  1. Boundary-Scale Monitoring and Correction
  • Improvement: Add a runtime diagnostic that tracks per-token RMS at block boundaries during training. If pre-norm drift exceeds a threshold (e.g., RMS > 20 or kurtosis > 1000), the system can either switch to post-norm or insert an additional normalization layer at the drifting boundary.

  • What the improved system can do: Prevent catastrophic activation spikes (e.g., document-boundary RMS growing 37×) that cause unstable training, especially in long-context or multi-document tasks, by dynamically correcting scale drift.

  1. Transition-Aware Learning-Rate Scheduling
  • Improvement: Modify the curriculum growth protocol to use asymmetric learning-rate recovery: slower warmup after the first block append (where post-norm shows slower recovery) and faster warmup after subsequent appends.

  • What the improved system can do: Reduce the immediate loss spike at phase transitions (e.g., from 0.03 to <0.01 CE) and stabilize training when adding new blocks, leading to faster convergence and better final performance.

  1. Block-Importance-Weighted Distillation
  • Improvement: Use block-removal sensitivity (measured via ∆CE and angular distance) to weight the distillation loss for each block during curriculum growth. Blocks with near-identity mappings (like pre-norm’s final block) should receive higher regularization or be re-initialized.

  • What the improved system can do: Ensure every appended block contributes non-trivially to the model’s function, avoiding degenerate “pass-through” blocks that waste parameters and reduce representational capacity.

  1. Structural-Token-Aware Normalization
  • Improvement: Implement a token-type-conditional normalization that applies stronger scaling or separate normalization statistics for structural tokens (e.g., document boundaries, special tokens) when using pre-norm, since these tokens drive the heavy-tailed drift.

  • What the improved system can do: Maintain stable activations for both regular and structural tokens, improving robustness on long documents, multi-turn dialogues, or code with many delimiters, without sacrificing pre-norm’s optimization benefits.

  1. Curriculum-Phase-Adaptive Architecture Search
  • Improvement: Extend the block-stack framework to search over normalization placement per phase (e.g., pre-norm in Phase 0, post-norm in Phases 1–2) based on the observed crossover.

  • What the improved system can do: Automatically discover phase-specific normalization strategies that outperform any single fixed placement, potentially yielding further loss reductions beyond the uniform post-norm result.

  1. Freeze-Aware Initialization for Appended Blocks
  • Improvement: When freezing earlier blocks during curriculum growth, initialize the new block’s normalization parameters to match the boundary scale of the frozen prefix (e.g., set post-norm’s scale to the observed RMS of the incoming representation).

  • What the improved system can do: Reduce the freeze-vs-retrain performance gap (currently 0.144 CE) by ensuring the new block starts with well-conditioned inputs, enabling more efficient fine-tuning of only the appended layers.

Sources

Related papers