Quantifying the Memorization-to-Generalization Transition: Scaling Laws and Phase Structure in Grokking

arXiv:2609.10657 · cs.AI, cs.LG · Submitted 2026-09-09 · Read on arXiv

cs.AI, cs.LG

Submitted: 2026-09-09

Updated: 2026-09-09

License: http://creativecommons.org/licenses/by/4.0/

The gist: Neural networks trained past memorization frequently undergo a delayed transition to generalization, a phenomenon known as grokking.

Terminology

Abstract

Neural networks trained past memorization frequently undergo a delayed transition to generalization, a phenomenon known as grokking. Despite theoretical progress on why this transition occurs, the quantitative structure of when it occurs in hyperparameter space remains uncharacterized. We map the memorization-to-generalization boundary across 384 configurations of two-hidden-layer MLPs on modular arithmetic, fitting a power-law scaling relation for generalization onset time: T grok proportional to H-0.27, D-2.04, η-0.50, λ-0.64 (R squared = 0.732; 0.821 with interactions). The exponent hierarchy reveals that data complexity (D-2.04) is the dominant driver of regime transition, not model capacity (H-0.27): doubling data accelerates generalization by about 4 times, while doubling width yields only about 1.2 times. A sharp phase boundary at weight decay λ 1.0 separates grokking from non-grokking configurations, and weight norm trajectories show monotonic compression during the transition, consistent with implicit regularization selecting low-complexity solutions. These results provide a quantitative foundation for predicting and controlling regime transitions in overparameterized networks.

Related papers