Neural Quadratic Forms: A Unified Minimal Model for Sudden Learning and Scaling Laws

arXiv:2608.13335 · cs.LG, cond-mat.dis-nn, cond-mat.stat-mech · Submitted 2026-08-13 · Read on arXiv

Liu Ziyin, Yizhou Xu, Tomaso Poggio, Isaac Chuang

Massachusetts Institute of Technology · École Polytechnique Fédérale de Lausanne

cs.LG, cond-mat.dis-nn, cond-mat.stat-mech

Submitted: 2026-08-13

Updated: 2026-08-14

Code: https://github.com/xu-yz19/Neural-Quadratic-Forms

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 100/100

The gist: This paper introduces the Neural Quadratic Form (NQF), a universal minimal model that unifies sudden learning and neural scaling laws across diverse neural network architectures.

Terminology

Summary

This paper introduces the Neural Quadratic Form (NQF), a universal minimal model that unifies sudden learning and neural scaling laws across diverse neural network architectures. The central insight is that permutation symmetry—the invariance of a network layer under relabeling of its interchangeable units—combined with smoothness and a zero-gradient condition at the origin, enforces a universal leading-order quadratic expansion about the near-zero weights present at the start of training.

The paper proves that any module f x(w 1,, w d) that is S d-symmetric in its components, three times continuously differentiable, and satisfies the zero-gradient-at-zero condition obeys:

f x(W) - f x(0) = mu g(x) + Tr[W W A(x)] + O(W 3)

where W = (w 1,, w d) in R p times d collects the d components, mu = sum i=1 d w i, and A(x) in R p times p is a symmetric matrix fixed by the architecture and input alone. The paper calls A(x) the structure matrix. The term mu g(x) vanishes for most practical architectures, so the NQF can often be written with only the second-moment term.

The paper computes A(x) explicitly for several standard architectures:

  • Two-layer MLP: A(x) = phi'(0) over 2 0 k times k & x x & 0

  • Single-layer CNN: A(X) = phi'(0) over 2 0 m times m & sum p=1 P x p sum p=1 P x p & 0

  • Multi-head attention: A(X) = 1 over 2 0 & 0 & 0 & 0 0 & 0 & 0 & 0 0 & 0 & 0 & x avg I d v 0 & 0 & x avg I d v & 0, where x avg = 1 over N X 1 N

  • Mixture of experts: A(x) = 1 over 2 0 & psi'(0) x grad i E(x; 0) psi'(0) grad i E(x; 0) x & 0

  • Phase retrieval: A a = aa

  • Diagonal linear network: A x = 1 over 4 Diag(x)

A striking finding is that for attention, only its value and readout blocks appear at this order, the query and key matrices entering first at quartic order.

Under stochastic gradient descent, the dynamics of any NQF close on the pair (M, mu) = (sum i w i w i, sum i w i), regardless of the number of trainable parameters. The update equations are:

mu = -eta(d times v + H mu)

M = -eta(v mu + mu v + HM + MH) + eta 2(d times vv + v mu H + H mu v + HMH)

where v = sum x in B 'x (g(x) + 2B(x) mu) and H = sum x in B 2 'x A(x).

Any NQF with d neurons can be compressed to a smaller NQF with d' = k V + 1 neurons, where k V is the dimension of the joint subspace spanned by the data vectors and column spaces of A(x) and B(x). The compressed model reproduces the original's predictions on training data at every step under a rescaling.

Under a commutativity assumption (Assumption 1: data matrices share a common eigenbasis), the dynamics reduce to the Generalized Lotka-Volterra (GLV) equation from population ecology:

k(t) = 8 over m (sum mu=1 m y mu lambda mu,k - sum j=1 p sum mu=1 m lambda mu,k lambda mu,j z j(t)) z k(t)

The paper solves this in four regimes:

  1. Proportional samples (Theorem 4): A(x mu) = c mu A, yielding W(t) = P (-xi(t))P W(0)

  2. Orthogonal samples (Theorem 5): A(x mu)A(x nu) = 0 for mu not equal to nu, yielding independent per-sample dynamics

  3. Orthogonal features (Theorem 6): sum mu lambda mu,k lambda mu,j = 0 for k not equal to j, yielding logistic growth z k(t) = r k z k(0) over C kkz k(0) + (r k - C kkz k(0)) (-r k t)

  4. Isotropic samples (Theorem 7): second-moment tensor proportional to identity, yielding matrix Riccati solution M(t) = (Yt)M(0)[I + 8c (t)M(0)]-1 (Yt)

As the initialization scale epsilon to 0, the characteristic time for feature z k to be activated is:

t k* about 1 over a zeta k 1 over epsilon

where zeta k = r k (Theorem 6) or zeta k = gamma k (Theorem 7). The gaps between different feature activations diverge: epsilon to 0 t k'* - t k* = infinity if zeta k' not equal to zeta k.

When zeta k proportional to k-alpha 2 and V k proportional to k-alpha 1 (power-law decays), the limiting excess loss follows:

E(tau) = (tau-alpha 1 - 1 over alpha 2)

Under Theorem 5's conditions, with effective growth rate zeta mu = 8 over m lambda mu, y mu > 0, the empirical MSE converges to a saddle-to-saddle trajectory:

epsilon to 0+ L (tau 1 over epsilon) = 1 over m sum mu=1 m y mu squared times I(tau < tau mu*)

where tau mu* = 1/zeta mu. When zeta mu proportional to mu-gamma 2 and y mu proportional to mu-gamma 1 for gamma 1 > 1/2, the limiting MSE follows:

L(tau) = (tau-2 gamma 1 - 1 over gamma 2)

The paper validates the theory numerically across:

  1. NQF approximation: Two-layer MLPs, CNNs, query-key-only attention, and multi-head attention trained with GD, momentum, and Adam all track their NQF approximations almost exactly at small initialization (sigma = 0.01), while departing visibly at large initialization (sigma = 0.2).

  2. Feature-wise saddle-to-saddle dynamics: Numerical trajectories match Theorem 6 predictions, with sequential plateaus at predicted characteristic timescales t k*.

  3. Sample-wise saddle-to-saddle dynamics: Per-sample losses remain stagnant then undergo sharp decay near t mu*, producing m sequential plateaus in total loss.

  4. Power laws on Fourier MLP: A Fourier MLP with s k = k-theta and target coefficients b k = k-beta exhibits power-law loss decay matching the predicted exponent t-(2 beta-1)/(theta+ beta), while a standard tanh MLP on raw input does not.

The paper draws explicit analogies to physics:

  • Symmetry group: S d (permutation of components)

  • Order parameter: M = WW, which is S d-invariant and vanishes in the unlearned state (z k = 0) but not in the learned state (z k > 0)

  • Control parameter: epsilon (initialization scale), which controls the sharpness of crossovers; (1/epsilon) plays the role of system size N, and epsilon to 0 is the analogue of the thermodynamic limit

  • Sudden learning: ignition of individual modes of M, analogous to phase transitions emerging only in the infinite-system limit

The paper honestly states five limitations:

  1. Small initialization: The NQF is accurate only when initialization scale is small compared to feature length scale

  2. Smoothness: Requires three-times differentiability; rectified activations, biases, normalization layers violate this

  3. Other plateaus: Some plateaus observed experimentally are not accounted for by the NQF (attributed to cubic and higher terms)

  4. Power-law spectra are an assumption: The theory predicts exponents given power-law spectra but does not predict that spectra have power-law tails

  5. Adaptive optimizers: The linear-update corollary covers GD, weight decay, and momentum but not Adam (though experiments show Adam is tracked accurately)

Improvements for AI systems

Based on this paper, here are specific improvements I can make to AI systems, along with what the improved systems can do:

Improvement: I can now predict the exact early-training dynamics of any neural network module that satisfies permutation symmetry, smoothness, and zero-gradient-at-origin conditions. Instead of treating training as a black box, I can compute the structure matrix A(x) analytically for the architecture and use the closed-form update equations for (M, mu).

What the improved system can do:

  • Predict the loss curve trajectory before training begins, given only architecture and data statistics

  • Determine whether a given architecture will exhibit sudden learning (sharp phase transitions) or gradual learning

  • Estimate the characteristic timescale for each feature to be ignited as a function of initialization scale epsilon, without running training

Sources

Related papers