Structured Features Overfit Where Random Features Grok

arXiv:2609.15047 · cs.LG, stat.ML · Submitted 2026-09-14 · Read on arXiv

cs.LG, stat.ML

Submitted: 2026-09-14

Updated: 2026-09-14

Comments: 13 pages, 2 figures, 3 tables. Submitted to OPT 2026 (18th Annual Workshop on Optimization for Machine Learning)

License: http://creativecommons.org/licenses/by/4.0/

The gist: Xu, Vardi and Safran (ICML 2026) prove that over-parameterized ridge regression over an unstructured random Gaussian feature map groks, with the delay between memorization and generalization growing

Terminology

Abstract

Xu, Vardi and Safran (ICML 2026) prove that over-parameterized ridge regression over an unstructured random Gaussian feature map groks, with the delay between memorization and generalization growing as 1/λ in the weight decay. We show that on a structured feature map the same delay does not appear. For a band-limited Fourier feature map over Z p squared carrying a single-character target that lies inside the expressible class, enlarging the band at fixed positive weight decay drives peak held-out accuracy monotonically from 1.00 to 0.07, with no memorize-then-generalize regime anywhere along the sweep. The degradation is not an interpolation effect. It sets in at capacity ratio q/n = 0.638, far below the interpolation threshold, on separate grounds from the exact null space that appears above it. What does have a sharp boundary is the active support. Holding the nominal dimension fixed and masking the band back to 1089 active modes restores held-out accuracy of 1.000 with zero variance across seeds, while the full 4225-mode band collapses to 0.185. The number of active modes acts through the teacher-weighted spectrum of the empirical Gram matrix and not through the capacity ratio, which makes this a statement about feature geometry and not a restatement of double descent.

Sources

Related papers