Structured Features Overfit Where Random Features Grok
cs.LG, stat.ML
Submitted: 2026-09-14
Updated: 2026-09-14
Comments: 13 pages, 2 figures, 3 tables. Submitted to OPT 2026 (18th Annual Workshop on Optimization for Machine Learning)
License: http://creativecommons.org/licenses/by/4.0/
The gist: Xu, Vardi and Safran (ICML 2026) prove that over-parameterized ridge regression over an unstructured random Gaussian feature map groks, with the delay between memorization and generalization growing
Terminology
Abstract
Xu, Vardi and Safran (ICML 2026) prove that over-parameterized ridge regression over an unstructured random Gaussian feature map groks, with the delay between memorization and generalization growing as 1/λ in the weight decay. We show that on a structured feature map the same delay does not appear. For a band-limited Fourier feature map over Z p squared carrying a single-character target that lies inside the expressible class, enlarging the band at fixed positive weight decay drives peak held-out accuracy monotonically from 1.00 to 0.07, with no memorize-then-generalize regime anywhere along the sweep. The degradation is not an interpolation effect. It sets in at capacity ratio q/n = 0.638, far below the interpolation threshold, on separate grounds from the exact null space that appears above it. What does have a sharp boundary is the active support. Holding the nominal dimension fixed and masking the band back to 1089 active modes restores held-out accuracy of 1.000 with zero variance across seeds, while the full 4225-mode band collapses to 0.185. The number of active modes acts through the teacher-weighted spectrum of the empirical Gram matrix and not through the capacity ratio, which makes this a statement about feature geometry and not a restatement of double descent.
Sources
- Grokking Modular Polynomials
- Grokking modular arithmetic
- Algebraic Representability as the Limiting Regime of Grokking: An Exactly Solvable Model with Holomorphic Activations
- Grokking as the Transition from Lazy to Rich Training Dynamics
- Grokking in Linear Estimators -- A Solvable Model that Groks without Understanding
- Towards Understanding Grokking: An Effective Theory of Representation Learning
- Why Do You Grok? A Theoretical Analysis of Grokking Modular Addition
- Progress measures for grokking via mechanistic interpretability
- Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks