Small Enough to Know Everything: The Fully-Enumerable Transformer as an Instrument for the Science of Delayed Generalization
cs.LG
Submitted: 2026-08-03
Updated: 2026-09-18
Comments: 9 pages, 2 figures. Code and per-seed run records will be released publicly
License: http://creativecommons.org/licenses/by/4.0/
The gist: Tiny transformers trained on fully-enumerable tasks occupy an unusual position in the study of grokking: every input can be evaluated, every generalization ceiling can be computed exactly, and
Terminology
Abstract
Tiny transformers trained on fully-enumerable tasks occupy an unusual position in the study of grokking: every input can be evaluated, every generalization ceiling can be computed exactly, and hundreds of seeds cost minutes. We argue this regime is a scientific instrument with four capabilities that approximate settings cannot offer: (a) exact, falsifiable generalization ceilings; (b) task surgery that manipulates one structural variable while provably fixing all others; (c) direct observation of every weight; and (d) survival-time statistics over many seeds that recast "does not grok" as a censored observation. The obvious objection is that laws characterized at 10 4 parameters may not mean anything beyond them. We answer it with a preregistered conservation study: three task-side laws established at 12K parameters -- a recoverability-ceiling law, a role-conflict delay law, and a weight-decay response law -- are re-measured under an identical from-scratch protocol at 12K, 1M, and 50M parameters (a 4,000x span; 360 runs plus a 44-run control arm). The ceiling law and the delay law are conserved (0/144 Holm-corrected ceiling violations; Spearman rho >= 0.75 at every scale, permutation p < 1e-4), while the weight-decay law deforms systematically, steepening with scale. Preregistered controls show the 50M role-conflict deficit survives learning-rate adjustment and a tripled budget. Conservation was tested against criteria frozen before data collection, and one law's deformation shows the test could have failed. These results license the fully-enumerable transformer as a model organism for the task-side laws of delayed generalization: what it measures exactly, larger models largely obey -- and where they deviate, the deviation is itself lawful and measurable.
Sources
- Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets
- Grokking Is Conditional and Fragile: A Fully-Tractable, Multi-Seed Study at 12K Parameters
- At-Grok Is Not Converged:A Measurement-Validity Audit for Grokking Representation Metrics
- Grokking on the Weight-Decay Clock: A Rate Hierarchy from Softly Broken Symmetries
- Input Pathways Shape Few-Shot, Not Zero-Shot, Binding in Tiny Transformers: A Fully-Enumerable Study
- A Pre-Training Analogue of Grokking in Language Models: Tracing Delayed Grammatical Generalization
- Structure-Specific Representational Priors Causally Control the Grokking Delay
- What Makes a Representational Prior Work? Feature Families, Label-Free Invariances, and Critical Windows in Grokking
- Explaining grokking through circuit efficiency
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks