Coupled Scaling: A Representational Accessibility Framework for Neural Scaling Laws
cs.LG
Submitted: 2026-09-03
Updated: 2026-09-13
Comments: 35 pages, 2 figures. Code and reproducibility artifacts: https://github.com/quintonvina/coupled-scaling/tree/v1.0-reanalysis
Code: https://github.com/quintonvina/coupled-scaling
License: http://creativecommons.org/licenses/by/4.0/
The gist: Existing theories derive neural scaling from data geometry or a specified data-model spectrum, but systems trained on the same data can scale differently when architecture or optimization changes the
Terminology
Abstract
Existing theories derive neural scaling from data geometry or a specified data-model spectrum, but systems trained on the same data can scale differently when architecture or optimization changes the representations they can efficiently reach. We introduce Coupled Scaling, a task-conditioned framework in which finite-budget scaling depends on the relation between task structure and the geometry accessible to an architecture-optimization system. In a solvable mode-truncation model, loss separates into target energy outside architectural support and an unresolved supported tail. For an arbitrary priority order, the residual lies between the best-N supported tail and the tail beyond the largest completed high-value prefix. If the cumulative-tail and coverage log-rates are γ A,T and ρ A,O,T, the residual exponent lies in [ρ A,O,Tγ A,T,γ A,T]. Under bounded off-prefix gain, the completed prefix is rate-determining and α A,O,T=ρ A,O,Tγ A,T; for a A,T,j j-b A,T, this gives α A,O,T=ρ A,O,T(b A,T-1). A fixed-kernel specialization derives the training-time exponent from the near-zero tail of a task-weighted spectral measure defined independently of the loss fit. The framework separates architectural support from finite-budget acquisition and motivates two tests: static task-relevant geometry should track loss at a common budget, while multiscale geometry should track coupling-specific exponent ordering, including reversal across contrasting tasks. An audit of released emergence trajectories identifies the controls needed for a direct factorial test that measures geometry separately from the scaling fit.
Sources
- A Qualitative Test-Risk Mechanism for Scaling Behavior in Normalized Residual Networks
- Optimal scaling laws in learning hierarchical multi-index models
- Deep Learning Scaling is Predictable, Empirically
- Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention
- Same Architecture, Different Capacity: Optimizer-Induced Spectral Scaling Laws
- Scaling Laws for Neural Language Models
- Spectral Lens: Activation and Gradient Spectra as Diagnostics of LLM Optimization
- What do Language Models Learn and When? The Implicit Curriculum Hypothesis
- Neural Scaling Universality: If Exponents Are Fixed, Time to Understand Coefficients
- Superposition Yields Robust Neural Scaling
- A Solvable Model of Neural Scaling Laws
- Spectral Reach: Understanding Neural Scaling as Progress into the Spectral Tail
- On the Optimizer Dependence of Neural Scaling Laws
- Data Scaling as Progressive Coverage of a Predictive Contribution Spectrum
- Towards Robust Scaling Laws for Optimizers
- SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers
- Mechanisms of Width Scaling in Normalized Residual Networks: The Effective Alignment Dimension
- Effective Frontiers: A Unification of Neural Scaling Laws
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks