Predicting When Random Low-Dimensional Reparameterizations Train Neural Networks

arXiv:2608.12597 · cs.LG, cs.AI · Submitted 2026-08-12 · Read on arXiv

Andrew Cheng, Ali Eslamian, Jie Cheng, Mehdi Zargham, Qiang Cheng

Tsinghua University · The University of Manchester · University of Kentucky · Miami University · University of Dayton

cs.LG, cs.AI

Submitted: 2026-08-12

Updated: 2026-08-14

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: This paper addresses the question of how large a random low-dimensional search space must be to access parameters that achieve low loss when training neural networks through random low-dimensional

Terminology

Summary

This paper addresses the question of how large a random low-dimensional search space must be to access parameters that achieve low loss when training neural networks through random low-dimensional reparameterization. The authors first express the known accessibility transition in an equivalent conic form, centered at the statistical dimension of the polar cone for compact convex targets. Their main theoretical contribution is an orientation-resolved master formula that predicts the random-slice residual from both the curvature spectrum and the reference-to-solution displacement profile, making explicit how displacement orientation relative to curvature affects the required latent dimension. This formula yields a new self-consistent isotropic-orientation predictor and, in its conservative radius-only specialization, recovers the earlier Gaussian-width quadratic bound. Building on this analysis, they introduce Random Mapping Networks (RaMaN), a scalable framework that instantiates the predicted latent dimension using structured Hadamard mappings or seed-regenerated Gaussian maps, eliminating the O(dP) frozen-map storage and reducing optimizer-state memory from O(P) to O(d). The paper also develops matrix-free curvature approximations and sweep-free dimension-selection procedures, and designs a benchmark measuring accessibility transitions across tasks, architectures, and map families.

Theory. The paper recasts the accessibility transition in an equivalent conic/statistical-dimension form and derives a new orientation-resolved quadratic master formula. Unlike prior radius-only quadratic bounds, the master-formula predictor depends jointly on the curvature spectrum and the reference-to-solution displacement profile. Two useful specializations follow: a new self-consistent isotropic-orientation predictor, and an orientation-uniform predictor that recovers the earlier radius-only bound as a special case.

Method. The authors develop Random Mapping Networks (RaMaN), which couple curvature- and displacement-informed latent-dimension selection with scalable frozen random reparameterizations. RaMaN supports orientation-resolved and radius-based sweep-free selection together with nested adaptive dimension expansion. To instantiate the selected dimension without storing a dense P×d map, they develop a seeded Structured Hadamard Mapping (SHM) and a seed-regenerated Gaussian construction.

Measurement. They design a benchmark for measuring the phase transition of random low-dimensional training across tasks, architectures, and mapping families, reporting not only trainable parameters but also frozen-map storage, optimizer-state memory, checkpoint size, and wall-clock cost.

The paper considers low-dimensional reparameterizations of the form θ = gω(z), where z ∈ R d is trainable, d ≪ P, and ω denotes random, structured, or non-trainable parameters. For fine-tuning, they write θ(z) = θref + Δθ(z), where θref is a reference parameter vector. The low-loss sublevel set is defined as Sε = θ: L(θ) ≤ L* + ε, and the conic hull Cε:= cone(Sε − θ0) contains the rays from the reference point toward the low-loss region.

The paper proves Theorem 1, which states that for a localized quadratic target with positive-semidefinite Hessian H, the slice reaches the target set S if and only if E ∩ C ≠ 0. The intersection probability is at least 1−η when d ≥ P − δ(C) + aη√P and at most η when d ≤ P − δ(C) − aη√P, where aη = √(8 log(4/η)). The conic transition is centered at dconic:= P − δ(C) = δ(C°), within an O(√P) transition window.

The master-formula predictor is defined as:

ρ̂MF(d; Δ0):= Σi [κ(d)λi/(λi + κ(d))] Δi2

where κ(d) > 0 is the unique solution to d = tr[H(H + κ(d)I)−1] = Σi λi/(λi + κ(d)), and Δi = ⟨Δ0, vi⟩ is the component of the reference-to-solution displacement along the ith curvature eigenvector. The predicted critical dimension is d̂MF:= min d ∈ 1,...,r−1: ρ̂MF(d; Δ0) ≤ εq, where εq = 2ε.

Isotropic-orientation predictor: Under the equal-energy model Δi2 = R2/P, the master formula reduces to ρ̂iso(d) = (R2/P)κ(d)d, and d⋆iso is determined by d⋆iso = Σi λi/(λi + κ⋆) with κ⋆d⋆iso = 2εP/R2.

Orientation-uniform predictor: For every displacement satisfying ∥Δ0∥2 ≤ R, the master formula obeys ρ̂MF(d; Δ0) ≤ κ(d)R2. Imposing κ(d)R2 ≤ εq gives dunif(ε, R) = Σi λiR2/(λiR2 + εq), equivalently denoted reff(ε, R).

The paper notes that this orientation-uniform predictor recovers the earlier radius-only quadratic expression of Larsen et al., while the orientation-resolved predictor extends it by retaining the full displacement profile.

RaMaN instantiates the selected latent dimensions using frozen matrix-free random maps. Two default instantiations are considered:

Structured Hadamard Mapping (SHM): Uses a seeded random selection of Hadamard basis directions implemented through a fast Walsh–Hadamard transform. The map is defined by R SHM l,dl (u) = √(P̃l/Pl) crop Pl H P̃l D l I Ωl(dl)(u), requiring O(P̃l log P̃l) computation and O(P̃l) transform working memory.

Seed-regenerated Gaussian RaMaN: Uses a Gaussian random map that is never stored, with entries regenerated on demand using a counter-based pseudorandom number generator with fixed layer seed. The persistent map state is the random seed, while temporary block storage is controlled by the chosen block size.

The paper develops a three-mode dimension-selection procedure (Algorithm 1):

  • Mode 1 (orientation-resolved): Uses curvature and displacement estimates to compute the practical master-formula estimate d̃MF.

  • Mode 2 (radius selection): Uses only a displacement-radius estimate R with the equal-tail approximation to the orientation-uniform effective-rank predictor.

  • Mode 3 (adaptive): Uses nested adaptive expansion when neither estimate is reliable.

Across five synthetic spectra spanning low-rank, rapidly decaying, outlier-plus-bulk, and slowly decaying regimes, the isotropic-orientation predictor d⋆iso had absolute relative error at most 1.1% with median 0.36% relative to empirical midpoints. The orientation-uniform predictor did not underestimate d50 but ranged from approximately 1.1 to 17.5 times the empirical midpoint.

Varying displacement direction while keeping Euclidean norm fixed showed that orientation changes the transition midpoint by as much as a factor of 37 across three spectra. The master-formula predictor d̂MF had median absolute relative error of 0.45% and maximum absolute relative error of 5.36% over 30 orientation–spectrum combinations.

Applying the quadratic hit test to curvature operators from trained neural networks, the MF prediction differed from measured midpoint by at most 0.91% across six verified cases. The isotropic predictor d⋆iso underestimated the required dimension by as much as 21.0× and returned zero in one case. The anisotropy score mstiff ranged from 0.105 to 1.0.

On TinyMLP/MNIST, all evaluated RaMaN constructions displayed clear optimization transitions. Global and layer-wise seed-regenerated Gaussian maps had the same empirical midpoint (d50 = 939), while layer-wise SHM required d50 = 1,195. All constructions attained best observed test accuracies within approximately two percentage points of the 92.55% full-parameter baseline. On SmallCNN/CIFAR-10, layer-wise SHM transitioned earlier (d50 = 3,072) than layer-wise seed-regenerated Gaussian (d50 = 6,144).

Seed regeneration differed from dense-Gaussian accuracy by only 0.34 percentage points on TinyMLP and 0.85 points on SmallCNN while eliminating storage of the P×d random map. Checkpoint size fell from 45.41 to 0.01 MB on TinyMLP and from 267.73 to 0.04 MB on SmallCNN. SHM differed from dense Gaussian by only 0.31 points on TinyMLP but by 3.75 points on SmallCNN.

The conservative linear-tail correction selected d = 6,518 (oracle) and d = 4,602 (pilot), approximately 98% and 69% of P = 6,634. The equal-tail default reduced selected dimensions to 3,353 and 2,779 respectively, a reduction of approximately 40–49% while retaining successful retraining. Adaptive expansion reached the target at d = 2,048.

ViT-Tiny on CIFAR-10 and CIFAR-100 exhibited sharp transitions with midpoints approximately 2.1×10−4 to 1.8×10−3 of P. The CIFAR-100 midpoint was approximately 5.1× higher than CIFAR-10 for Gaussian maps and 8.7× higher for SHM. The map-family ordering reversed across tasks.

For bert-tiny on SST-2, the empirical midpoint used only 0.43–0.62% of the reparameterized budget (Prep = 413,314), approximately 35–75× smaller as a fraction than TinyMLP and SmallCNN trained from scratch.

Both tasks exhibited sharp training-success transitions at latent dimensions that are small fractions of the full parameter count (d50 ≈ 6×103 on CIFAR-10 and d50 = 10,752 on CIFAR-100). Under a protocol-matched comparison with widened learning-rate tuning, RaMaN at d = 1,048,576 (9.4% of P) reached 90.33% test accuracy, 0.38 percentage points above the 89.95% full-parameter reference, with a smaller generalization gap (6.0 versus 8.9 percentage points).

The sharp end-to-end transition persisted under both AdamW and SGD with momentum 0.9 on SmallCNN, but SGD shifted the transition toward larger latent dimensions (approximately 1.6–2.0× more). On TinyMLP, the two optimizers agreed within grid resolution.

Across a 20× tolerance range (ε ∈ 0.01, 0.02, 0.05, 0.10, 0.20), the AdamW midpoint changed by factors of 1.3–2.6 depending on model and map family. The existence of a sharp transition was robust to the tolerance choice.

Within the GGN–Ritz surrogate, the practical master-formula predictor tracked the measured midpoint to within 0.04–0.79% across four cells. The isotropic-orientation predictor evaluated to zero, while the practical orientation-uniform estimate exceeded the measured midpoint by factors of 4.9–47.3. However, only 0.2–6.0% of squared displacement was resolved on computed Ritz directions, so this is an internal consistency test rather than independent validation.

The paper establishes that low-loss geometry, rather than ambient dimension, governs accessibility. The orientation-resolved master formula makes explicit how displacement orientation relative to curvature affects the required latent dimension. The experiments show that geometric accessibility and end-to-end optimization are distinct: the measured transition is conditional on the optimizer, tuning protocol, and loss-excess tolerance. Training accessibility does not imply full-model generalization—the paper distinguishes dtrain (dimension needed to reach training-loss target) from dgen (dimension needed to approach full-parameter predictive performance). The ResNet-18 experiments show that substantially larger latent dimensions may be required for test accuracy to approach the full-parameter reference, though this gap is not intrinsic to random reparameterization.

The paper notes several limitations: the exact theory applies to uniformly random Gaussian subspaces while practical constructions introduce additional structure; Hessian-based predictors rely on local quadratic descriptions; and the ViT-scale analysis relies on GGN–Ritz approximation with modeled spectral tail. Future directions include extending theory to product and structured random maps, developing more reliable matrix-free curvature estimators, and characterizing dgen rather than only dtrain.

Improvements for AI systems

Improvement 1: Adaptive Low-Dimensional Training with Curvature-Aware Dimension Selection

The improved AI system can automatically determine the minimal trainable parameter dimension needed to reach a loss target by analyzing the local curvature spectrum and the displacement profile from a reference initialization. Instead of using fixed or heuristic latent dimensions, the system computes the orientation-resolved master formula predictor (ρ̂MF) in real time, selecting d such that the predicted residual falls below a tolerance. This reduces optimizer memory from O(P) to O(d) and eliminates frozen-map storage, enabling training of large models on memory-constrained devices. The system can also switch between isotropic, radius-only, and adaptive modes depending on the reliability of curvature estimates.

Improvement 2: Matrix-Free Random Reparameterization with Seed-Regenerated Maps

The improved AI system can train neural networks using random low-dimensional subspaces without ever storing the full P×d projection matrix. By using counter-based pseudorandom generators with fixed seeds, the system regenerates map entries on demand, reducing checkpoint size from tens of megabytes to kilobytes (e.g., 45.41 MB → 0.01 MB on TinyMLP). This enables deployment of large models on edge devices with limited storage, while maintaining accuracy within 0.34–0.85 percentage points of dense-Gaussian baselines. The system also supports structured Hadamard mappings for faster forward/backward passes via fast Walsh–Hadamard transforms when computational speed is prioritized over storage.

Improvement 3: Orientation-Aware Fine-Tuning for Transfer Learning

The improved AI system can fine-tune pre-trained models by exploiting the orientation of the reference-to-solution displacement relative to the loss landscape curvature. It identifies which curvature directions contribute most to the loss reduction and allocates latent dimensions accordingly, rather than treating all parameter directions equally. This allows the system to reach a target loss with up to 37× fewer dimensions when displacement aligns with high-curvature directions, as demonstrated in synthetic experiments. For fine-tuning tasks like BERT-tiny on SST-2, the system can achieve training success using only 0.43–0.62% of the reparameterized budget, significantly reducing computational and memory overhead.

Improvement 4: Sweep-Free Dimension Selection with Adaptive Expansion

The improved AI system can select the latent dimension without exhaustive sweeps, using either curvature-displacement estimates (mode 1), radius-only estimates (mode 2), or nested adaptive expansion (mode 3). When estimates are unreliable, the system starts with a small dimension and iteratively expands it based on observed loss progress, reaching the target dimension efficiently (e.g., d = 2,048 vs. oracle d = 6,518 in experiments). This reduces training time by avoiding unnecessary high-dimensional optimization while guaranteeing convergence to the loss target.

Improvement 5: Generalization-Aware Latent Dimension Scheduling

The improved AI system can distinguish between the dimension needed to reach a training-loss target (dtrain) and the dimension needed to approach full-parameter test performance (dgen). By monitoring the generalization gap during training, the system can dynamically increase the latent dimension after training convergence to improve test accuracy. For example, on ResNet-18/CIFAR-10, the system can start at d = 6,000 for training success and expand to d = 1,048,576 to achieve 90.33% test accuracy, surpassing the full-parameter baseline by 0.38 percentage points while maintaining a smaller generalization gap (6.0 vs. 8.9 points).

Improvement 6: Optimizer-Aware Dimension Scaling

The improved AI system can adjust its latent dimension selection based on the optimizer choice. Since SGD with momentum requires 1.6–2.0× larger dimensions than AdamW to reach the same loss target (as shown on SmallCNN), the system can automatically inflate the predicted dimension when SGD is used, preventing premature convergence to poor local minima. This ensures consistent training success across different optimizers without manual tuning.

Improvement 7: Task-Adaptive Map Family Selection

The improved AI system can choose between structured Hadamard mappings and seed-regenerated Gaussian maps based on the task and architecture, since the optimal map family varies (e.g., SHM transitions earlier on CIFAR-10 but later on CIFAR-100 for ViT-Tiny). The system can run a quick probe (e.g., 100 steps) with each map family at a fixed dimension, then select the one with the lowest loss trajectory for full training. This avoids the 3.75 percentage point accuracy drop observed when using SHM on SmallCNN/CIFAR-10, while retaining its speed advantages when appropriate.

Improvement 8: Robust Loss-Tolerance Calibration

The improved AI system can automatically calibrate its loss-excess tolerance (ε) based on the task difficulty and available compute budget. Since the required dimension scales by factors of 1.3–2.6 across a 20× tolerance range, the system can estimate the sensitivity and choose a tolerance that balances training speed against final performance. For high-stakes tasks, it uses a tighter tolerance (ε = 0.01) to ensure near-optimal loss, while for rapid prototyping it uses a looser tolerance (ε = 0.20) to reduce dimension and speed up training.

Improvement 9: Anisotropy-Aware Early Stopping

The improved AI system can detect when the loss landscape is highly anisotropic (using the mstiff score, which ranges from 0.105 to 1.0 in experiments) and adjust its dimension selection strategy accordingly. For highly anisotropic landscapes, it relies more heavily on the orientation-resolved master formula rather than isotropic predictors, which can underestimate required dimensions by up to 21×. This prevents premature convergence in ill-conditioned problems and ensures reliable training across diverse architectures.

Improvement 10: Memory-Scalable Vision Transformer Training

The improved AI system can train Vision Transformers (e.g., ViT-Tiny) on CIFAR-10/100 using only 0.02–0.18% of the full parameter count as trainable dimensions (midpoints at 2.1×10−4 to 1.8×10−3 of P). By combining seed-regenerated Gaussian maps with the master-formula dimension predictor, the system reduces optimizer memory from O(P) to O(d) (e.g., from 1M parameters to 2,000 trainable dimensions), enabling ViT training on devices with less than 1 GB of memory while maintaining competitive accuracy. The system also automatically adapts to task difficulty (CIFAR-100 requires 5–9× larger dimensions than CIFAR-10) without manual intervention.

Sources

Related papers