PowerStep: Memory-Efficient Adaptive Optimization via p-Norm Steepest Descent

arXiv:2605.10335 · cs.LG, cs.AI, cs.CL, cs.NA, math.NA, math.OC · Submitted 2026-05-11 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "PowerStep: Memory-Efficient Adaptive Optimization via p-Norm Steepest Descent".

Tom: Adaptive optimizers like Adam are standard for training large neural networks but suffer from substantial memory overhead due to storing running estimates of first and second moments.

Jane: First, who's behind it and why it matters.

Title and authors: Tom: We've just been discussing the paper "PowerStep: Memory-Efficient Adaptive Optimization via p-Norm Steepest Descent." It’s clear this work is focused on solving a major scalability problem in training large models by changing how we handle the optimizer state.

Jane: That's right, Tom. The authors, Yao Lu, Dengdong Fan, and Shixun Zhang, are looking at moving away from the standard Adam architecture because of the substantial memory overhead associated with storing running estimates of first and second moments.

Lu: Their motivation stems from the idea that there’s a parallel line of work exploring optimization under an p-norm geometry, which is a different mathematical perspective than traditional adaptive methods.

Meng: I'm trying to grasp the simple takeaway for our team; what exactly is the core problem they are solving in plain terms beyond just "memory overhead"?

Lalam: They are essentially addressing the memory wall where standard adaptive optimizers take up too much space, which limits how many of those huge models we can practically train and study.

Tom: Precisely. The title itself points to the solution: using p-norm geometry to enable memory-efficient adaptive optimization, which is a clever way to get adaptation without the heavy state storage.

Jane: It’s about achieving coordinate-wise adaptivity, meaning each parameter gets its own personalized update logic based on its history, but doing it in a way that doesn't require storing those full second-moment statistics.

Lu: By framing the problem as finding the steepest descent direction within an p-norm constraint, they establish a geometric framework that dictates the optimal direction for movement.

Meng: So, instead of storing huge matrices for moments, they are using this geometric definition to figure out where to step next. That feels like a much more manageable computational task for large-scale training infrastructure.

Lalam: It’s an important conceptual shift because it suggests that the adaptation isn't solely dependent on accumulating historical statistics in a dense vector, but rather on applying a specific transform to the current momentum state.

The paper's summary: Tom: So, looking at the summary of "PowerStep: Memory-Efficient Adaptive Optimization via p-Norm Steepest Descent," they are proposing an optimizer that applies a nonlinear transform directly to the heavy-ball momentum buffer to achieve coordinate-wise adaptivity without needing those second-moment statistics.

Jane: That's the central mechanism again; they take the standard momentum accumulation, m t = gamma m t-one + g t, and then apply this signed power transform beta(m t) to get the parameter update theta t = theta t-one - eta t beta(m t).

Lu: The paper details how this update can be rephrased as a preconditioned momentum step, where the diagonal preconditioner D t is defined by m t,i beta-one which ties the adaptivity directly to the local momentum magnitude.

Meng: That diagonal preconditioner sounds like a very localized control mechanism; it allows each parameter to react differently based on whether its corresponding momentum is large or small.

Lalam: This local control is what makes it coordinate-wise adaptive; when a direction has accumulated a lot, the preconditioner dampens the step to prevent overshooting, and when it's flat, it amplifies the step to speed things up.

Tom: So they are achieving this adaptivity purely through this transform and its resulting diagonal preconditioner without ever looking at the second-moment accumulator during updates.

Jane: That is a key distinction; you aren't storing v t for moments, which saves massive amounts of memory, yet you still get adaptive behavior based on the momentum history.

The paper's improvements: Tom: The paper lays out several improvements, and one of the most significant is that PowerStep achieves the optimal convergence rate of O(one/sqrt T) for non-convex stochastic optimization under standard regularity conditions.

Jane: That optimal rate is achieved while still matching Adam’s convergence speed on large Transformer models, which they show through extensive experiments on models ranging from 124M to 235B parameters.

Lu: Beyond the convergence proof, they demonstrate that when combined with aggressive int8 quantization, PowerStep remains numerically stable and reduces optimizer memory by about eight times compared to full-precision AdamW.

Meng: That eight times reduction in memory is huge for our infrastructure because it directly impacts how much data we can process and how much model size we can manage on our current setups.

Lalam: The stability under int8 quantization is a major win, as it shows the optimizer's behavior isn't overly sensitive to the low precision used for weight updates during deployment.

Tom: So, what about hyperparameter tuning? The study found that choosing a power exponent beta = zero point one provided the best balance between rapid early progress and long-run stability across different model sizes.

Jane: That specific choice of beta = zero point one shows that the method is quite robust; it’s not overly sensitive to minor variations in the hyperparameter range, which simplifies the process of finding a reliable setting for performance.

Conclusion: Tom: Wrapping up this discussion on "PowerStep: Memory-Efficient Adaptive Optimization via p-Norm Steepest Descent," we see an optimizer that successfully achieves coordinate-wise adaptivity without storing second-moment statistics by leveraging geometric principles.

Jane: In short, they prove that applying a nonlinear transform to the momentum buffer yields this adaptivity and proves optimal convergence at O(one/sqrt T), which is matched against Adam’s speed on very large Transformer models.

Lu: The implications are that we have a new mathematical framework for adaptive optimization grounded in p-norm geometry that can be applied broadly to complex, high-dimensional learning landscapes.

Meng: From an engineering standpoint, the practical impact is huge because it allows us to train larger models with significantly reduced memory footprints while maintaining competitive convergence rates.

Lalam: The ultimate vision is that this kind of efficiency enables us to build more powerful AI systems that are accessible and deployable on a wider variety of hardware by decoupling optimization complexity from the sheer size of the model state.

Tom: So, we’ve seen how PowerStep tackles memory and convergence simultaneously with a novel approach based on p-norm geometry.

Jane: It really shows that efficiency doesn't have to come at the expense of performance when you use a principled mathematical foundation like this for your adaptive processes.

Lu: This paper provides a solid foundation for future research exploring how geometric constraints can guide learning in ways that are currently unexplored in standard stochastic optimization literature.

Meng: We need to focus on integrating these memory savings into our next generation of training pipelines immediately, because the practical gains there are substantial and immediate.

Lalam: This work points toward a future where AI systems are built not just on raw parameter count, but on smart, memory-efficient methods that allow for truly massive scale and accessibility.

Yao Lu, Dengdong Fan, Shixun Zhang, Yonghong Tian

Peking University

cs.LG, cs.AI, cs.CL, cs.NA, math.NA, math.OC

Submitted: 2026-05-11

Updated: 2026-09-29

Code: https://github.com/yaolubrain/PowerStep

Importance score: 88/100

The gist: Adaptive optimizers like Adam are standard for training large neural networks but suffer from substantial memory overhead due to storing running estimates of first and second moments.

Key concepts

Coordinate-wise Adaptivity
This means the optimizer adjusts the learning rate differently for every single parameter in the model based on its own historical movement. PowerStep achieves this by using a diagonal preconditioner derived from momentum, allowing it to scale updates locally without needing complex statistics like those in Adam.
Signed Power Transform (Φβ(x))
This is a mathematical function that takes the sign and magnitude of the momentum buffer. It is used to scale the update step based on how large or small the past movement was for each parameter, providing a nonlinear way to adjust learning rates.
p-Norm Steepest Descent Geometry
The algorithm is rooted in finding the best direction for descent under an l_p-norm geometry. This mathematical framework defines the optimal update direction based on minimizing a specific norm constraint, which dictates how the parameters should be adjusted for faster convergence.

Terminology

Summary

Adaptive optimizers like Adam are standard for training large neural networks but suffer from substantial memory overhead due to storing running estimates of first and second moments. PowerStep introduces a memory-efficient optimizer that achieves coordinate-wise adaptivity without storing these second-moment statistics, demonstrating that applying a nonlinear transform directly to a momentum buffer yields this adaptivity, leading to an optimal convergence rate of O(1/√T) for non-convex stochastic optimization while matching Adam's speed on large Transformer models.

The gist

PowerStep achieves coordinate-wise adaptivity without storing second-moment statistics by applying a nonlinear transform directly to a heavy-ball momentum buffer, proving convergence at the optimal O(1/√T) rate for non-convex stochastic optimization.

Algorithm Design and Geometric Foundation

The algorithm is derived from the principles of steepest descent under an lp-norm geometry. The optimal direction in this framework is defined as:

/v

v∗ = arg min v∈Rd ⟨∇f(θ), v⟩ s.t.∥v∥ ≤ 1 (Equation 2)

The paper defines the signed power transform as:

Φβ(x) = sign(x)⊙xβ (Equation 5), where β = 1/(p−1). PowerStep updates parameters using this transform on the momentum buffer: θt = θt−1 − ηtΦβ(mt) (Equation 7). This update is equivalent to a preconditioned momentum step, where the diagonal preconditioner Dt is defined as Dt = diag(mtβ−1) (Equation 9).

Adaptivity Mechanism

The adaptivity of PowerStep stems from the coordinate-wise scaling provided by the diagonal preconditioner Dt. The update rule can be expressed as: θt = θt−1 − ηtDtmt (Equation 8). For β ∈ (0, 1), the exponent β − 1 is strictly negative, meaning each entry in Dt is inversely related to the local momentum magnitude: when mt,i is small (a flat direction), the preconditioner amplifies the step to accelerate progress; when mt,i is large (a steep direction), it dampens the update to prevent overshooting. This coordinate-wise scaling is computed instantaneously from mt without maintaining a second-moment estimator.

Convergence Analysis

The convergence analysis establishes an optimal O(1/√T) rate for PowerStep in non-convex stochastic optimization under standard regularity conditions (Assumptions 1–5). The proof relies on several key lemmas:

/Lemma 1 (Induced Norm Structure):

⟨m, Φβ(m)⟩ =∥m∥1+β1+β (Equation 43)

The convergence proof involves showing that the expected squared gradient norm is bounded by O(1/√T). This is achieved by bounding the expected second moment of the momentum, E[∥mt∥2], and applying Lemma 5 (Descent Inequality), which relates the function decrease to the inner product with Φβ(mt). The final result proves that min t∈[T] E[∥∇f(θt−1)∥2] = O(1/√T) (Equation 96).

Empirical Validation and Robustness

Extensive experiments on Transformer models ranging from 124M to 235B parameters demonstrate that PowerStep matches Adam’s convergence speed while using only half the optimizer state memory. Furthermore, when combined with aggressive int8 quantization, PowerStep remains numerically stable and reduces optimizer memory by ∼8× compared to full-precision AdamW. The analysis shows that PowerStep's robustness to int8 quantization is due to four factors: (i) the absence of a reciprocal square root singularity, (ii) the Hölder continuity of Φβ which attenuates quantization noise with exponent β = 0.1, (iii) the elimination of the second-moment buffer and use of linear accumulation which preserves dynamic range, and (iv) heavy-ball momentum whose full-weight gradient updates make first-moment stalling far less likely than under EMA. This configuration preserves both convergence speed and numerical stability while reducing the optimizer memory footprint by approximately 8×.

Hyperparameter Sensitivity

An ablation study on GPT-2-Medium showed that the power exponent β = 0.1 provides the best trade-off, balancing rapid early progress with long-run stability, whereas β = 0 (equivalent to SignSGD) collapses, and β = 0.2 converges more slowly due to insufficient nonlinearity. The method is also shown to be robust across the momentum coefficient γ in the range [0.85, 0.95], with γ = 0.9 providing a reliable default setting for performance stability.

Improvements for AI systems

As a fastidious researcher, I have analyzed PowerStep: Memory-Efficient Adaptive Optimization via lp-Norm Steepest Descent. This paper introduces an optimizer that achieves coordinate-wise adaptivity without storing second-moment statistics, leveraging the geometry of the lp-norm.

Based on this research, here are the specific improvements to AI systems and what those improved systems can achieve:


)1. Memory Efficiency & Scalability (The Core Improvement)

The primary improvement is a massive reduction in optimizer memory footprint by eliminating the need to store second-moment statistics (which typically doubles the state size compared to SGD with momentum).

  • Inference/Training systems for large language models (LLMs) or complex neural networks.

  • PowerStep reduces optimizer memory by approximately 8x compared to full-precision AdamW, and further by using aggressive int8 quantization on the momentum buffer.

  • This allows training and fine-tuning of extremely large models (up to 235B parameters, as shown in experiments) that were previously infeasible due to GPU VRAM constraints for optimizer states.

  • It enables the deployment of larger, more complex models on resource-constrained hardware (e.g., mobile GPUs or edge devices) where memory is a critical bottleneck.

)2. Faster Convergence at Optimal Rate

PowerStep achieves the optimal convergence rate of 1/√T for non-convex stochastic optimization, matching AdamW's convergence speed while using half the memory.

  • In training complex models on high-dimensional, ill-conditioned loss landscapes (like Transformers), PowerStep will reach a stationary point in fewer iterations than standard SGD with momentum or potentially other memory-efficient optimizers like pbSGDM.

  • This translates directly into faster research cycles and reduced computational costs for achieving a target level of model performance.

)3. Enhanced Numerical Stability via Quantization Robustness

The design inherently makes the optimizer robust to aggressive low-precision quantization (int8). Unlike AdamW, which suffers catastrophic divergence due to errors in its reciprocal square root operations, PowerStep remains stable when the momentum buffer is quantized.

  • In production deployment pipelines where model weights might be compressed or where training infrastructure uses mixed-precision arithmetic, PowerStep ensures that the optimization process itself does not become unstable due to accumulated quantization noise.

  • This provides a safe path for deploying optimized models, as the optimizer's behavior is decoupled from the precision of the weight updates.

)4. Principled Adaptivity without Second Moments

PowerStep achieves coordinate-wise adaptivity through a geometric interpretation (steepest descent on lp-norm geometry) rather than relying on historical gradient statistics (second moments).

  • This provides a principled alternative to Adam when second-moment accumulation is computationally prohibitive or numerically unstable.

  • It allows the AI system to adapt its learning rate and direction based purely on the current momentum state, which is an instantaneous reflection of past gradients, rather than requiring a complex, high-dimensional state vector.

)5. Optimized Hyperparameter Tuning

The research highlights that PowerStep's performance is robust across a range of hyperparameters (e.g., power exponent β = 0.1 provides the best trade-off).

  • This suggests that AI researchers can use PowerStep with confidence when tuning optimization schedules, as the method is less sensitive to minor hyperparameter variations than some prior methods might be.

  • It simplifies the process of finding a good learning rate schedule for a given model size, reducing trial-and-error experimentation time.

Sources

Related papers