Gefen: Optimized Stochastic Optimizer

arXiv:2606.13894 · cs.LG, cs.AI, cs.CL, cs.CV · Submitted 2026-06-11 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Gefen: Optimized Stochastic Optimizer".

Jane: AdamW is a default optimizer for modern deep learning, but its first and second moment states add roughly two parameter-sized buffers to training memory,

Tom: First, who's behind it and why it matters.

Paper summary: Tom: We've covered the initial claims of Gefen, which centers on reducing AdamW's memory usage by about eight times through automatic sharing of second-moment estimates across parameter blocks and quantizing the first moment using a learned codebook.

Jane: The paper is motivated by a theoretical result suggesting that large mixed Hessian entries constrain the ratio of squared gradients toward one, which points to Hessian-aligned parameters as natural candidates for sharing those statistics.

Lu: In terms of what Gefen actually does, it infers block structure from the initial squared gradients without computing Hessians directly and then learns an exact histogram-based dynamic programming quantization codebook to reuse those blocks.

Meng: The goal is to reduce that memory cost, specifically aiming for a reduction of six point five GiB per billion parameters while keeping the performance similar to AdamW <ref:2606.13894#pg0,a reduction of 6.5 GiB per billion parameters>.

Lalam: The method updates the first moment using dequantize and then quantize it back up, which ensures those second-moment estimates are shared within the inferred blocks.

Tom: This is interesting because they're not adding any new learning rate adjustments or extra hyperparameters; they rely on AdamW defaults to get this effect.

Jane: They also discuss the optimizer desiderata, which sets out criteria like peak memory footprint and distributability, showing that this isn't just about a trick for one model architecture.

Meng: The practical aspect is that it needs to be compatible with sharding strategies like DDP and FSDP, because if it doesn't fit in those setups, the memory saving is moot.

Lu: The theoretical justification comes from two theorems: Theorem four point one shows that high Hessian affinity implies similar squared gradients, and Theorem four point two links layer affinity to comparable raw gradient magnitudes up to the same curvature scale <ref:2606.13894#pg1,that high Hessian affinity implies similar squared gradients>.

Tom: So it’s a combination of solid math—those Hessian relationships—and clever automated inference from first-iteration data to group parameters efficiently for memory reuse.

Jane: It really speaks to how we can use theoretical properties of optimization dynamics, like those related to the Hessian, to create practical optimizations for large-scale AI training.

Meng: From an engineering standpoint, the fact that it requires no complex architecturespecific metadata is a huge win because it means we don't have to constantly update our model structure documentation just for this optimizer.

Lalam: If this works as claimed, it means future large models could be trained on hardware with much tighter memory constraints without sacrificing the quality of the training run.

Conclusion: Tom: So we've talked about Gefen: Optimized Stochastic Optimizer, and what it does by automating memory sharing based on gradient similarity guided by Hessian theory.

Jane: It seems like a really practical approach because it maintains AdamW-level performance while delivering a significant reduction in the persistent memory footprint, specifically showing that an eight times reduction is achievable for large models.

Lu: The authors are putting forward a method that relies on inferring block structure from just the initial squared gradients to group parameters and then using learned quantization to share second moments effectively.

Meng: It's a solid tool because it's general, compatible with standard distributed training frameworks, meaning researchers don't have to rewrite their entire optimization setup just for this optimizer.

Lalam: This means that for the next generation of large models, we can expect to train them on hardware with much tighter memory limits without sacrificing the quality of the training run.

Tom: The paper shows that by leveraging theoretical insights into how gradients relate to curvature, we can create an optimizer that is both efficient and performs comparably to established ones like AdamW.

Jane: Gefen takes a complex problem—managing massive second moments—and solves it by grouping parameters where the theory predicts they should be grouped together based on their Hessian properties.

Lu: It’s a powerful demonstration of how theoretical results about Hessian structure can translate directly into practical, memory-saving algorithms for deep learning optimization.

Meng: I think the main implication is that we get a way to train bigger models more affordably without having to constantly worry about memory bottlenecks in the optimizer state.

Nadav Benedek, Tomer Koren, Ohad Fried

Reichman University · Tel Aviv University · Google Research

cs.LG, cs.AI, cs.CL, cs.CV

Submitted: 2026-06-11

Updated: 2026-10-04

Code: https://github.com/ndvbd/Gefen

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 89/100

The gist: AdamW is a default optimizer for modern deep learning, but its first and second moment states add roughly two parameter-sized buffers to training memory, increasing the already substantial cost of

Key concepts

Second-Moment Estimates
These are statistical summaries used by optimizers to track the history of gradients, helping the algorithm adapt its learning rate. In AdamW, these estimates consume significant memory. Gefen shares these estimates across related parameters to save space.
Automatic Block Partitioning
This process automatically divides model parameters into groups or 'blocks' based on their initial squared gradients. It does this using only the first-iteration gradient information, avoiding the need to compute complex Hessian matrices directly.
Dynamic Programming Quantization Codebook
The method learns a specific way to quantize and reuse these parameter blocks for scaling the first moment. This learned codebook allows the optimizer to efficiently apply scaling operations while ensuring that second-moment statistics are correctly shared within the identified parameter groups.

Terminology

Summary

AdamW is a default optimizer for modern deep learning, but its first and second moment states add roughly two parameter-sized buffers to training memory, increasing the already substantial cost of large-scale pretraining. Gefen proposes a memory-efficient optimizer that automatically shares second-moment estimates across parameter blocks and quantizes the first moment using a learned codebook, thereby reducing AdamW’s memory footprint by ∼8× while maintaining the same performance, corresponding to a reduction of 6.5 GiB per billion parameters.

How it works

Gefen is designed to reduce the optimizer's peak memory footprint by automatically grouping parameters whose squared gradients are similar and sharing second-moment estimates within these groups, which is motivated by a theoretical result showing that large mixed Hessian entries constrain the ratio of squared gradients toward one, suggesting that Hessian-aligned parameters are natural candidates for sharing secondmoment statistics.

The method achieves this grouping by inferring block structure from the initial squared gradients, requiring no architecturespecific metadata or hyperparameters beyond AdamW defaults.

The process involves several automated steps:

  1. AUTOMATIC BLOCK PARTITIONING: This step uses a one-dimensional squared-gradient-based algorithm for partitioning parameters into blocks, without computing the Hessian directly, by analyzing the first-iteration gradient only to find candidate periods based on the average contrast between adjacent block means.

  2. EXACT DP QUANTIZATION CODEBOOK LEARNING: This step learns an exact histogram-based dynamic-programming quantization codebook to reuse the same blocks for first-moment scaling, which is described as a learnable and exact dynamic programming quantization that automatically reuses the block partitioning described earlier.

  3. Sharing Second Moments: The algorithm updates the first moment using DEQUANTIZE and then applies QUANTIZE to obtain m¯ t, mt∞, which is then used to update parameters, ensuring that second-moment estimates are shared within the inferred blocks.

Theoretical Grounding

The paper provides theoretical justification for the grouping strategy through two main theorems. Theorem 4.1 demonstrates that large mixed Hessian entries contract squared-gradient ratios toward one, which implies that parameters with high Hessian affinity have similar squared gradients, making them suitable candidates for sharing second-moment estimates.

Theorem 4.2 further shows that within a layer, if the affinity ρij is close to one, then the Hessian-normalized gradient magnitudes of coordinates i and j are close, which leads to the conclusion that the raw gradient magnitudes gi and gj are comparable up to the same curvature scale.

Empirical Results

Empirically, Gefen achieves the lowest peak optimizer memory among the compared AdamW-like methods while maintaining AdamW-level performance, and it reduces AdamW’s peak and persistent 32-bit memory footprints by ∼8× across various models and datasets.

For instance, for the Llama 3-1.5B model, the total optimizer state is 15 GiB with AdamW and 1.5 GiB with Gefen, representing a reduction of 9.7 GiB that can be crucial in memory-constrained settings.

In terms of throughput, Gefen improves FSDP throughput by 56% compared to AdamW and enables the DDP setting where AdamW cannot fit even a microbatch size of 1, as it allows for a larger microbatch to fit on each GPU in distributed training configurations.

Distributability and Compatibility

Gefen is designed to be general-purpose, requiring minimal optimizerspecific hyperparameters or manual configuration choices, and it is compatible with common distributed-training frameworks including DDP, FSDP, and others.

Table 3 confirms its compatibility: Gefen supports all four configurations (DDP, FSDP, FSDP2 Sharded Init, DeepSpeed ZeRO-3), whereas Adam4bit and Adam8bit do not support FSDP.

Furthermore, the method can be combined with other optimization methods; for example, combining it with Muon reduces Muon’s persistent memory state by 4× and decreasing peak memory by approximately 38% compared to Muon as the model gets larger, while maintaining nearly identical training curves.

Conclusion

Gefen successfully reduces AdamW’s second-moment memory by sharing estimates in Hessian-aligned regions, achieves performance similar to AdamW, and requires no learning-rate adjustment or additional hyperparameters, making it a practical drop-in replacement for AdamW with a lower memory footprint (up to 8× less than AdamW) and high throughput.

REFERENCES

Daniel Aloise, Amit Deshpande, Pierre Hansen, and Preyas Popat. Np-hardness of euclidean sumof-squares clustering. Machine learning, 75(2):245–248, 2009.

Rohan Anil, Vineet Gupta, Tomer Koren, and Yoram Singer. Memory efficient adaptive optimization. Advances in Neural Information Processing Systems, 32, 2019.

Axolotl. Axolotl: Open source llm post-training, 2023. URL https://github.com/axolotl-ai-cloud/axolotl.

Ake Bj˚ orck and Clary Bowie. An iterative algorithm for computing the best estimate of an orthog- ¨onal matrix. SIAM Journal on Numerical Analysis, 8(2):358–364, 1971.

Xiangning Chen, Chen Liang, Da Huang, Esteban Real, Kaiyuan Wang, Hieu Pham, Xuanyi Dong, Thang Luong, Cho-Jui Hsieh, Yifeng Lu et al. Symbolic discovery of optimization algorithms. Advances in neural information processing systems 36:49205–49233, 2023.

Ronan Collobert. Large scale machine learning. 2004.

Sanjoy Dasgupta and Yoav Freund. Random projection trees for vector quantization. IEEE Transactions on Information Theory, 55(7):3229–3242, 2009.

Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition pp. 248–255. IEEE, 2009.

Tim Dettmers, Mike Lewis, Sam Shleifer, and Luke Zettlemoyer. 8-bit optimizers via block-wise quantization. arXiv preprint arXiv:2110.02861, 2021.

John Duchi, Elad Hazan, and Yoram Singer. Adaptive subgradient methods for online learning and stochastic optimization. Journal of machine learning research, 12(7), 2011.

Vineet Gupta, Tomer Koren, and Yoram Singer. Shampoo: Preconditioned stochastic tensor optimization. In International Conference on Machine Learning pp. 1842–1850 PMLR, 2018.

Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition pp. 770–778, 2016.

Nicholas J Higham. Functions of matrices: theory and computation. SIAM, 2008.

Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Advances in neural information processing systems, 33:6840–6851, 2020.

Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Chen, HyoukJoong Lee et al. Gpipe: Efficient training of giant neural networks using pipeline parallelism. Advances in neural information processing systems 32, 2019.

Shuoran Jiang, Qingcai Chen, Youcheng Pan, Yang Xiang, Yukang Lin, Xiangping Wu et al.

Improvements for AI systems

  1. The Gefen optimizer can be used as a drop-in replacement for AdamW with an 8× reduction in memory footprint and maintaining AdamW-level performance, enabling training of larger models or use of larger global batch sizes, as noted in the abstract.

  2. Gefen significantly improves throughput by enabling larger microbatches; specifically, it can increase throughput by 56% over AdamW in FSDP settings and enable microbatch sizes of 2 where AdamW is infeasible for Llama 3-1.5B on C4 (Figure 2).

  3. The method reduces peak memory footprint by ∼8× while maintaining the same performance, corresponding to a reduction of 6.5 GiB per billion parameters, allowing models like Llama 8B to be trained on fewer GPUs or with smaller GPUs (e.g., training Llama 8B on four RTX A6000s instead of H100s) (Section 6.1).

  4. Gefen leverages the theoretical result showing that large mixed Hessian entries constrain the ratio of squared gradients toward one, allowing it to automatically group parameters and share second moments within these groups without requiring architecture-specific metadata or hyperparameters beyond AdamW defaults (Section 1 and Section 4).

  5. The exact dynamic programming quantization codebook learns an optimal codebook by computing the minimum cost using a recurrence relation, ensuring that the exact DP method achieves the lowest quantization error compared to heuristic methods like Lloyd-Max, while requiring no initialization (Algorithm 3 and Section C.1).

  6. Gefen's partitioning mechanism is robust during training; as shown in Figure 9, 99% of the entries are green, indicating that the inferred period remains relatively stable throughout training, suggesting that the partitioning obtained from the first step appears to be a good proxy for the partitioning that would be obtained later (Section C.2).

  7. The Gefen-Muon combination allows for memory reduction in matrix parameter optimization; it reduces Muon’s persistent memory state by 4× and decreasing peak memory by approximately 38% compared to Muon as the model gets larger while preserving its optimization behavior (Table 4 and Table 5).

Abstract

AdamW is a default optimizer for deep learning, but its moment states add two parameter-sized buffers to training memory, increasing the cost of large-scale pretraining. We propose Gefen, a memory-efficient optimizer that automatically shares second-moment estimates across parameter blocks and quantizes the first moment using a learned codebook. Gefen reduces AdamW's optimizer memory footprint by up to 8x while maintaining performance, saving 6.5 GiB per billion parameters. Prior work shares second moments across parameters grouped along the Hessian's block-diagonal structure, but relies on hand-specified architectural rules and leaves unexplained why such grouping works. We prove that large mixed Hessian entries constrain the ratio of squared gradients toward one, explaining why shared second moments are accurate when the squared gradients they pool are similar. The Hessian need not be computed: its block structure is inherited by squared gradients, allowing blocks to be found directly. Gefen therefore infers block structure from initial squared gradients, requiring no architecture-specific metadata or user-tuned hyperparameters beyond AdamW defaults. Gefen learns an exact histogram-based dynamic-programming quantization codebook and reuses the blocks for first-moment scaling. Across diverse pretraining experiments, Gefen achieves the lowest peak optimizer memory among compared methods that maintain AdamW-level performance. In single-machine and distributed training, the reduced footprint enables larger microbatches and substantially improves throughput over AdamW, making Gefen a drop-in replacement that can train larger models or use larger global batch sizes. We provide the complete Python implementation, including fused CUDA kernels at https://github.com/ndvbd/Gefen

Sources

Related papers