Gefen: Optimized Stochastic Optimizer
summary
The gist
AdamW is a default optimizer for modern deep learning, but its first and second moment states add roughly two parameter-sized buffers to training memory, increasing the already substantial cost of
In short
Gefen is a memory-efficient optimizer designed to reduce AdamW's large memory footprint by about 8 times. It achieves this by automatically grouping parameters whose squared gradients are similar and sharing second-moment estimates within these groups. This method maintains AdamW's performance while requiring no extra hyperparameters or architecture-specific metadata.
Key concepts
- Second-Moment Estimates
- These are statistical summaries used by optimizers to track the history of gradients, helping the algorithm adapt its learning rate. In AdamW, these estimates consume significant memory. Gefen shares these estimates across related parameters to save space.
- Automatic Block Partitioning
- This process automatically divides model parameters into groups or 'blocks' based on their initial squared gradients. It does this using only the first-iteration gradient information, avoiding the need to compute complex Hessian matrices directly.
- Dynamic Programming Quantization Codebook
- The method learns a specific way to quantize and reuse these parameter blocks for scaling the first moment. This learned codebook allows the optimizer to efficiently apply scaling operations while ensuring that second-moment statistics are correctly shared within the identified parameter groups.
Terminology used across episodes
This episode discusses
- Gefen: Optimized Stochastic Optimizer · Paper Radio
- 8-bit Optimizers via Block-wise Quantization
- Adam: A Method for Stochastic Optimization
- PyTorch Distributed: Experiences on Accelerating Data Parallel Training
- Muon is Scalable for LLM Training
- Decoupled Weight Decay Regularization
- On the Convergence of Adam and Beyond
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
- Qwen3 Technical Report
- PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel
- APOLLO: SGD-like Memory, AdamW-level Performance
The paper
Gefen: Optimized Stochastic Optimizer · Read on arXiv
Nadav Benedek, Tomer Koren, Ohad Fried
Reichman University · Tel Aviv University · Google Research
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Gefen: Optimized Stochastic Optimizer".
Jane: AdamW is a default optimizer for modern deep learning, but its first and second moment states add roughly two parameter-sized buffers to training memory,
Tom: First, who's behind it and why it matters.
Paper summary: Tom: We've covered the initial claims of Gefen, which centers on reducing AdamW's memory usage by about eight times through automatic sharing of second-moment estimates across parameter blocks and quantizing the first moment using a learned codebook.
Jane: The paper is motivated by a theoretical result suggesting that large mixed Hessian entries constrain the ratio of squared gradients toward one, which points to Hessian-aligned parameters as natural candidates for sharing those statistics.
Lu: In terms of what Gefen actually does, it infers block structure from the initial squared gradients without computing Hessians directly and then learns an exact histogram-based dynamic programming quantization codebook to reuse those blocks.
Meng: The goal is to reduce that memory cost, specifically aiming for a reduction of six point five GiB per billion parameters while keeping the performance similar to AdamW <ref:2606.13894#pg0,a reduction of 6.5 GiB per billion parameters>.
Lalam: The method updates the first moment using dequantize and then quantize it back up, which ensures those second-moment estimates are shared within the inferred blocks.
Tom: This is interesting because they're not adding any new learning rate adjustments or extra hyperparameters; they rely on AdamW defaults to get this effect.
Jane: They also discuss the optimizer desiderata, which sets out criteria like peak memory footprint and distributability, showing that this isn't just about a trick for one model architecture.
Meng: The practical aspect is that it needs to be compatible with sharding strategies like DDP and FSDP, because if it doesn't fit in those setups, the memory saving is moot.
Lu: The theoretical justification comes from two theorems: Theorem four point one shows that high Hessian affinity implies similar squared gradients, and Theorem four point two links layer affinity to comparable raw gradient magnitudes up to the same curvature scale <ref:2606.13894#pg1,that high Hessian affinity implies similar squared gradients>.
Tom: So it’s a combination of solid math—those Hessian relationships—and clever automated inference from first-iteration data to group parameters efficiently for memory reuse.
Jane: It really speaks to how we can use theoretical properties of optimization dynamics, like those related to the Hessian, to create practical optimizations for large-scale AI training.
Meng: From an engineering standpoint, the fact that it requires no complex architecturespecific metadata is a huge win because it means we don't have to constantly update our model structure documentation just for this optimizer.
Lalam: If this works as claimed, it means future large models could be trained on hardware with much tighter memory constraints without sacrificing the quality of the training run.
Conclusion: Tom: So we've talked about Gefen: Optimized Stochastic Optimizer, and what it does by automating memory sharing based on gradient similarity guided by Hessian theory.
Jane: It seems like a really practical approach because it maintains AdamW-level performance while delivering a significant reduction in the persistent memory footprint, specifically showing that an eight times reduction is achievable for large models.
Lu: The authors are putting forward a method that relies on inferring block structure from just the initial squared gradients to group parameters and then using learned quantization to share second moments effectively.
Meng: It's a solid tool because it's general, compatible with standard distributed training frameworks, meaning researchers don't have to rewrite their entire optimization setup just for this optimizer.
Lalam: This means that for the next generation of large models, we can expect to train them on hardware with much tighter memory limits without sacrificing the quality of the training run.
Tom: The paper shows that by leveraging theoretical insights into how gradients relate to curvature, we can create an optimizer that is both efficient and performs comparably to established ones like AdamW.
Jane: Gefen takes a complex problem—managing massive second moments—and solves it by grouping parameters where the theory predicts they should be grouped together based on their Hessian properties.
Lu: It’s a powerful demonstration of how theoretical results about Hessian structure can translate directly into practical, memory-saving algorithms for deep learning optimization.
Meng: I think the main implication is that we get a way to train bigger models more affordably without having to constantly worry about memory bottlenecks in the optimizer state.
More episodes
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck
- 2407.14562-Thought-Like-Pro: Enhancing Reasoning of Large Language Models through Self-Bootstrapped Prolog-based Chain-of-Thought