Gefen: Optimized Stochastic Optimizer

summary

Video file (mp4)

The gist

AdamW is a default optimizer for modern deep learning, but its first and second moment states add roughly two parameter-sized buffers to training memory, increasing the already substantial cost of

In short

Gefen is a memory-efficient optimizer designed to reduce AdamW's large memory footprint by about 8 times. It achieves this by automatically grouping parameters whose squared gradients are similar and sharing second-moment estimates within these groups. This method maintains AdamW's performance while requiring no extra hyperparameters or architecture-specific metadata.

Key concepts

Second-Moment Estimates
These are statistical summaries used by optimizers to track the history of gradients, helping the algorithm adapt its learning rate. In AdamW, these estimates consume significant memory. Gefen shares these estimates across related parameters to save space.
Automatic Block Partitioning
This process automatically divides model parameters into groups or 'blocks' based on their initial squared gradients. It does this using only the first-iteration gradient information, avoiding the need to compute complex Hessian matrices directly.
Dynamic Programming Quantization Codebook
The method learns a specific way to quantize and reuse these parameter blocks for scaling the first moment. This learned codebook allows the optimizer to efficiently apply scaling operations while ensuring that second-moment statistics are correctly shared within the identified parameter groups.

Terminology used across episodes

This episode discusses

The paper

Gefen: Optimized Stochastic Optimizer · Read on arXiv

Nadav Benedek, Tomer Koren, Ohad Fried

Reichman University · Tel Aviv University · Google Research

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Gefen: Optimized Stochastic Optimizer".

Jane: AdamW is a default optimizer for modern deep learning, but its first and second moment states add roughly two parameter-sized buffers to training memory,

Tom: First, who's behind it and why it matters.

Paper summary: Tom: We've covered the initial claims of Gefen, which centers on reducing AdamW's memory usage by about eight times through automatic sharing of second-moment estimates across parameter blocks and quantizing the first moment using a learned codebook.

Jane: The paper is motivated by a theoretical result suggesting that large mixed Hessian entries constrain the ratio of squared gradients toward one, which points to Hessian-aligned parameters as natural candidates for sharing those statistics.

Lu: In terms of what Gefen actually does, it infers block structure from the initial squared gradients without computing Hessians directly and then learns an exact histogram-based dynamic programming quantization codebook to reuse those blocks.

Meng: The goal is to reduce that memory cost, specifically aiming for a reduction of six point five GiB per billion parameters while keeping the performance similar to AdamW <ref:2606.13894#pg0,a reduction of 6.5 GiB per billion parameters>.

Lalam: The method updates the first moment using dequantize and then quantize it back up, which ensures those second-moment estimates are shared within the inferred blocks.

Tom: This is interesting because they're not adding any new learning rate adjustments or extra hyperparameters; they rely on AdamW defaults to get this effect.

Jane: They also discuss the optimizer desiderata, which sets out criteria like peak memory footprint and distributability, showing that this isn't just about a trick for one model architecture.

Meng: The practical aspect is that it needs to be compatible with sharding strategies like DDP and FSDP, because if it doesn't fit in those setups, the memory saving is moot.

Lu: The theoretical justification comes from two theorems: Theorem four point one shows that high Hessian affinity implies similar squared gradients, and Theorem four point two links layer affinity to comparable raw gradient magnitudes up to the same curvature scale <ref:2606.13894#pg1,that high Hessian affinity implies similar squared gradients>.

Tom: So it’s a combination of solid math—those Hessian relationships—and clever automated inference from first-iteration data to group parameters efficiently for memory reuse.

Jane: It really speaks to how we can use theoretical properties of optimization dynamics, like those related to the Hessian, to create practical optimizations for large-scale AI training.

Meng: From an engineering standpoint, the fact that it requires no complex architecturespecific metadata is a huge win because it means we don't have to constantly update our model structure documentation just for this optimizer.

Lalam: If this works as claimed, it means future large models could be trained on hardware with much tighter memory constraints without sacrificing the quality of the training run.

Conclusion: Tom: So we've talked about Gefen: Optimized Stochastic Optimizer, and what it does by automating memory sharing based on gradient similarity guided by Hessian theory.

Jane: It seems like a really practical approach because it maintains AdamW-level performance while delivering a significant reduction in the persistent memory footprint, specifically showing that an eight times reduction is achievable for large models.

Lu: The authors are putting forward a method that relies on inferring block structure from just the initial squared gradients to group parameters and then using learned quantization to share second moments effectively.

Meng: It's a solid tool because it's general, compatible with standard distributed training frameworks, meaning researchers don't have to rewrite their entire optimization setup just for this optimizer.

Lalam: This means that for the next generation of large models, we can expect to train them on hardware with much tighter memory limits without sacrificing the quality of the training run.

Tom: The paper shows that by leveraging theoretical insights into how gradients relate to curvature, we can create an optimizer that is both efficient and performs comparably to established ones like AdamW.

Jane: Gefen takes a complex problem—managing massive second moments—and solves it by grouping parameters where the theory predicts they should be grouped together based on their Hessian properties.

Lu: It’s a powerful demonstration of how theoretical results about Hessian structure can translate directly into practical, memory-saving algorithms for deep learning optimization.

Meng: I think the main implication is that we get a way to train bigger models more affordably without having to constantly worry about memory bottlenecks in the optimizer state.

More episodes

← Home