Lions and Muons: Optimization via Stochastic Frank-Wolfe under Heavy-Tailed Noise

arXiv:2506.04192 · math.OC, stat.ML · Submitted 2025-06-04 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Lions and Muons: Optimization via Stochastic Frank-Wolfe under Heavy-Tailed Noise".

Jane: The paper was written by Jun-Kun Wang and Maria-Eleni Sfyraki from University of California San Diego.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Paper discussion segment 2: Tom: So, we've established that "Lions and Muons: Optimization via Stochastic Frank-Wolfe under Heavy-Tailed Noise" provides a unifying perspective, but let's look at what this means for the practical performance of these algorithms.

Jane: The paper gives us a solid theoretical underpinning by establishing convergence guarantees in terms of the Frank-Wolfe gap. That's a standard way to measure progress in non-convex optimization, which is huge for theoretical support.

Lu: And it goes further by showing that when convergence to this gap implies convergence to a KKT point under the norm constraint, that provides a very specific condition we can aim for during training.

Meng: I appreciate the focus on KKT points; it tells us precisely what kind of solution we are reaching, which is helpful for optimization stability.

Lalam: It also clarifies how these methods behave when the stochastic gradients exhibit heavy-tailed distributions, providing a robust foundation for large-scale AI training.

Tom: That robustness is critical, Jane, especially given that real-world data often results in those heavy tails instead of simple bounded variance assumptions.

Paper discussion segment 3: Tom: Now that we've understood the core unification and the general convergence guarantees for "Lions and Muons: Optimization via Stochastic Frank-Wolfe under Heavy-Tailed Noise," let's discuss what makes this paper truly novel in terms of improvements.

Jane: The authors introduce two specific robust variants to handle this heavy-tailed noise, which is a significant step beyond just having the initial theoretical framework.

Lu: They developed a clipping mechanism and also explored variance reduction techniques within the Stochastic Frank-Wolfe structure, which is something we haven't seen applied much before.

Meng: I see how practical impact improves when you can handle those unexpected spikes in gradient noise without losing performance, which is a major headache in real engineering problems.

Lalam: It’s great to see these new variants of Lion and Muon, Lalam, because it means we can adapt the theory to address real-world complexity while maintaining better performance.

Tom: So, we' have seen the foundational theory and the practical improvements for "Lions and Muons: Optimization via Stochastic Frank-Wolfe under Heavy-Tailed Noise." --- SEGMENT four: Paper discussion segment four ---

Tom: We've covered a lot of ground, from the initial conceptual link between Lion, Muon, and Stochastic Frank-Wolfe to the specific variants designed for heavy-tailed noise. Let's wrap up our conversation.

Jane: I think we can all agree that this paper offers a very general framework for analyzing these optimizers across different settings and constraint types.

Lu: It's a significant theoretical contribution, providing a common language to understand the dynamics of several state-of-the-art algorithms in deep learning.

Meng: The practical implication is that by using more robust methods like LION+ and MUON+, we are making training more reliable when facing messy real-world data.

Lalam: I believe "Lions and Muons: Optimization via Stochastic Frank-Wolfe under Heavy-Tailed Noise" will help us build more resilient AI systems, Lalam.

Tom: That's a great way to conclude. We've really seen how this paper unifies these algorithms and provided strong theoretical guarantees for the whole team, Tom.

Paper discussion segment 3: Tom: So, if we wrap up our discussion on this paper, the main theme we're taking away is how robust these optimization techniques are when faced with messy, unpredictable data in the real world.

Jane: Exactly! What I want everyone to grasp is that traditional optimization methods assume noise follows a nice bell curve—the normal distribution—but real systems are way messier than that, which is where heavy-tailed noise comes into play.

Lu: It's wild to think about the implications; if these algorithms can handle truly heavy tails, we're talking about optimizing complex AI models running on unreliable sensor networks, like deep space probes or autonomous vehicles in extreme weather.

Meng: From an engineering standpoint, the biggest practical hurdle is always stability; a small burst of noise that’s way outside the expected range can derail a standard algorithm entirely, so fixing that robustness is huge for deployment.

Lalam: Considering how much modern culture relies on massive amounts of data coming from imperfect sources, improving optimization under heavy tails fundamentally makes our digital infrastructure more trustworthy and less prone to catastrophic failure due to unexpected inputs.

Tom: That’s right, Meng brings up the key point about stability; it’s not just about getting a decent average result, it's about knowing that when things go sideways—when the noise spikes—the system doesn't crash.

Jane: And what I found fascinating is how the Stochastic Frank-Wolfe structure adapts to this messiness, giving us convergence guarantees even when the noise distribution is way more volatile than we usually assume.

Lu: It suggests a paradigm shift in how we model real-world uncertainty; instead of trying to smooth out the noise until it looks normal, we're building algorithms that explicitly account for the occasional massive outlier.

Meng: But Lu, making it stable sounds one thing, but implementing a framework that correctly models and accounts for *unknown* heavy tails—that requires significant computational overhead that I’m curious about.

Lalam: The ability to treat uncertainty as an inherent feature rather than a bug changes the whole philosophy of AI design; it encourages resilience in our systems, which builds public trust and accelerates adoption across critical fields like healthcare.

Tom: So, while the theory is amazing—and you all laid out how much better this robustness makes things—I’m really curious about one thing: what does this mean for optimization problems that aren't even formulated as standard machine learning tasks?

Conclusion: Tom: So, to wrap up our discussion on "Lions and Muons: Optimization via Stochastic Frank-Wolfe under Heavy-Tailed Noise," we're looking at how this powerful framework unifies several popular optimizers while also providing robust solutions for real-world noise.

Jane: It’s truly remarkable that the authors managed to bridge these seemingly disparate methods, showing us they aren't just random algorithms but special instances of a foundational Stochastic Frank-Wolfe approach.

Lu: I think the biggest impact is in how it allows us to design more resilient AI systems; knowing we have theoretical guarantees under heavy tails means we can push the boundaries of what's possible with confidence.

Meng: From a practical standpoint, this means that when deploying these optimizers, like LION+ or MUON+, in massive distributed training jobs, the risk of those high-variance noise spikes causing a complete failure is significantly reduced.

Lalam: It’s about building reliable cultural infrastructure; if our foundational AI algorithms are more robust against unexpected inputs, the resulting applications—from medical diagnostics to climate modeling—become more trustworthy.

Tom: I agree with Lalam; we're moving toward a future where AI systems are not just fast, but inherently stable.

Jane: It’s reassuring to know that the convergence rate for achieving an epsilon-gap is also well-defined, which gives us concrete targets for optimization efforts.

Lu: The complexity results, particularly the improved rates achieved by Algorithm six and LION++, are a testament to how much we've advanced in understanding non-convex dynamics.

Meng: And that really translates into better resource utilization during training phases, saving time and computational power on those large GPU clusters.

Lalam: It’s a huge step forward in making the very foundation of our AI culture more dependable and ultimately safer for everyone using these systems.

Tom: Well, I think we've seen enough today; it's been an incredibly insightful look at how this research unifies theory with practical engineering needs.

Jane: We’re excited to see what the next paper brings to our show, but we hope you found the discussion on "Lions and Muons: Optimization via Stochastic Frank-Wolfe under Heavy-Tailed Noise" as enlightening as we did.

Jun-Kun Wang, Maria-Eleni Sfyraki

University of California San Diego · Pethick University of California San Diego, msfyraki@ucsd.edu†, jkw005@ucsd.edu

math.OC, stat.ML

Submitted: 2025-06-04

Updated: 2026-08-25

Comments: Accepted at the Forty-Third International Conference on Machine Learning (ICML) 2026

Code: https://github.com/karpathy/nanoGPT

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 100/100

The gist: The paper investigates advanced optimization techniques—specifically comparing algorithms like Lion and Muon—designed to handle complex optimization landscapes corrupted by heavy-tailed noise,

Key concepts

Stochastic Frank-Wolfe
This is a foundational optimization approach discussed in the paper that provides a unifying framework for analyzing various optimizers. It helps understand the dynamics of state-of-the-art algorithms by providing a common language for analysis.
Heavy-Tailed Noise
Unlike traditional assumptions of normal noise, heavy tails account for real-world data where occasional massive outliers or spikes occur. Handling this noise is critical for building resilient AI systems that don't fail due to unexpected inputs.
KKT Point
A specific condition and solution type that the algorithms aim to reach during training. Discussing convergence to a KKT point helps define precisely what kind of stable solution the optimization process is achieving.
Convergence Guarantees
The theoretical support provided by the paper, establishing how well and how quickly an algorithm approaches its optimal solution. These guarantees are measured using metrics like the Frank-Wolfe gap.

Terminology

Summary

The paper investigates advanced optimization techniques—specifically comparing algorithms like Lion and Muon—designed to handle complex optimization landscapes corrupted by heavy-tailed noise, such as those encountered in stochastic gradient descent. By evaluating these methods across diverse experimental settings, including synthetic function minimization and multi-class image classification, the research demonstrates how incorporating robust noise handling mechanisms can significantly improve model performance compared to standard approaches.

Performance on Synthetic Function Minimization

The study rigorously evaluates optimization convergence using synthetic functions F(x) = 2 x squared and F(X) = 2 X squared. These tests assess algorithm stability under varying levels of noise corruption, comparing performance across Normal noise (e.g., n=1 vs. n=30) and Pareto noise (e.g., p=1.5 and p=2.5). For instance, in the case of Normal noise on F(x), the comparison between algorithms like Lion and Lion++ across different sample sizes shows systematic differences in convergence behavior, as detailed in Table 4. Similarly, Table 5 confirms that these comparisons hold for the higher-dimensional function F(X). The results highlight how algorithm choices and hyperparameter settings—such as adjusting the learning rate or gradient clipping threshold—are critical for achieving optimal convergence regardless of the noise distribution encountered.

Multi-class Image Classification Results (CIFAR-10)

The algorithms are applied to a practical deep learning task: multi-class image classification using a ResNet18 model trained on the CIFAR-10 dataset. Figure 9 presents the test loss and test accuracy curves over 200 epochs, demonstrating clear comparative performance. The findings show that LION++ consistently outperforms Lion in terms of test loss throughout training. Furthermore, M UON ++ exhibits an overall improvement in test loss compared to Muon, with this advantage being particularly noticeable during the first half of training. In terms of accuracy, both L ION ++ and M UON ++ show improved performance relative to their baselines, with L ION ++ demonstrating a more consistent and pronounced advantage over Lion during the final epochs.

Hyperparameter Optimization Strategies

To maximize performance across both optimization tasks, extensive hyperparameter searches were conducted. For the ResNet18 training on CIFAR-10, the search space covered multiple parameters:

  • Learning rate: 1e-1, 5e-2, 1e-2, 5e-3, 1e-3, 5e-4, 1e-4, 5e-5, 1e-5

  • Gradient clipping threshold: 1, 2, 3, 4, 5, infinity

  • Weight decay: 1e-1, 1e-2, 1e-3

Table 7 reports the specific hyperparameters selected by grid search for both LION++ and M UON ++. These optimized settings—including a learning rate of 5e-2 and gradient clipping thresholds of infinity or 5 —are crucial for establishing the best balance between optimization speed and stability across different noise regimes.

Impact of Noise Distribution on Convergence

The study provides visual evidence comparing algorithm behavior under various noise types. By presenting figures for Normal noise (e.g., Figure 6) versus heavy-tailed Pareto noise (e.g., Figure 7), the paper illustrates the robustness required of modern optimizers. For example, comparing Pareto (p=1.5) and Normal noise in optimization plots reveals how different stochastic processes affect the convergence curve F(X t). The visualization of these differences across varying sample sizes (e.g., n=2 versus n=30) underscores the necessity for algorithms that can maintain stable gradients even when faced with highly volatile, heavy-tailed noise sources.

Improvements for AI systems

Based on the provided scientific paper excerpts, which extensively compare advanced optimization algorithms (Lion vs. Lion++, Muon vs. Muon++) under various noise conditions and demonstrate their application to complex tasks like image classification, I can suggest several critical improvements for current AI systems.

These improvements focus on enhancing robustness, generalization capabilities, and efficiency during the training phase.


The most significant finding is the superior performance of the enhanced optimizers (Lion++ and Muon++) compared to their baselines, especially when facing non-Gaussian noise distributions.

  • Improvement: Integrate an adaptive optimization module that dynamically adjusts gradient updates to mitigate the negative impact of heavy-tailed noise (e.g., Pareto distribution) and significant outliers in the loss landscape.

  • Mechanism: The system should employ a mechanism similar to Lion++ or Muon++, which effectively dampens large, outlier gradients that characterize heavy-tailed noise. This is superior to standard optimizers that assume Gaussian noise (Normal distribution).

  • Improved System Capability: The resulting AI model will exhibit significantly enhanced robustness when deployed in real-world environments where data collection or measurement processes are prone to extreme, non-Gaussian outliers (e.g., sensor readings, noisy financial data, adversarial attacks). The training process will become less susceptible to catastrophic forgetting or divergence caused by rare but large gradient spikes.

The comparative results across different dimensions (d=1 and d=1000) suggest that the optimization method's effectiveness scales with the complexity of the parameter space.

  • Improvement: Develop a meta-optimization layer that assesses the effective dimensionality of the current loss landscape (d eff) during training. If d eff is high (approaching 1000), the system should automatically switch to optimization parameters optimized for high-dimensional spaces.

  • Mechanism: This module would dynamically adjust hyperparameter sets—specifically the learning rate decay schedule and gradient clipping thresholds—to maintain optimal convergence speed and stability in complex, high-dimensional parameter spaces.

  • Improved System Capability: The AI system can efficiently train massive models (e.g., large language models or deep generative networks) with a guaranteed level of stability and convergence speed, regardless of whether the underlying data manifold is simple (d=1) or extremely complex (d=1000).

The tables (Table 4 and Table 5) show meticulous hyperparameter tuning across various noise types and dimensions. This suggests that a single, static set of hyperparameters is insufficient.

  • Improvement: Implement an automated, multi-objective hyperparameter search routine (beyond simple grid search) that profiles the expected noise characteristics of the target deployment environment before training begins.

  • Mechanism: The system should ingest metadata about potential data corruption sources (e.g., Expected Noise Profile: Pareto, p=1.5, or Expected Noise Profile: Normal, sigma=0.8 ). It would then use this profile to select the optimal combination of learning rate decay schedule, weight decay strength, and gradient clipping thresholds (LR, WD, GC).

  • Improved System Capability: This leads to Adaptive Training Configuration. The AI system can achieve state-of-the-art performance on unseen datasets simply by knowing the statistical nature of the expected noise, drastically reducing manual tuning time and increasing reliability across different domains.

The application to ResNet18 on CIFAR-10 demonstrates that these optimized methods improve overall performance metrics (Test Loss and Test Accuracy).

  • Improvement: Integrate the LION++ or M UON++ optimization frameworks directly into the initialization phase of transfer learning pipelines.

  • Mechanism: When fine-tuning a pre-trained model (like ResNet18), instead of relying on standard SGD/Adam, the system should use the advanced optimizers. Specifically, it should prioritize the stability benefits derived from handling complex noise profiles during fine-tuning layers.

  • Improved System Capability: This results in higher generalization performance and faster convergence during transfer learning. The AI system can achieve superior accuracy (e.g., consistently surpassing 92%+) on niche tasks with less labeled data and fewer epochs, maximizing the utility of pre-trained weights while maintaining robustness against domain shift noise.

Sources

Related papers