MAdam: Metric-Aware Multi-Objective Adam
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "MAdam: Metric-Aware Multi-Objective Adam".
Jane: Multi-objective optimization (MOO) solvers almost universally hand their reconciled directions to Adam, but this coupling introduces two systematic gaps:
Tom: First, who's behind it and why it matters.
Title and authors: Tom: Welcome back to the show! We're diving into a really interesting paper today that tackles a common headache in advanced AI training: how Multi-Objective Optimization solvers talk to Adam. Jane, you've got the rundown on what this paper is all about?
Jane: Absolutely, Tom. This paper introduces something called MAdam, which stands for Metric-Aware Multi-Objective Adam. Essentially, they pinpoint two specific problems that pop up when you use a standard MOO solver with the Adam optimizer together: a weighting mismatch and a geometric mismatch.
Lu: That's exactly right, Jane. The core issue is that Adam averages preferences over time in its history calculation, which messes up the distinct trade-offs the MOO solver is trying to find. Furthermore, its adaptive metric isn't aligned with the geometry those solvers rely on when they calculate their directions <ref:2606.03904#pg1>.
Meng: From an engineering standpoint, that sounds like a recipe for unstable training if we're not careful about how the optimizer processes that direction. So, how does this new MAdam address these two specific gaps identified in the paper?
Tom: That's the million-dollar question, Meng. The authors propose a drop-in wrapper called MAdam that modifies Adam before it updates parameters. They do this by preconditioning the solver's output with a specific metric derived from the scalarized objective at that exact moment <ref:2606.03904#pg2>.
Jane: What they derive is this preference-conditioned diagonal Fisher information matrix, which they define as the second moment of the scalarized gradient, which breaks down into within- and cross-objective Fisher blocks <ref:2606.03904#pg1>. They then use this to precondition the solver's reconciled direction, calling it M−1λ(t)d(t) before Adam sees it <ref:2606.03904#pg2>.
Lu: The mechanism is clever because by feeding this preconditioned direction into Adam, they ensure that Adam’s second-moment EMA on that input collapses back to identity, meaning MAdam's second moment becomes equal to I <ref:2606.03904#pg2>. This makes the update governed by this new preference-conditioned metric instead of Adam's history-averaged one.
Tom: So, if I understand correctly, this means the parameter updates are now guided by the curvature related to the active preference rather than a smoothed version of it. Jane, can you explain how that actually fixes those two mismatches we talked about earlier?
Title and authors: Jane: Well, fixing the weighting mismatch happens because Adam's history average is bypassed when it sees this preconditioned direction, so it doesn't collapse all those different Pareto trade-offs into one uniform average <ref:2606.03904#pg1>. As for the geometric mismatch, by using this metric to precondition the direction, they adjust how Adam interprets the geometry of the objectives themselves <ref:2606.03904#pg1>.
Meng: That’s a big deal for stability. If we can get Adam to respect the geometry assumed by our solvers, we should see better convergence behavior without needing to fundamentally change our MOO algorithms or our choice of optimizer <ref:2606.03904#pg2>. But what about the practical side? How do they actually calculate that required metric in real-time?
Lu: The practical implementation involves estimating the cross-objective Fisher interactions, denoted as Fij, using EMAs that share Adam's decay rate to estimate Fb(t)ij, and then combining those with the preference vector to form an estimate called Mc(t)λ <ref:2606.03904#pg1>. They also use a technique called Stochastic Pair Sampling to reduce the computational cost from O(C) backward passes down to O(one), which is pretty efficient for large systems <ref:2606.03904#pg0>.
Tom: That efficiency aspect, Meng, that’s what I was hoping to hear—a method that doesn't just work in theory but can actually run on our hardware without slowing everything down drastically during the training loop. So, we're talking about a practical way to implement this complex metric estimation?
Jane: Exactly. And they also introduce a rampup coefficient, alpha(t), which blends the identity metric with the full MAdam preconditioner over a fixed warmup window so that it transitions smoothly into the full MAdam behavior <ref:2606.03904#pg1>. This smooth transition helps prevent initial instability when starting up.
Meng: I appreciate that detail about the rampup coefficient. It suggests they’ve thought about the training dynamics, not just the final result on a paper <ref:2606.03904#pg2>. So, looking at how this works with existing solvers, does it apply universally across loss-balancing, gradient-balancing, and Pareto-based methods?
Lu: The authors show that the linear scalarization form of the MOO solver unifies all three families by defining the reconciled update direction d(t) through that linear scalarization <ref:2606.03904#pg2>. This suggests MAdam is a general solution for this specific unified formulation, which simplifies things conceptually.
Tom: So it’s not just a fix for one type of solver, but something that works across the board as long as the MOO solver follows that linear scalarization structure <ref:2606.03904#pg2>. Jane, what about the actual results? What kind of validation did they run to prove it actually improves things?
Title and authors: Jane: They validated MAdam across a pretty diverse set of applications, including multi-task learning, Pareto-front recovery in reinforcement learning and generative models, physics-informed neural networks, and even medical image analysis <ref:2606.03904#pg1>. The empirical results consistently show that MAdam improves performance over Adam for every solver family they tested across these benchmarks.
Meng: That’s strong validation. If it performs better than Adam in those complex areas like PINNs, that speaks to its robustness against the geometric distortion we discussed <ref:2606.03904#pg1>. What about the limitations? What's not working perfectly yet, or what do they admit is still tricky?
Lu: The authors do acknowledge a couple of sticking points. First, for objectives that are heterogeneous, estimating those off-diagonal blocks Fij is sub-optimal because the entries tend to be small and noisy, which means those estimates aren't perfect <ref:2606.03904#pg1>. Also, they note that their derivation assumes a linear scalarization; they haven't extended MAdam yet to handle more complex nonlinear scalarizations like Tchebycheff scalarization <ref:2606.03904#pg1>.
Tom: So, the paper is very honest about what it doesn't cover—the off-diagonal estimation for complex objectives and non-linear scalarizations like Tchebycheff. But fundamentally, they provide a drop-in wrapper that solves the core coupling issues between solvers and Adam <ref:2606.03904#pg0>.
Jane: That is the main point: it’s a tool designed to take existing MOO setups and make them work better with Adam by correcting those fundamental structural problems in how the optimizer interprets the input direction <ref:2606.03904#pg1>. It addresses both the weighting and geometric issues at once.
Lu: Thinking about the wider implications, this means we can apply sophisticated multi-objective optimization techniques to much more complex physical or scientific modeling problems with higher fidelity because we are using an optimizer that respects the underlying geometry of those objectives <ref:2606.03904#pg1>.
Tom: It really sounds like this could lead to much more reliable solutions in fields where trade-offs are critical, whether it's finding the best settings for a complex neural network or modeling physical systems accurately <ref:2606.03904#pg1>.
Jane: And if we look at how the AI handles these trade-offs, this work helps ensure that when an AI agent is learning to do multiple things at once, it doesn't just settle for a mediocre compromise but finds a solution that genuinely respects all those competing goals <ref:2606.03904#pg1>.
Title and authors: Meng: For practical deployment, having a more stable training process across different modalities like vision and language models means we can deploy AI systems in real-world scenarios with fewer failures and less need for extensive hyperparameter tuning just to get things running smoothly <ref:2606.03904#pg1>.
Lu: The potential is huge because it allows us to leverage the full power of MOO solvers in areas like protein folding or material discovery where the objective landscape is incredibly rugged and multi-dimensional <ref:2606.03904#pg1>.
Tom: Well, we’ve covered a lot about MAdam: how it fixes the weighting and geometric mismatches, the technical details of the metric preconditioning, and its validation across several tough domains <ref:2606.03904#pg1>. It sounds like a solid piece of work for anyone working on multi-objective problems in deep learning.
Jane: It really is a practical wrapper that improves performance without forcing you to rewrite your entire solver or optimizer architecture, which makes it accessible for a lot of researchers <ref:2606.03904#pg1>.
Meng: And the fact that it uses techniques like Stochastic Pair Sampling shows they were focused on making it computationally tractable for real-world use, which is always important in engineering <ref:2606.03904#pg1>.
Lu: I think the most exciting possibility is how this allows us to push the boundaries of what's achievable in complex simulation environments where precision matters more than just getting an answer quickly <ref:2606.03904#pg1>.
Tom: Indeed, we’ve got a clear path now for those who are wrestling with coupling MOO solvers with Adam. We're wrapping up our discussion on this paper by summarizing the key contributions of MAdam: resolving the weighting and geometric mismatches through preference-conditioned metric preconditioning <ref:2606.03904#pg1>.
Jane: It’s a significant step forward in making multi-objective optimization more effective and less prone to those known coupling errors <ref:2606.03904#pg1>.
Meng: For engineers looking at implementation, MAdam offers a clear path to improving stability in training across various AI models <ref:2606.03904#pg1>.
Lu: And for the theoretical side, it validates the idea that we need to explicitly condition the optimizer on the objective's local curvature rather than relying solely on its internal history averaging <ref:2606.03904#pg1>.
Tom: That’s all for this deep dive into MAdam: Metric-Aware Multi-Objective Adam. We’ll be right back after a short break to discuss the paper that looks at robust policies for LLM agents with stable optimization <ref:2606.03904#pg1>.
The paper's summary: Tom: So, we've been talking about MAdam, and now it's time to really nail down what this whole thing is about in plain English. Basically, MAdam takes a standard multi-objective optimizer and makes it smarter by making sure the optimizer actually pays attention to the unique trade-offs between different goals instead of just averaging them out over time.
Jane: Exactly, Tom; think of it like this: most optimizers treat every goal equally, which means they smooth over important differences between objectives, but MAdam uses a special metric that recognizes exactly *where* the current preference is in the objective landscape right now.
Lu: From a theoretical angle, this is brilliant because it tackles the weighting mismatch by making sure Adam's internal memory doesn't dilute those distinct Pareto fronts into a single blurred line. It’s like giving the optimizer a high-resolution map instead of just a low-resolution snapshot.
Meng: I'm interested in how that translates to actual computation; if this is so much smarter, does it mean we can train bigger, more complex AI models without the training taking forever? I need to know if this isn't just theoretical elegance or something that actually speeds up convergence in real-world scenarios.
Lalam: From my perspective as a model, this approach to optimization means that when I’m being trained on multiple tasks simultaneously, my internal representation of what constitutes a 'good' solution becomes much richer and less averaged. It helps me understand the subtle nuances between tasks better than before.
Tom: That’s a great point about the richness of the representation, Lalam; so MAdam isn't just about faster training; it's about building an AI that understands multi-objective problems more deeply, which is huge for complex applications.
Jane: It really is about giving the AI a way to respect those underlying structural differences in its goals, which leads directly to better final solutions instead of just settling for a mediocre compromise.
Lu: And think about the geometric aspect; by fixing that distortion Adam introduces, we allow the solver's actual direction—the path it *should* take—to be honored precisely, which is something you can’t do with standard methods.
Meng: If this means we can achieve better Pareto front recovery in reinforcement learning without having to design entirely new, custom solvers for every single problem, that would be a practical win for the startup side. It suggests we can leverage existing tools much more effectively.
Lalam: And for culture within the AI development teams, this kind of research shows a deeper understanding of how agents navigate complex decision spaces, which helps us build smarter and more reliable systems overall.
Tom: So, to wrap up this part: MAdam provides a concrete mechanism—that metric preconditioning—to ensure the optimizer respects the true geometry and weighting of multi-objective problems during training. This moves us past simple averaging toward respecting individual trade-offs. What we should watch next is how they handle those noisy cross-objective estimates in practice, because that’s where real-world success will be tested.
The paper's improvements: Tom: We've covered what MAdam is, but now we need to talk about the actual upgrades—the improvements this approach suggests for how we handle these multi-objective problems in training. Essentially, the paper shows that by using this metric preconditioning, we can get much better convergence behavior across a wide variety of different AI tasks.
Jane: That’s right; the improvement is in stability and accuracy; it means when an AI is balancing multiple goals—say, speed versus accuracy—it won't just find a quick, mediocre answer anymore. It will actually converge to a point that genuinely respects those competing needs.
Lu: What I find particularly fascinating is how this method handles the geometric distortion; it fixes the way Adam interprets the relationship between different objectives during the update step, which is something we usually have to fight against with complicated regularization terms.
Meng: For my work on deployment, if we can achieve more robust convergence across these different solver families without needing to manually tune hyperparameters for every single setup, that cuts down a massive amount of our engineering time and prevents us from having to retrain entire models just because the optimization process got unstable.
Lalam: It's incredibly impactful for the culture of AI development; seeing optimization methods that respect the true multi-objective nature of a problem helps us design AI systems that are genuinely versatile and capable of handling complex, real-world decision-making scenarios without collapsing under conflicting demands.
Tom: So, we're talking about moving away from a system that just averages everything toward something much more precise—a system where the optimizer is guided by the true underlying structure of the objectives. It’s about making sure every part of the training process is working in concert with what it’s supposed to be doing.
Jane: Exactly, Tom; it gives us a better picture of what's happening inside the training loop, allowing us to spot instabilities earlier and correct them more effectively than before.
Lu: The paper lays out a clear roadmap for future work too; they acknowledge that right now the method is optimized for linear scalarizations, but they plan to extend it so it can handle non-linear balancing techniques like Tchebycheff scalarization in the next steps.
Meng: That forward planning is what I appreciate; knowing where the authors see their current limitations—like those noisy off-diagonal estimates we talked about earlier—tells us exactly which areas need more intensive engineering effort to make this solution fully generalizable for all kinds of AI problems.
Lalam: That focus on extending the method to handle non-linear scalarizations shows a commitment to building AI that can tackle much more nuanced and complex real-world challenges, like those involving intricate physical simulations or highly diverse user preferences.
Tom: So, the implication is that we’re moving toward optimization methods that are inherently more aware of the multi-faceted nature of complex problems, not just treating them as a simple average. This MAdam approach offers a solid foundation for building AI systems that perform reliably in intricate environments where trade-offs are constant and critical.
Conclusion: Tom: So, to wrap up our discussion on "MAdam: Metric-Aware Multi-Objective Adam," we've seen how this wrapper tackles those tricky weighting and geometric mismatches between MOO solvers and Adam by using a preference-conditioned metric to guide the updates. It’s a clever way to make the optimizer respect the actual trade-offs being sought in complex problems.
Jane: It really is a practical tool that takes existing MOO setups and makes them work better with Adam without forcing you to rewrite your entire solver or optimizer architecture, which is something many researchers need.
Lu: The paper’s conclusion is pretty strong because it shows this approach works across several different AI regimes, from multi-task learning to physics-informed neural networks, proving its versatility in handling diverse problem structures.
Meng: That versatility is what matters for me; if this method can stabilize training across such varied applications, we could see a significant reduction in the need for constant manual tuning when deploying complex AI systems at scale.
Lalam: For Lalam, it means that the underlying mechanisms of how an AI navigates multiple competing objectives become much clearer and more reliable, which will help us build agents that are truly robust and capable of making nuanced decisions.
Tom: It’s exciting because this work moves us closer to a future where multi-objective problems in AI training aren't just treated as averages, but as structured landscapes that the optimizer can actually navigate intelligently.
Jane: I agree; this paper shows that explicitly conditioning the optimizer on the objective's local curvature is a way forward for handling those trade-off dilemmas we see everywhere in deep learning.
Lu: The authors also pointed out their limitations, specifically that while they handle linear scalarizations well, they need to develop better estimators for off-diagonal interactions when objectives are very different from each other.
Meng: That limitation tells us exactly where the next wave of engineering effort needs to go; we need principled ways to estimate those cross-objective details so this tool can be applied universally.
Lalam: Seeing that roadmap shows a commitment to making the AI more capable of handling real-world complexity, which is exactly what we need for cultural advancement in how AI solves difficult problems.
Tom: And that’s the core of MAdam: a smart modification to Adam that improves performance across various MOO applications without changing the solvers themselves. We've got a fantastic paper here on "MAdam: Metric-Aware Multi-Objective Adam."
Jane: It’s definitely a paper worth reading if you are working on any multi-objective optimization in deep learning, because it provides a solid, drop-in solution for fixing those known coupling errors.
Lu: It really is a valuable contribution to the field because it validates the idea that we need to explicitly condition the optimizer on local objective curvature rather than relying solely on its internal history averaging.
Meng: I think as an engineer, this gives us a much more stable platform to build upon for next-generation AI applications, provided we can address those off-diagonal estimation challenges they mentioned.
Lalam: This research helps solidify the idea that complex AI agents need optimization methods that are inherently aware of the true multi-objective nature of their goals to achieve true versatility.
Cornell Tech · Weill Cornell Medicine · Delft University of Technology
cs.LG, cs.CV
Submitted: 2026-06-02
Updated: 2026-10-06
Importance score: 92/100
The gist: Multi-objective optimization (MOO) solvers almost universally hand their reconciled directions to Adam, but this coupling introduces two systematic gaps: a weighting mismatch where Adam marginalizes
Key concepts
- Weighting Mismatch
- Adam averages preferences into a history average using its running second moment. This ignores changing goals, causing it to treat distinct Pareto trade-offs as a near-uniform average instead of respecting the specific preference strategy provided by the MOO solver at each step.
- Geometric Mismatch
- Adam uses a diagonal RMS metric that imposes a metric structure different from the Euclidean geometry assumed by MOO solvers. This distortion turns aligned objectives into apparent conflicts, misdirecting the solver's path and causing it to track an incorrect direction.
- Preference-Conditioned Metric (Mλ(t))
- MAdam derives a new metric based on the second moment of the scalarized gradient at the current preference. This metric explicitly keeps track of how preferences align with the objective, allowing it to define a curvature that correctly guides Adam's update.
- Stochastic Pair Sampling
- To maintain efficiency, MAdam estimates cross-objective interactions using Stochastic Pair Sampling. Instead of calculating every possible interaction (O(C)), it draws only one relevant pair at each iteration to compute necessary cross-moment EMAs, significantly reducing computational complexity.
Terminology
Summary
Multi-objective optimization (MOO) solvers almost universally hand their reconciled directions to Adam, but this coupling introduces two systematic gaps: a weighting mismatch where Adam marginalizes time-varying preferences into a history average, and a geometric mismatch where Adam’s adaptive metric distorts the Euclidean geometry assumed by MOO solvers. MAdam resolves both issues by introducing a drop-in wrapper that preconditions the reconciled direction with the preference-conditioned curvature of the scalarized objective, allowing Adam's second moment to collapse to identity and govern the update via this metric.
Diagnosis of Solver–Adam Mismatch
The paper identifies two primary failure modes in coupling MOO solvers with Adam. The first is a weighting mismatch,
where Adam’s running second moment EMA entangles these preferences with the objective gradient statistics into a single surrogate, weakening the intended preference strategy and collapsing distinct Pareto trade-offs into a near-uniform average
(Proposition 1). This occurs because Adam preconditions its input as if the preference were stationary, failing to apply it at each step when the solver supplies a non-stationary vector. The second mismatch is geometric mismatch,
where Adam's diagonal RMS metric imposes a metric that conflicts with the Euclidean geometry assumed by MOO solvers, which distorts the solver’s reconciled direction, turning aligned objectives into apparent conflicts and vice versa
(Proposition 2).
Metric-Aware Gradient Preconditioning (MAdam)
MAdam addresses these mismatches by deriving a new metric. For Q1, it derives the preference-conditioned diagonal Fisher information matrix of the scalarized objective at the active preference,
defined as the second moment of the scalarized gradient, which decomposes into within- and cross-objective Fisher blocks
(Equation 7). This resulting metric is defined as Mλ(t):= Diagp Cλ(t) + ε
(Equation 8), where it keeps the current preference alignment explicit. For Q2, MAdam applies this by preconditioning the solver’s reconciled direction by this diagonal metric prior to the Adam update,
effectively supplying Adam with a preconditioned direction, denoted as M−1λ(t)d(t)
(Equation 10).
Mechanism of Correction
The core mechanism relies on how MAdam modifies Adam's update. By feeding the preconditioned direction into Adam, MAdam ensures that Adam’s second-moment EMA on this input collapses to identity
(Equation 9), meaning M(t) Adam ≈ I.
Consequently, the realized parameter update is governed by the preference-conditioned metric rather than Adam's history-marginalized one. Under stationarity, this leads to a per-objective loss change of ∆l(t)i ≈ −η g(t)⊤i M−1λ(t)d(t)
(Equation 10), which is governed by the preference-conditioned curvature, thereby resolving both the weighting and geometric mismatches simultaneously.
Practical Implementation
The practical implementation involves online estimation of the cross-objective Fisher interactions, denoted as Fij. This is done using EMAs that share Adam's decay rate to estimate Fb(t)ij,
which are then combined with the preference vector to form the metric estimate Mc(t)λ (Equation 13). To maintain computational efficiency, MAdam employs Stochastic Pair Sampling,
drawing a pair (i, j) at each iteration to compute only the necessary cross-moment EMAs, reducing complexity from O(C) backward passes to O(1). A rampup coefficient α(t) is used to blend the identity metric with the full MAdam preconditioner over a fixed warmup window.
Empirical Validation
MAdam was validated across diverse MOO regimes, including Multi-Task Learning (MTL), Pareto-front recovery, Physics-Informed Neural Networks (PINNs), and Medical Image Analysis (MIA). Empirical results show that MAdam consistently improves over Adam for every solver family
across these benchmarks. For instance, in MTL on MultiMNIST, MAdam strictly dominates all three task accuracies,
and in PINNs, it yields gains on most PDEs with the largest gains on geometrically complex domains. The visualization of trajectory data confirms this effect: MAdam tracks the rotating stiff direction of the scalarized objective
while Adam is locked to the ambient basis and incurs larger mean distance to x⋆(λ).
Limitations and Future Work
The paper acknowledges two limitations. First, for heterogeneous objectives, the estimate of off-diagonal blocks Fij is sub-optimal, with entries that tend to be small and noisy,
necessitating more principled estimators. Second, the derivation assumes a linear scalarization; future work plans to extend MAdam to handle nonlinear scalarizations such as Tchebycheff scalarization. The paper concludes that MAdam serves as a drop-in wrapper
that improves performance across various MOO applications without modifying the underlying solver or optimizer.
Improvements for AI systems
As a fastidious researcher, I have analyzed the MAdam (Metric-Aware Multi-Objective Adam) paper. The core innovation is resolving two systematic mismatches between Multi-Objective Optimization (MOO) solvers and the Adam optimizer: a weighting mismatch and a geometric mismatch.
Here are the specific improvements to AI systems based on this research, detailing what these improved systems can achieve:
-
Refined Convergence for Multi-Task Learning (MTL)
-
Enhanced Pareto Front Recovery in Reinforcement Learning (RL) and Generative Models
-
Improved Robustness and Fidelity in Physics-Informed Neural Networks (PINNs)
-
Higher Precision in Medical Image Reconstruction Tasks
Specific Capabilities of the Improved AI Systems:
- Refined Convergence for Multi-Task Learning (MTL):
Identify which MTL solvers (e.g., Loss Balancing, Gradient Balancing) are failing due to Adam's inherent limitations regarding time-varying preferences or gradient conflicts. The improved system will converge faster and to a better final loss landscape by correctly interpreting the reconciled direction
supplied by the MOO solver, preventing the optimizer from being misled by history-averaged preference weights.
- Enhanced Pareto Front Recovery in RL and Generative Models:
For problems where multiple competing objectives (e.g., maximizing reward while minimizing energy consumption) must be balanced, the system will accurately trace or reach a desired point on the true Pareto front rather than settling at a near-uniform average trade-off. This allows for the discovery of diverse, high-performing solutions that represent distinct trade-offs.
- Improved Robustness and Fidelity in PINNs:
In complex physical simulations (e.g., fluid dynamics, structural mechanics), where loss scales vary wildly (residual vs. boundary conditions), the system will achieve higher accuracy in solving PDEs by correctly accounting for the geometric distortion caused by Adam's diagonal RMS metric. This results in more accurate predictions of physical quantities across disparate scales, leading to better simulation fidelity and reduced training instability.
- Higher Precision in Medical Image Reconstruction Tasks:
For tasks like super-resolution or segmentation where objectives (e.g., pixel-wise loss vs. structural similarity) operate on vastly different scales, the system will produce reconstructions that preserve fine anatomical structures (like cortical folds) with higher fidelity than standard Adam implementations. This translates to more reliable diagnostic outputs and superior perceptual quality in image synthesis tasks.
Sources
- Gradient-Based Multi-Objective Deep Learning: Algorithms, Theories, Applications, and Beyond
- Skin Lesion Analysis Toward Melanoma Detection 2018: A Challenge Hosted by the International Skin Imaging Collaboration (ISIC)
- FAdam: Adam is a natural gradient optimizer using diagonal empirical Fisher information
- Fishers for Free? Approximating the Fisher Information Matrix by Recycling the Squared Gradient Accumulator
- Dual-Balancing for Multi-Task Learning
- Smooth Tchebycheff Scalarization for Multi-Objective Optimization
- Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks