Momba: Network Modernization Improves Multi-Objective Reinforcement Learning
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Paper Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Momba: Network Modernization Improves Multi-Objective Reinforcement Learning".
Jane: The paper was written by Adam Štafa, Santeri Heiskanen, Petr Novotný and Joni Pajarinen from.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: So we're looking at a paper from Masaryk University in Brno and Aalto University in Finland, and the title makes a bold claim — modernizing the network, not the algorithm, is what improves multi-objective reinforcement learning. The name Momba is playful, but the argument behind it is very pointed. I think the title is the whole thesis in seven words.
Jane: It really is. In single-objective RL, there's been a wave of work showing you can get enormous gains just by changing the neural network architecture — normalization, better critics, residual connections — while keeping the learning rule untouched. And this group looks at the multi-objective side and says, why is that wave missing here? They treat the field's favorite algorithms as fixed and ask what the network is costing them.
Lu: The author mix tells part of the story. Adam Štafa and Petr Novotný work on the formal and theoretical side at Masaryk, while Santeri Heiskanen and Joni Pajarinen bring the Aalto robotics and machine learning perspective. You can see both flavors — rigorous benchmarking and careful ablation — all through the paper. That combination matters because the claim is empirical, so the experimental discipline has to be solid.
Meng: And that discipline is exactly what the title demands. They're not proposing a new way to handle multiple objectives. They take an existing algorithm called CAPQL and give its function approximators a serious upgrade, and the performance jumps. So "modernization" is doing a lot of work in that title.
Tom: Right, and it's a provocation to a field that has spent years on clever preference sampling and specialized update rules. This paper says the bottleneck might be much more mundane — simple feedforward networks that aren't expressive enough to represent value functions across different trade-offs. If they're right, a lot of recent MORL machinery has been polishing the wrong part of the pipeline.
Lalam: What I find exciting is the wider implication. If architecture modernization transfers from single-objective to multi-objective RL this cleanly, then a whole family of problems — robot control, treatment planning — could get better without any new theory, just by borrowing what deep RL already learned about building networks. That's a cheap win for a lot of application areas.
Jane: The title "modernization" is exactly the right word, then. It's like upgrading the engine instead of changing the route you drive. And since the authors are building directly on SimbaV2, a recent single-objective architecture, the next question is what exactly they borrowed and what they had to adapt.
Tom: That's our next stop — the paper's summary, and the core question of why multi-objective RL has been leaving performance on the table.
Summary: Tom: We've established the provocation — architecture over algorithm. Now the paper's summary lays out the problem clearly: multi-objective RL wants a set of policies that balance conflicting objectives, and most methods handle that by conditioning a single policy on a preference vector, which is the trade-off knob. That framing immediately tells you why representation matters.
Jane: And the key observation is that even though the optimal policy can look very different depending on the trade-off, the standard choice is still a plain feedforward network conditioned on that preference. You're asking one network to cover the entire Pareto front, and then you give it a fairly weak function approximator to do that job. That mismatch is where the paper starts.
Lu: Right, and they connect this to a real gap in the literature. Single-objective RL has a string of recent papers showing that normalization and distributional critics improve sample efficiency and final performance, with analyses of why they help — better conditioning of the optimization, less overfitting to early data. The multi-objective side kept its simple feedforward networks, and this paper is essentially importing that whole toolbox into a field that hasn't touched it.
Meng: So what they do is take an entropy-regularized MORL algorithm called CAPQL, which is essentially the multi-objective cousin of Soft Actor-Critic, and they bolt on three things from the recent single-objective toolbox: observation and feature normalization, weight normalization, and a distributional critic that models the distribution of returns instead of just their expected value. That last part turns out to be the most interesting.
Tom: And the punchline of the summary is that these changes substantially improve the quality of the solution sets without requiring major changes to the underlying algorithm. That's a strong statement. It means the gains come from representation, not from new learning rules.
Jane: I also like that they framed it as an open question first — could MORL benefit from these advances? — and then actually tested it. A lot of papers assume the answer before running the experiments. The summary is refreshingly honest about what they did and what they found.
Lalam: There's a bigger pattern here. In deep RL, we keep rediscovering that how you parameterize the problem matters as much as the objective you optimize. This paper extends that lesson to a field that was overdue for it, and if the gains hold across tasks, it changes where researchers should spend their effort.
Tom: And that leads us to the actual improvements — the three components, and especially the distributional critic, which needed some clever adaptation to work in the multi-objective setting.
Improvements: Tom: We've seen why architecture was the missing piece — now let's get concrete about what the modernization actually contains. The first two pieces are normalizations: they normalize observations with running statistics, they normalize hidden features onto a hypersphere, and they normalize the weights after every gradient update. All of that keeps training stable and prevents the network from overfitting to early experiences.
Jane: And the third piece is the distributional critic. Instead of predicting the expected return, the critic predicts a full distribution over returns, which is known to stabilize learning and improve the conditioning of the optimization. But there's a catch: in multi-objective RL, returns are vectors, and categorical distributional RL was designed for scalar returns.
Lu: That's where the paper's most interesting idea comes in. They notice that the policy only ever sees the critic through a scalarized value — the preference vector dotted with the Q-values. So instead of modeling the multivariate joint distribution of returns, like some earlier work attempted, they model the distribution of the scalarized return directly. One conditioned univariate distribution per preference.
Meng: It's a simpler adaptation, and honestly a pragmatic one. Previous attempts at multivariate distributional RL required kernel functions to measure distance between distributions, or were limited to tabular settings. By targeting linear scalarization — which is already the common assumption in MORL — they sidestep all of that complexity. And in the appendix they show that if you need the non-scalarized vector returns, you can learn the marginal distribution of each component separately, and it performs just as well.
Tom: There's also a practical detail: they normalize each reward component by the maximum return seen so far during training. That matters because different preferences can produce very different return scales, and the categorical critic needs a fixed support interval to work well.
Jane: They also changed how preferences are sampled — once per episode from a uniform distribution, instead of at every timestep like the original CAPQL. That puts them in line with most other MORL work, which is smart because it means the gains can't be credited to a fancy preference selection trick.
Lalam: All together, that's a real philosophy difference. The field has been trying to solve multi-objective RL with better search over preferences. This paper says, give the network a better internal representation and the search problem gets easier on its own.
Tom: And the obvious question now is whether it actually worked. So in the next segment, we look at the first page and the empirical evidence — and the numbers are pretty dramatic.
First page: Tom: We've walked through the three components and the clever scalarized critic. Now the first page of the paper formalizes all of that into two contributions: first, expressive architectures with a distributional critic substantially improve MORL performance without complex preference selection or MORL-specific update rules, and second, a simple adaptation of the categorical critic to learn scalarized return distributions. Both claims are testable, and they test them thoroughly.
Jane: And those two contributions aren't just claims — the results section backs them up. Aggregated over seven continuous control tasks, Momba gets roughly 35 percent higher hypervolume and 16 percent higher expected utility than PGMORL, the runner-up. Against CAPQL, the algorithm it's built on, that's about 132 percent and 35 percent improvement respectively.
Lu: Those are large gaps, especially the comparison to its own backbone. And I appreciate that they report both metrics because hypervolume captures convergence and coverage while EUM captures actual user utility. The two metrics agreeing makes the result more convincing.
Meng: The sample efficiency story is just as strong. In Ant and Humanoid, Momba reaches PGMORL's final performance after only 100,000 and 200,000 steps, while PGMORL was trained for tens of millions of steps. And when they match update-to-data ratios against GPI-LS, the method famous for sample efficiency, Momba beats it using plain uniform preference sampling.
Tom: Then the ablations show where the gains come from. The distributional critic is the most important component — even a plain MLP with the categorical loss matches or beats the fancy Simba architecture with a standard MSE critic. And among the normalizations, observation normalization contributes the most, around 55 percent according to their Shapley value analysis.
Jane: That ablation is the most convincing part of the paper for me. They systematically varied every component, and they showed that parameter count alone doesn't explain the gains — bigger MLPs with the old loss stay bad. It's the combination of representation and loss that unlocks the performance.
Lalam: The broader implication is that a lot of the sophisticated machinery in recent MORL papers — the preference prioritization, the self-consistency losses — may be solving a symptom rather than the cause. If the function approximator is the limiting factor, architecture research deserves a much bigger share of attention in this field.
Tom: And that sets up our final segment — what this means going forward, where the approach falls short, and what we should watch for next.
Conclusion: Tom: So let's wrap this up. The paper took a standard multi-objective algorithm, gave it a modern network with normalization and a distributional critic, and got dramatic gains in both final performance and sample efficiency across continuous control benchmarks. If you work in MORL, you don't need to adopt a new learning rule here — you need a better network under the hood. That message is simple and it travels well.
Jane: And the analysis holds up. The distributional critic carries most of the weight, observation normalization comes second, and the whole thing works without fancy preference selection. That's a clean result, and one that practitioners can pick up almost immediately. I think this will show up in a lot of future MORL codebases.
Lu: But the authors are honest about the limits. They only consider linear scalarization, which is the common assumption in MORL but doesn't cover every setting, like nonlinear or learned scalarization. And the benchmarks are continuous control tasks, mostly with two objectives — only one environment has three. So we don't know yet how this scales to many objectives or to discrete state and action spaces. That's a real open question, and I'd love to see someone run those experiments.
Meng: Still, the practical path is wide open. Anyone using SAC-style methods can take these changes almost off the shelf, and the vector version of the distributional critic widens the reach even further. This is the kind of paper where you read the appendix and immediately start thinking about which of your own problems could benefit from the same treatment.
Lalam: For me, the lasting contribution is that the field should stop treating architectures as an afterthought. This paper doesn't just improve one algorithm — it opens a research direction of modernizing all the existing MORL methods that were built on older network designs. That could shift where a lot of research effort goes over the next few years.
Tom: And that's a great note to end on. The rigorous ablations, the honest limitations, the genuinely practical payoff — this made for a really satisfying discussion. Thanks everyone for the conversation, and we're ready to move on to the next paper.
Jane: Goodbye from all of us, and see you on the next episode.
Adam Štafa, Santeri Heiskanen, Petr Novotný, Joni Pajarinen
cs.LG, cs.AI
Submitted: 2026-08-07
Updated: 2026-08-10
Comments: 21 pages, 10 figures; Accepted to RLC 2026
Code: https://github.com/adamstafa/momba
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 66/100
Key concepts
- Multi-objective reinforcement learning (MORL)
- A type of reinforcement learning where an agent must balance multiple conflicting objectives, like speed and energy use. Instead of a single optimal policy, it seeks a set of policies that trade off between objectives, often guided by a preference vector.
- Distributional critic
- A neural network component that predicts the full distribution of possible returns rather than just the average. This provides richer learning signals and stabilizes training. In this paper, it's adapted to model the distribution of scalarized returns (preference-weighted sums) for multi-objective settings.
- Normalization
- Techniques that adjust the scale or distribution of data or network parameters to improve training stability. The paper uses observation normalization, feature normalization, and weight normalization, which help prevent overfitting and improve optimization conditioning.
- Hypervolume and expected utility
- Metrics used to evaluate the quality of a set of policies in MORL. Hypervolume measures the volume of the objective space covered by the policies, while expected utility measures how well the policies satisfy a user's preferences. Both improved significantly with the proposed method.
Terminology
Summary
Keywords: Reinforcement Learning, Multi-objective Reinforcement Learning, Distributional Reinforcement Learning, Neural Architectures
The paper addresses an underexplored area in Multi-Objective Reinforcement Learning (MORL). While recent advances in single-objective deep RL have shown that improving neural network architectures can yield substantial gains in sample efficiency and asymptotic performance without altering the underlying algorithms,
work in MORL has predominantly focused on algorithmic innovations, leaving the area of architectures underexplored.
The authors note that while the optimal policies and value functions can differ significantly depending on the trade-offs, MORL algorithms commonly represent them with simple feedforward networks conditioned on the trade-off.
This raises the question of whether MORL algorithms could benefit from more expressive function approximators.
The paper proposes Momba, which integrates recent advances in neural network design into the MORL setting: (i) observation and feature normalization, (ii) weight normalization, and (iii) modeling of distributional returns with an entropy-regularized MORL algorithm.
The empirical results across 7 standard continuous control benchmarks demonstrate that these changes substantially improve the quality of the produced solution sets without requiring major changes to the underlying algorithm.
The paper states two primary contributions:
-
Demonstrating that expressive architectures with a distributional critic improve MORL performance: "We showcase that employing more expressive neural network architectures with a distributional critic in a MORL algorithm substantially improves the performance on continuous control tasks without requiring complex preference selection algorithms or MORL-specific update rules.
The authors contrast this with previous MORL works that achieved improvements via
advanced preference selection algorithms (Alegre et al., 2023) or specialized update rules (Yang et al., 2019; Li et al., 2025).They apply principles from single-objective RL research, where
the combination of (i) feature normalization, (ii) weight normalization, and (iii) categorical critic loss can improve the sample efficiency and asymptotic performance of existing reinforcement learning algorithms, without altering them.They also identify
the distributional critic as a major component." -
A simple adaptation of the categorical critic to the multi-objective domain: "We propose a simple adaptation of the categorical critic to the multi-objective domain by directly learning to predict the scalarized return distribution, motivated by the commonplace scalarized expected return objective.
The authors note that previous research applying distributional critics to multi-valued returns
has been limited to tabular cases (Wiltzer et al., 2024), or required specifying a kernel function for measuring distance between distributions (Zhang et al., 2021).Instead, under the prevalent linear scalarization assumption, they
propose a simple approach where the Q-network directly predicts the scalarized return distribution."
Base algorithm — CAPQL: Momba is built on top of CAPQL (Lu et al., 2023), an entropy-regularized algorithm akin to Soft Actor-Critic (SAC).
The authors chose this for two reasons: the entropy bonus ensures that the induced solution sets are strictly convex, promoting numerically stable optimization under linear scalarization,
and SAC has been used as the backbone in many recent works on neural architectures in deep RL.
Architecture — SimbaV2: The authors adopt the SimbaV2 architecture (Lee et al., 2025b), which uses: "(i) l2 (hyperspherical) hidden feature normalization and running statistics observation normalization; (ii) the weights are normalized after every gradient update; and (iii) a distributional critic. Additionally, the architecture features scaling layers and learnable residual connections."
Multi-objective distributional critic: The key methodological innovation is how the categorical critic is adapted to the MORL setting. The authors observe that "the critic predictions only affect the policy optimization via the scalarized Q-value ω T Q(s, a, ω). Thus, we can avoid learning the multivariate return distribution, and propose to use a conditioned distributional critic Zθ(s, a, ω), which learns the univariate distribution of the scalarized returns ω T G(s,a). This is trained by minimizing the cross-entropy loss: LZ(θ) = E(s,a,r,s′,ω)∼D H(Ẑ(ω T r, s′), Zθ(s, a, ω)). The actor then maximizes
E a∼πϕ[E[Zθ(s, a, ω)]] + αH(πϕ(· s, ω))."
Reward normalization: Since the magnitude of returns can change significantly based on the preference,
the authors propose to normalize the rewards component-wise by the maximum returns encountered throughout the training
to bound the Q-values to the distributional critic's support.
Preference selection: To isolate the effect of architectural changes, the authors choose a fairly unsophisticated approach: We sample a new preference at the beginning of each episode from a static uniform distribution over the preference space.
This differs from CAPQL's per-timestep preference sampling, but aligns with other MORL work (Abels et al., 2019; Alegre et al., 2023; Xu et al., 2020; Li et al., 2025).
Benchmarks: The authors use 7 continuous control tasks from Xu et al. (2020), implemented in the MuJoCo physics engine
: Ant, Swimmer, HalfCheetah, Humanoid, Walker2D, and Hopper (with two variants: Hopper-v4 with 2 objectives and Hopper-v3 with 3 objectives). They note that since SimbaV2 targets continuous control domains, we omit evaluation on discrete environments,
and existing discrete MORL benchmarks have small state/action spaces where function approximation is unlikely to be the limiting factor.
Baselines: Four baselines are used: CAPQL (the underlying algorithm without architectural improvements), PGMORL and DPMORL (multi-policy methods training separate policies for each scalarization function), and GPI-LS (which uses General Policy Improvement for preference selection, offering best-in-class sample efficiency
but at high computational cost — 300k timesteps taking more than 120 hours
).
Metrics: The authors use Hypervolume (HV) as their main metric, noting hypervolume can detect improvements in uniformity, spread, and convergence of the solution set,
and additionally report the Expected Utility Metric (EUM). Results are reported using interquartile mean (IQM) and stratified bootstrapped confidence intervals (SBCIs) as recommended by Agarwal et al. (2021).
They also visualize solution sets for qualitative assessment, noting that standard diversity metrics like sparsity are ill-suited for comparing fronts with varying convergence rates.
Overall performance: "The proposed approach outperforms baselines, achieving ∼35% improvement in HV and ∼16% improvement in EUM over the runner-up, PGMORL. When compared to CAPQL, the underlying algorithm behind Momba, we achieve ∼132% and ∼35% improvements in aggregate HV and EUM, respectively." The authors note that DPMORL performed worse than expected based on the original paper, and discuss possible reasons (code discrepancies, unavailable normalization values).
Sample efficiency: In Ant and Humanoid, Momba reaches the final performance of PGMORL after training only for 100k and 200k steps respectively.
While CAPQL could not match PGMORL in any shown environment, Momba outperforms PGMORL in 2 of the 3 environments, and matches it in the challenging 3-objective Hopper environment, while requiring a fraction of the training steps.
Comparing against GPI-LS at similar UTD ratios, Momba can beat the sample efficiency of GPI-LS
— noteworthy because "this result is achieved by training Momba using preferences sampled from a static distribution, whereas GPI-LS uses a sophisticated strategy for preference selection, highlighting the importance of an expressive neural network."
Solution set quality: In Walker2d and Hopper-v4, Momba produces solution sets with good coverage, outperforming other methods.
In HalfCheetah, Momba has a more limited coverage, while still producing a solution set that dominates most of the baselines.
The authors attribute gaps in some solution sets to linear scalarization, noting it is possible that a stronger regularization would be required to reliably obtain policies in these regions.
Distributional critic: Comparing the cross-entropy (CE) loss to standard MSE loss while controlling for architecture and parameter count, the authors find that the proposed approach (Momba=Simba+CE) clearly outperforms the Simba+MSE variant across all critic sizes.
Strikingly, the standard MLP with categorical loss (MLP+CE) matches or even outperforms Simba+MSE, showcasing the effectiveness of the proposed categorical critic adaptation to the multi-objective domain.
The default MLP with MSE loss (a parameter-scaled CAPQL) "retains poor performance, regardless of the critic size, implying that the performance gains are not due to the increase in the parameter counts, but rather stem from the improvements to function approximation and critic loss."
Normalization components: Using Shapley values, the authors find that "observation normalization was the most effective technique on our benchmarks, accounting for ≈55% of the improvements according to the Shapley values. While feature and weight normalization provide more modest improvements, all normalization schemes synergistically improve the performance."
Vectorized vs. scalarized returns (Appendix B): The authors also validated an alternative vector distributional critic that learns marginal distributions of each return component separately, finding that the performance of the two approaches is nearly identical, demonstrating that our proposed adaptation of the distributional critic could also be considered in cases where the vector-values are required.
The paper concludes that "the recent advances in neural network design, such as i) feature & observation normalization, ii) weight normalization, and iii) distributional critic, can improve the performance of an existing MORL algorithm — while keeping the algorithm
close to the original, only changing the preference sampling rate and critic loss. Their
extensive experiments demonstrated that one can achieve competitive performance while matching or improving the sample efficiency of the base algorithm in multi-objective continuous control tasks."
The authors acknowledge several limitations: (1) linear scalarization is assumed — while common in MORL, nonlinear or learned scalarization may be preferable in some settings, and our distributional critic does not assume a specific scalarization form; the constraint arises from TD learning via the Bellman error, which may not hold under nonlinear scalarization
; (2) benchmarks are limited to continuous control tasks, with 2 objectives, with only one with 3 objectives,
making it difficult to assess applicability to discrete settings or scalability with more objectives; (3) more expressive function approximators can increase wall-clock time, though recent systems advances (e.g., JIT compilation) can mitigate this.
Improvements for AI systems
Based on the paper, an AI system can be improved in these ways:
-
Use a distributional critic for scalarized returns instead of learning raw multi-objective value distributions. The Q-network directly predicts the distribution of the scalarized return ωTZ(s, a, ω), trained with cross-entropy loss. This captures return uncertainty under linear scalarization without needing complex kernel-based distributional losses.
-
Apply feature normalization and observation normalization (hyperspherical normalization for hidden features, running statistics for observations) to stabilize training and accelerate convergence in multi-objective continuous control.
-
Apply weight normalization after every gradient update to stabilize the critic and actor training under changing preference vectors.
-
Normalize rewards component-wise by the maximum returns encountered during training, bounding Q-values and keeping them within the distributional critic's support, which handles the varying reward scales across objective preferences.
-
Combine these with an entropy-regularized, SAC-style MORL algorithm (e.g., CAPQL). The entropy bonus maintains strictly convex solution sets and avoids degenerate policy collapse as preferences vary.
-
Sample a new preference uniformly at the start of each episode (rather than per timestep) to align with multi-policy training and reduce computational overhead while still producing a diverse solution set.
The improved AI system (e.g., Momba
) can:
-
Achieve 35% higher hypervolume and 16% higher expected utility than the previous state-of-the-art (PGMORL) across continuous control benchmarks, and 132% higher hypervolume than the base algorithm CAPQL.
-
Match or surpass the performance of sophisticated preference-selection methods like GPI-LS with a much simpler uniform preference sampling, demonstrating that expressive architectures matter more than complex preference scheduling.
-
Attain final performance of strong multi-policy baselines in a fraction of training steps (e.g., reaching PGMORL's final performance in Ant after just 100k steps, and in Humanoid after 200k steps).
-
Produce higher-quality, better-covered Pareto fronts in environments like Walker2d and Hopper, while maintaining competitive sample efficiency—especially useful for real-world continuous-control tasks where interactions are costly.
Abstract
Recent advances in deep reinforcement learning (RL) have shown that improving neural network architectures can yield substantial gains in sample efficiency and asymptotic performance without altering the underlying algorithms. In contrast, work on multi-objective reinforcement learning (MORL), which aims to discover a set of policies that balance trade-offs among conflicting objectives, has predominantly focused on algorithmic innovations, leaving the area of architectures underexplored. While the optimal policies and value functions can differ significantly depending on the trade-offs, MORL algorithms commonly represent them with simple feedforward networks conditioned on the trade-off. This raises the question of whether the performance of the algorithms could be improved with more expressive function approximators. In this paper, we integrate recent advances in neural network design: (i) observation and feature normalization, (ii) weight normalization, and (iii) modeling of distributional returns with an entropy-regularized MORL algorithm. The empirical results across standard continuous control benchmarks demonstrate that these changes substantially improve the quality of the produced solution sets without requiring major changes to the underlying algorithm.
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks