Critic Architecture Matters: Dual vs. Unified Critics for Humanoid Loco-Manipulation

arXiv:2606.11891 · cs.RO, cs.LG · Submitted 2026-06-10 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "Critic Architecture Matters".

Dev: Multi-objective reinforcement learning for humanoid robots must coordinate locomotion and manipulation within a single policy,

Rosa: First, who's behind it and why it matters.

Title and authors: Rosa: So, what the paper boils down is that this decision about the critic design isn't just an arbitrary choice; it’s a design variable that needs to be measured directly instead of assumed.

Dev: They are pointing out that the unified critic tends to let the locomotion reward dominate very early in training, which ends up suppressing how much movement the arm actions actually need to take.

Taro: That suppression idea makes sense; if you heavily weight the walking goal initially, it could result in a robot that walks perfectly but struggles to reach or grasp anything effectively when things deviate from the expected path.

Rosa: Exactly, and they observed that the unified critic produced actions with a mean magnitude of one point two two, which is roughly half of what the dual critics produced, which were around two point five four and three point zero four.

Dev: That difference in action magnitudes makes sense from a control standpoint; if the critic has to satisfy two competing demands at once, it might settle for a safer, less ambitious action than if it had dedicated critics for each task.

Taro: That suggests that the dual-critic approach might be better at finding the true optimal trade-off between movement and manipulation when those two goals aren't perfectly aligned during the initial learning phase.

Rosa: Furthermore, they also highlighted some findings on reward hacking, noting that adding five anti-gaming mechanisms didn't actually provide an extra benefit when used with the dual critics in this specific setup.

Dev: That’s a bit surprising because I thought those extra reward mechanisms might help guard against unintended behavior, but it seems the architectural change itself was the bigger factor for efficiency here.

Taro: So the paper suggests that sometimes simplifying the architecture by using separate critics might be a more effective way to guide reinforcement learning when dealing with multi-objective problems in robotics.

The paper's summary: Rosa: Looking at what the authors suggest moving forward, they are really pushing us to treat this critic design choice as a variable worth measuring instead of just adopting it by default.

Dev: They are advocating for a single-variable ablation study to really establish the causal contribution of the critic architecture, trying to isolate it from other factors like curriculum schedule or action space dimensionality.

Taro: That focus on isolating the variable is crucial for rigorous research because without that control, you can't be sure if a performance gain actually comes from the critic or just a lucky combination of other settings.

Rosa: They hypothesize that dual critics might protect imitation-learned behaviors during RL fine-tuning by reducing interference between objectives, which is a really interesting line of reasoning.

Dev: That idea—that separate critics can act like shields for pre-trained skills—is something we definitely need to test in our own systems when we fine-tune existing models.

Taro: If that hypothesis holds, it implies a way to blend pre-trained knowledge with new reinforcement learning objectives without causing the robot to forget how to walk or move correctly during fine-tuning.

Rosa: They also pointed out a methodological finding that training reward and reach counts actually mask these efficiency differences; the unified critic run accumulated three point three million training reaches while achieving only thirty-six point two reward.

Dev: That’s a huge point for us because it means that just looking at raw training metrics isn't enough to judge if one policy is genuinely better than another, which is a common pitfall in reinforcement learning evaluation.

Taro: So the paper suggests we need more rigorous testing protocols to properly assess these architectural differences than just looking at raw training counts.

The paper's improvements: Rosa: So, wrapping up our discussion on "Critic Architecture Matters: Dual vs. Unified Critics for Humanoid Loco-Manipulation," the main conclusion is that the dual-critic architecture shows clear advantages in terms of training speed and validated performance metrics.

Dev: I think what this paper really hammers home for us as control engineers is that we should be more deliberate about our critic design when we’re dealing with multi-objective problems in robotics, because the structure of the critic matters.

Taro: For autonomy, this means when the world throws a curveball at a humanoid robot, having separate critics might give it better internal decision-making pathways to prioritize stability over reaching in critical moments.

Rosa: Exactly; and I'm really excited about what this means for the future of these robots because if we can reliably separate those objectives, we open up new avenues for creating systems that are both highly mobile and incredibly dexterous.

Dev: I’m ready to see how this translates into practical loop rates and latency constraints in real-time systems, which is the next big question for me regarding deployment.

Taro: I think the implications are that we move closer to robots that can handle complex, dynamic environments much more intelligently than what a single unified learning system could manage alone.

Rosa: Well, team, this paper on "Critic Architecture Matters: Dual vs. Unified Critics for Humanoid Loco-Manipulation" has given us some very clear evidence that how we structure the critic architecture is a design choice that really impacts performance in multi-objective learning.

Dev: It's a solid piece of research showing the practical gains from separating those reward signals, even if the evaluation needs to be done under carefully controlled conditions.

Taro: We should definitely keep an eye on this and see how these insights apply when we start tackling systems with more unpredictable external forces.

Conclusion: Rosa: So we've covered a lot about "Critic Architecture Matters: Dual vs. Unified Critics for Humanoid Loco-Manipulation," and the main takeaway is that having two separate critics for locomotion and manipulation gives us a much more efficient way to train these robots.

Dev: I agree, Rosa; it really shows how crucial it is for us as control engineers to consider the internal architecture of the learning process when we're designing these complex systems, because that directly affects the training loop we have to manage.

Taro: From an autonomy standpoint, I think this confirms that for truly complex tasks, like navigating a cluttered room while picking up an object, having those distinct decision-making pathways is what allows the system to handle unexpected disturbances better than a single unified model.

Rosa: Exactly; I'm really excited about what this means for the future of these robots because if we can reliably separate those objectives, we open up new avenues for creating systems that are both highly mobile and incredibly dexterous.

Dev: I'm ready to see how this translates into practical loop rates and latency constraints in real-time systems, which is the next big question for me when we start thinking about deployment.

Taro: I think the implications are that we move closer to robots that can handle complex, dynamic environments much more intelligently than what a single unified learning system could manage alone.

Rosa: Well, team, this paper on "Critic Architecture Matters: Dual vs. Unified Critics for Humanoid Loco-Manipulation" has given us some very clear evidence that how we structure the critic architecture is a design choice that really impacts performance in multi-objective learning.

Dev: It's a solid piece of research showing the practical gains from separating those reward signals, even if the evaluation needs to be done under carefully controlled conditions.

Taro: We should definitely keep an eye on this and see how these insights apply when we start tackling systems with more unpredictable external forces.

Rosa: What a fantastic discussion; it’s clear that the structural design of the critic isn't just academic, it's fundamental to achieving high-performance humanoid robots.

Dev: I'm looking forward to seeing how these efficiency gains translate into lower latency in our next control system designs.

Taro: It’s compelling evidence that we need to think about these architectural choices proactively when we design autonomy systems for the real world, not just in simulation.

Mehmet Turan Yardımcı

cs.RO, cs.LG

Submitted: 2026-06-10

Updated: 2026-09-25

Comments: Accepted at the ICRA 2026 Workshop on Reinforcement Learning in the Era of Imitation Learning (RL4IL), Vienna. 7 pages, 2 figures. v3 fixes the workshop name and the unified run's description; with its own fingers driven, the standing-mode gap falls from 3.5x/2x to 1.3x/1.1x (speed/throughput; one re-evaluation). No retraining. https://mturan33.github.io/critic-architecture-matters/

Project page: https://mturan33.github.io/critic-architecture-matters

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 77/100

The gist: Multi-objective reinforcement learning for humanoid robots must coordinate locomotion and manipulation within a single policy, presenting a design choice between using a single (unified) critic that

Key concepts

Unified Critic
A single critic used for both locomotion and manipulation tasks. The hosts noted this approach often lets the locomotion reward dominate early training, which can suppress the necessary movement of arm actions.
Dual Critics
Using separate critics, one for each task (locomotion and manipulation). This approach was shown to produce larger action magnitudes than the unified critic, suggesting it finds a better trade-off between competing goals.
Critic Architecture as a Design Variable
The paper argues that the choice of critic design is not arbitrary but a measurable design variable. Researchers should measure this choice directly instead of assuming it will yield optimal results for multi-objective problems in robotics.

Terminology

Summary

Multi-objective reinforcement learning for humanoid robots must coordinate locomotion and manipulation within a single policy, presenting a design choice between using a single (unified) critic that estimates the combined value of all objectives or separate (dual) critics with disjoint reward signals. The paper compares these two architectures on the Unitree G1 humanoid in NVIDIA Isaac Lab, training loco-manipulation policies through sequential curricula progressing from stationary reaching to walking with variable-orientation targets.

The comparison reveals a substantial difference in efficiency: "the dual-critic policy reaches 3.5× faster (6.5 vs. 22.6 simulation steps), achieves 2× higher throughput (14.3 vs. 7.0 validated reaches per 1,000 steps), and attains a higher validated reach rate (65.2% vs. 53.8%) than the unified-critic run in a standardized evaluation."

The authors argue that this efficiency gap is not an isolated effect of the critic, noting that the two runs also differ in curriculum schedule, arm action dimensionality and one locomotion reward weight, and each is a single seed. They report this as an efficiency gap between two trained policies rather than an isolated effect of the critic.

The contributions include: "(1) A Dual Actor-Critic framework with frozen-branch transfer for humanoid locomanipulation on the Unitree G1 in NVIDIA Isaac Lab. (2) A three-way benchmark in which, under a matched compute budget, the dual-critic run reaches 3.5× faster and attains 2× higher throughput than the unified-critic run, while five additional anti-gaming reward mechanisms provide no further benefit. (3) A standardized benchmark methodology demonstrating that training reward and reach counts fail to capture efficiency differences between runs."

Mechanistic insights suggest that "the unified critic must estimate a combined value across competing objectives, so locomotion reward dominates early training and partially suppresses arm action magnitudes, producing a more conservative reaching strategy—consistent with the action-magnitude row of Table II. This is evidenced by the observation that the unified critic produces actions with mean magnitude 1.22, roughly half that of the dual critics (2.54 and 3.04)."

Furthermore, Anti-gaming reward mechanisms provide no additional benefit, as the dual-critic policy without them (S6s, 65.2%) slightly outperforms the variant with five anti-gaming mechanisms (S7, 60.9%). The authors suggest this is because S7 retrains its arm policy from scratch on top of a frozen locomotion branch, whereas S6s continues a jointly trained one.

The paper concludes that Critic architecture is a design variable worth measuring rather than adopting by default, and specifies the need for a single-variable ablation required to establish its causal contribution through identical curriculum, action space, and reward set comparisons across multiple seeds. The central hypothesis is that dual critics may protect imitation-learned behaviors during RL fine-tuning by reducing objective interference.

The authors also report a methodological finding: training reward and reach counts mask the efficiency difference, as the unified-critic run accumulated 3.3M training reaches and achieved reward 36.2, against 37.1 for the dual-critic run—figures a practitioner would read as comparable, while the 3.5× speed difference and 2× throughput gap are invisible in them and emerge only through standardized evaluation with time-to-reach and throughput measurements.

Limitations noted include "simulation-only evaluation, single-seed training with no seed set in either run, 5-DoF arm control, and—most importantly—the confounding of curriculum schedule, arm action dimensionality and one locomotion reward weight with the critic architecture (Sec. IV-D). The immediate next step is to conduct a single-variable ablation: identical curriculum, action space and reward set, only the critic swapped, across at least three seeds. The paper also plans to extend this work to 29-DoF dual-arm control with a Triple Actor-Critic" and integrate vision-language models.

Index Terms—reinforcement learning, imitation learning, humanoid robots, critic architecture, loco-manipulation, fine-tuning.

Improvements for AI systems

Here are specific improvements to AI systems based on the findings in this research, focusing on multi-objective humanoid robots:

  1. The core improvement is adopting a Dual Critic Architecture for any policy requiring simultaneous locomotion (walking/balance) and manipulation (reaching/grasping). Instead of a single critic estimating a combined value, use two separate critics with disjoint reward signals: one specifically for locomotion objectives (e.g., velocity tracking, balance) and one specifically for manipulation objectives (e.g., reach distance, displacement).

  2. The improved system will achieve significantly faster learning and higher throughput during the initial stages of training compared to a unified critic approach under the same compute budget (demonstrated 3.5x speedup and 2x throughput increase).

  3. The system's learned control actions for manipulation tasks will exhibit more consistent and committed movements, as evidenced by significantly larger arm action magnitudes (2.54x to 3.04x greater magnitude in the dual-critic run versus 1.22x for the unified run). This suggests the dual critic prevents the locomotion gradient from suppressing or diluting necessary manipulation efforts.

  4. For policies fine-tuned via Imitation Learning (IL), this architecture offers a crucial mechanism to prevent catastrophic forgetting. By isolating the RL signal to only one objective's reward stream, the policy is less likely to overwrite pre-trained, imitation-learned skills in other domains (e.g., locomotion).

  5. The system can be deployed in complex human environments where both coordinated movement and precise object interaction are required (e.g., navigating a cluttered room while picking up an object), achieving a validated reach rate of 65.2% compared to 53.8% for the unified critic run, indicating superior task success in real-world scenarios.

  6. The system can be designed with modularity by freezing the locomotion branch's policy and training only the arm policy using its dedicated dual critic, ensuring that fine-tuning for manipulation does not destabilize or degrade existing stable walking skills.

Sources

Related papers