BAT: Balancing Agility and Stability via Online Policy Switching for Long-Horizon Whole-Body Humanoid Control

arXiv:2604.01064 · cs.RO · Submitted 2026-04-01 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "BAT: Balancing Agility and Stability via Online Policy Switching for Long-Horizon Whole-Body Humanoid Control".

Dev: Developing a unified framework that can achieve agile, precise, and robust whole-body behaviors—particularly in long-horizon tasks—remains challenging due to conflicting control requirements such as stable manipulation versus highly dynamic responses.

Rosa: First, who's behind it and why it matters.

Paper summary: Rosa: To recap, this paper introduces BAT as an online policy-switching framework designed to address the difficulty in achieving agile, precise, and robust whole-body behaviors for long tasks by dynamically choosing between two complementary RL controllers.

Dev: Essentially, the thesis is that existing methods struggle with the trade-off between coupled policies for coordination and decoupled policies for precision without a systematic way to integrate them.

Rosa: BAT claims to solve this by proposing two modules: an option-guided HRL framework and an option-aware VQ-VAE that predicts option preference from motion tokens.

Dev: The paper highlights their contributions as an online policy switching framework, the option-aware VQ-VAE for rich latent space representations, and extensive validation on simulation and hardware.

Taro: I see how the authors are focusing on improving generalization by using this VQ-VAE structure to encode motion phase dependent features into discrete tokens that inform those switching decisions.

Rosa: That’s right; they're aiming for richer downstream inference by having these tokens capture the specific characteristics of different motion phases relevant for switching.

Dev: And they use offline supervision from sliding-horizon policy pre-evaluation to guide the HRL, which helps with training stability and sample efficiency in those long-horizon scenarios.

Taro: This structured guidance is important because it directly tackles the sample inefficiency that purely reward-based switching often suffers from when dealing with rare decision events.

Rosa: Exactly; by using this pre-evaluation, they are providing the system with prior knowledge to make smarter decisions before it even hits the main policy training loop.

Dev: So, if we boil it down, BAT is about orchestrating these two complementary whole-body RL controllers online based on a learned context that captures motion characteristics.

Conclusion: Rosa: Looking at "BAT: Balancing Agility and Stability via Online Policy Switching for Long-Horizon Whole-Body Humanoid Control," the authors Donghoon Baek, Sang-Hun Kim, and Sehoon Ha are presenting a method to dynamically manage the control strategy for humanoid robots during long sequences.

Dev: It boils down to having a system that can switch its underlying control paradigm on the fly—between stable manipulation mode and agile movement mode—based on what it’s currently doing.

Rosa: The implication is that we could see whole-body systems perform tasks that require both fine dexterity and dynamic action, which is currently hard because we're stuck choosing one or the other.

Dev: If this works reliably outside the lab, it means autonomous robots could handle much more complex environments where they have to transition between different physical demands smoothly without getting unstable or losing precision.

Taro: From an autonomy research standpoint, this suggests a path toward building systems that are fundamentally more adaptable to unpredictable real-world conditions because they aren't locked into a single control setting.

Rosa: Precisely; the framework offers a way for these robots to adapt their fundamental behavior based on the motion context, which is something we desperately need for true general-purpose autonomy.

Dev: It shows that combining structured guidance with hierarchical RL can lead to more effective switching strategies than just relying on purely reward-based signals alone for policy selection.

Taro: So, it points toward a future where whole-body control is not just about optimizing one objective, but about intelligently managing a set of competing objectives simultaneously in real time.

cs.RO

Submitted: 2026-04-01

Updated: 2026-10-05

Code: https://github.com/LeCAR-Lab/HumanoidVerse

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 83/100

The gist: Developing a unified framework that can achieve agile, precise, and robust whole-body behaviors—particularly in long-horizon tasks—remains challenging due to conflicting control requirements such

Key concepts

Decoupled Policy (πD)
This controller focuses on stability and robustness. It is designed to handle disturbances well, making it suitable for precise manipulation tasks where maintaining a steady posture and minimal movement is critical. It prioritizes safety and predictable behavior over high-speed agility.
Coupled Policy (πC)
This controller enables highly dynamic motions, allowing the robot to perform agile actions like jumping or rapid maneuvers. While excellent for speed and dynamism, it can sometimes lack the stability needed for precise tasks due to its tight coupling between body segments.
Option-Guided HRL
Hierarchical Reinforcement Learning (HRL) is used here with 'options'—sub-goals—to manage the complex decision of switching policies. The high-level policy uses these options to evaluate which controller (πD or πC) is best for the current situation, helping to solve credit assignment problems in long tasks.
Option-Aware VQ-VAE
This component encodes robot motions into discrete tokens that capture the essential structure of different movements. By training this encoder with an option prediction head, it learns a latent space where motions favoring one controller naturally separate from those favoring the other, informing the switching decision.

Terminology

Summary

Developing a unified framework that can achieve agile, precise, and robust whole-body behaviors—particularly in long-horizon tasks—remains challenging due to conflicting control requirements such as stable manipulation versus highly dynamic responses. This work proposes BAT, an online policy-switching framework that dynamically selects between two complementary whole-body Reinforcement Learning controllers to balance agility and stability across different motion contexts.

The gist

BAT is an online policy-switching framework that orchestrates complementary control policies to leverage their respective strengths for long-horizon, multi-task scenarios by dynamically selecting the most suitable policy based on the current motion context.

Framework Components and Motivation

The paper addresses the trade-off between decoupled whole-body policies, which provide stable and disturbance-robust behaviors, and coupled whole-body policies, which enable agile and dynamic motions but often sacrifice precision. The framework is motivated by the need for a unified approach that leverages both paradigms for long-horizon tasks requiring both highly dynamic motions (e.g., jumping over obstacles) and precise, stable manipulation (e.g., standing manipulation with minimal disturbance). To solve the challenge of online policy switching, the authors leverage hierarchical reinforcement learning (HRL) guided by sliding-horizon option evaluation to mitigate the absence of ground-truth signals and address temporal credit assignment issues where switching benefits are delayed.

Key Mechanisms for Policy Selection

BAT integrates three main modules to facilitate dynamic decision-making:

  1. An Option-Guided Hierarchical Reinforcement Learning (HRL) framework that selects a discrete controller index, choosing between the decoupled policy (πD) and the coupled policy (πC), based on an observation constructed from proprioception, motion reference, and previous state. The high-level switching policy is optimized using a composite reward function that includes velocity tracking rewards, pose tracking rewards, regularization terms for smoothness and energy efficiency, and a soft alignment term to guide switching behavior.

  2. An Option-Aware Vector-Quantized Variational Autoencoder (VQ-VAE) that encodes motions into discrete tokens. This VQ-VAE is trained with an objective that aligns the token space with switching decisions by introducing an option prediction head, encouraging latent tokens to separate motions favoring different controllers, thereby capturing controller-dependent structure relevant for switching.

  3. A Decision Fusion Module that selects the final controller by trusting the high-level policy when in-distribution and deferring to the option prediction head when outside its training support. This module estimates distributional shift using three signals: H(πsw), DKL(Pcb∥Ptrain), and H(pϕ), fusing them to determine the final selection, ensuring reliability even when local value differences are high.

Data Construction and Guidance

The framework relies on extensive offline data construction for structured guidance. This involves two primary steps:

  1. Sliding-Horizon Policy Pre-Evaluation: This step assesses switching decisions locally in time by evaluating each option (πD or πC) over a finite window of length H, isolating local motion segments and preventing short-lived instability from dominating long-horizon performance. The expert decision is defined as the one maximizing the option return within that window.

  2. Offline Data Collection: This step applies sliding-horizon evaluation over retargeted motion data from both policies (πD and πC), generating high-quality switching demonstrations data (DOp) via motion blending with inertialization to ensure smooth transitions at clip boundaries.

Experimental Validation

Extensive simulations and real-world experiments validate BAT across diverse scenarios. In simulation, the option-aware VQ-VAE representation was shown to achieve higher reward and success rates compared to raw motion representations, indicating that improved representation quality enhances policy identification. When tested on sequential multi-motion tasks, BAT achieved the best overall performance by learning improved switching strategies through exploration, demonstrating the effective integration of HRL and option guidance. Hardware deployment on the Unitree G1 humanoid confirmed that BAT enables successful execution of sequential tasks by dynamically switching between policies to handle diverse motion phases within a unified framework, consistently selecting the more suitable policy based on motion characteristics.

Conclusion

BAT successfully balances agility and stability for long-horizon whole-body control by orchestrating complementary controllers through an option-guided HRL framework and an option-aware VQ-VAE. The approach demonstrates that combining structured guidance with HRL provides a more effective solution than purely reward-based switching, leading to improved performance across single-motion and sequential multi-motion tasks in both simulation and hardware. BAT proves versatile by dynamically selecting the appropriate policy based on the motion context, achieving superior success rates over competing baselines. The framework is modular and validated for real-world deployment on humanoid platforms.


How it works

The framework is structured around an Option-Guided Hierarchical Reinforcement Learning (HRL) setup where a high-level switching policy selects between two frozen low-level controllers, πD (decoupled) and πC (coupled).

Improvements for AI systems

Here are the specific improvements and capabilities enabled by implementing the BAT framework, based on the provided paper:


) Robust Long-Horizon Whole-Body Control with Adaptive Agility/Stability Trade-off:

The improved system (BAT) can execute complex, multi-stage long-horizon tasks—such as navigating rough terrain while simultaneously performing precise manipulation—by dynamically switching between two complementary whole-body control paradigms.

  1. The system can switch to the decoupled policy for stable, disturbance-robust behaviors (e.g., standing on uneven ground or fine object placement).

  2. The system can switch to the coupled policy for highly dynamic, agile motions (e.g., jumping over obstacles or rapid locomotion).

) Enhanced Sample Efficiency via Option-Guided Hierarchical Reinforcement Learning (HRL):

By using offline option guidance derived from sliding-horizon policy pre-evaluation, the system drastically reduces the required number of interaction samples needed to learn effective switching policies.

  1. The system learns optimal switching patterns directly from expert demonstrations (DOp) without relying solely on sparse, delayed reward signals.

  2. This leads to significantly faster training convergence and better credit assignment for long-horizon decisions compared to standard HRL or value-function based switching methods, reducing the overall computational cost of policy learning by up to a factor related to the sample complexity reduction (Equation 6).

) Superior Motion Representation via Option-Aware VQ-VAE:

The system utilizes an option-aware VQ-VAE that explicitly learns latent representations tailored to distinguish between motion phases favored by different controllers.

  1. The learned discrete tokens capture controller-dependent structural features in the motion space, which the decision fusion module uses to make more informed switching choices.

  2. This results in higher accuracy in identifying task contexts, as evidenced by superior performance metrics (Table I and Fig 4), allowing the system to select the correct policy based on subtle differences in motion structure rather than just raw proprioception.

) Reliable Online Switching via Confidence-Weighted Fusion:

The final decision-making module employs a confidence-weighted fusion strategy, integrating three distinct signals: the high-level switching policy's output, the predicted option preference from the VQ-VAE, and a distributional shift measure (DKL divergence).

  1. When uncertainty is low (in-distribution), it trusts the learned switching policy.

  2. When uncertainty is high or the motion context is ambiguous, it falls back to the option prediction head, providing a reliable oracle estimate without requiring perfect transition exposure during training.

) Improved Task Performance and Generalization:

The integrated framework enables superior performance across diverse tasks compared to using either the decoupled or coupled policies in isolation.

  1. In simulation, BAT achieves higher success rates and lower tracking errors than both the pure decoupled and pure coupled baselines (Table III).

  2. On real hardware (Unitree G1), the system demonstrates versatility, successfully executing sequential long-horizon tasks that require a continuous interplay of stable manipulation, locomotion, and dynamic maneuvers.

Sources

Related papers