A Change of Frame Makes the Capture Point Proprioceptive: Distillation-Free Humanoid Single-Leg Balance
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "A Change of Frame Makes the Capture Point Proprioceptive".
Dev: Unified humanoid policies struggle to maintain clean single-leg balance, often resorting to recovery actions like stepping or hopping rather than prevention.
Rosa: First, who's behind it and why it matters.
Title and authors: Rosa: So, we're talking about this paper now, "A Change of Frame Makes the Capture Point Proprioceptive: Distillation-Free Humanoid Single-Leg Balance." Rosa here, I'm curious if we can actually take these kinds of policies out of the simulator and see them holding steady on real hardware for any meaningful amount of time.
Dev: From a control engineering standpoint, that’s exactly what I want to know; the latency and loop rate are critical when you move from simulation to reality. The authors mention they transfer directly to a Unitree G1 without distillation, which is promising, but we need proof it doesn't break under real-world sensor noise or unexpected dynamics.
Taro: As an autonomy researcher, my main concern is what happens when the world throws something completely unexpected at the system; does this policy just fall over because it doesn't anticipate novel disturbances?
Rosa: That’s a fair point, Taro; we’re looking for that root-level competence. The authors are essentially trying to move past policies that only recover from errors, focusing instead on the prevention aspect of balance using the capture point.
Dev: Exactly; they argue that current methods focus too much on just keeping the center of mass inside a polygon when motion starts, but this paper addresses the dynamic signals needed when you're actually moving.
Taro: And I wonder if this "change of frame" observation they introduce is robust enough to handle things like sudden pushes or uneven terrain without needing complex model-based reactive planning layered on top.
Rosa: That’s a good question about the deployability gap; they claim this specific observation, called support-relative dynamic-CoM, lets them get around not needing that unmeasurable base linear velocity signal.
Dev: If that observation is truly reconstructible from just encoders and IMU data on board, then the latency should be manageable for a real deployment at fifty Hertz.
The paper's summary: Rosa: So, to summarize what this paper proposes with "A Change of Frame Makes the Capture Point Proprioceptive: Distillation-Free Humanoid Single-Leg Balance," they are tackling the fundamental problem that unified humanoid policies struggle with: maintaining clean single-leg balance during agile motion. They show that standard policies often resort to recovery actions like hopping or stepping when they can't maintain a stable stance.
Dev: The core idea they present is using a deployable actor and a privileged critic shaped by what they call a human-science reward library, which translates postural control science directly into the AI's objective function, focusing on prevention instead of just repair.
Taro: I see how that connects to their work on posture; by encoding terms like stability margins and time-to-boundary thresholds right into the reward structure, they are trying to teach the policy *how* humans prevent falls rather than just mimicking successful balances.
Rosa: Precisely, and they demonstrate this approach works on real hardware with near perfect success rates across nine stratified pose classes and transfers directly to a Unitree G1 without needing any distillation from a teacher model.
Dev: That transferability is a big deal for deployment; if it works that well in simulation, it suggests the underlying control logic is robust enough to handle the realities of physical hardware constraints.
Taro: It’s interesting how they bridge the gap between theoretical postural control research and actual learned policy implementation through this reward library approach.
The paper's improvements: Rosa: Looking at the improvements detailed in "A Change of Frame Makes the Capture Point Proprioceptive: Distillation-Free Humanoid Single-Leg Balance," one major improvement is solving the deployability gap for crucial balance signals like the capture point, which usually requires unmeasurable base linear velocity. They solve this by using a change of frame observation that cancels out that unmeasurable velocity entirely.
Dev: That's the technical fix I was hoping to see; if they can derive this support-relative dynamic-CoM state purely from joint encoders and the gyroscope, it massively simplifies the required onboard sensing and reduces latency concerns for real hardware deployment.
Taro: From an autonomy perspective, this makes the policy much more self-contained because it doesn't rely on external or difficult-to-measure base velocity inputs to function correctly during dynamic maneuvers.
Rosa: Beyond that, they integrate a human-science reward library where they translate concepts like spatial stability margins and time-to-boundary thresholds into graded action penalties for the stance leg, which steers the robot toward smooth torque generation rather than just keeping it upright.
Dev: I also noticed they include a jerk penalty based on the second difference of action to suppress high-frequency motor chatter; that’s important because it directly relates to mechanical wear and noise in real actuators.
Taro: So, by combining the proprioceptive observation fix with these explicit, physics-informed reward terms, they are aiming for a behavior that feels like genuine prevention rather than just tracking a static goal.
Conclusion: Rosa: To wrap up the discussion on "A Change of Frame Makes the Capture Point Proprioceptive: Distillation-Free Humanoid Single-Leg Balance," the main implication is that we can move toward humanoid robots that don't just stumble and recover, but actually learn the fundamental principles of postural control to prevent falls.
Dev: The real impact here is demonstrating that these complex, high-level balance skills can be learned directly on real hardware from a policy trained in simulation, bypassing the need for extensive teacher-student distillation methods.
Taro: I think this work suggests that the path forward for generalist policies isn't just about absorbing more motion breadth, but about embedding root-cause balance design principles into the learning objective itself.
Rosa: And that’s what makes me really optimistic; we’re seeing a policy achieve ninety-eight point nine percent perfect success on held-out test sets, and that kind of performance suggests a significant step toward reliable deployment outside of highly controlled lab environments <ref:2608.00500#pg0>.
Dev: If this holds up under the continuous metrics they reported—specifically mentioning near-zero fore–aft margin and a capture point out-of-support duration of only zero point zero nine seconds—then we're looking at something genuinely biomechanically sound for real applications.
Taro: It certainly points toward systems that exhibit root-level competence, which is what we need if these robots are ever to interact with humans in dynamic settings.
Rosa: So, the DDC policy described in "A Change of Frame Makes the Capture Point Proprioceptive: Distillation-Free Humanoid Single-Leg Balance" shows that by providing a deployable observation and a human-science reward structure, we can move away from reactive recovery and toward proactive prevention.
Dev: It’s a significant step for control engineering because it shows how to build robust policies using only on-board sensory data for critical tasks.
Taro: This paper sets a high bar for what generalist policies need to achieve if they are ever going to handle complex, real-world physical challenges autonomously.
Peking University
cs.RO
Submitted: 2026-08-01
Updated: 2026-10-01
Code: https://github.com/amazonfar/holosoma
Project page: https://estoil.github.io/DDC
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 85/100
The gist: Unified humanoid policies struggle to maintain clean single-leg balance, often resorting to recovery actions like stepping or hopping rather than prevention.
Key concepts
- deployable dynamic Center of Mass observation
- This is a specific way to measure the robot's center of mass relative to its support foot. It solves the problem of needing unmeasurable base velocity by expressing the capture point in a frame relative to the support foot, allowing it to be reconstructed using only joint encoders and an IMU.
- human-science reward library
- This is a set of rewards derived directly from human postural control science. It encourages prevention by penalizing states that are near falling (stability margin) or close to crossing the support boundary in time (time-to-boundary). It also includes penalties for smooth movements and specific response hierarchies for the stance leg.
- change of frame observation
- This technique mathematically cancels out the unmeasurable base linear velocity when calculating the capture point relative to the support foot. By shifting the coordinate system, it transforms a complex balance signal into an observable state that can be derived entirely from on-board sensors like encoders and gyroscopes.
- method-agnostic benchmark
- This is a standardized testing environment used to evaluate different policies fairly. It scores every released policy in a simulator separate from its training environment. The results are categorized into three tiers: perfect hold, marginal success (recovery), or failure (fall), providing an objective measure of competence.
Terminology
Summary
Unified humanoid policies struggle to maintain clean single-leg balance, often resorting to recovery actions like stepping or hopping rather than prevention. This work introduces DDC, a distillation-free policy that achieves near-perfect single-leg stance on real hardware by incorporating a deployable dynamic Center of Mass observation and a human-science reward library.
The Gist
DDC holds clean single-leg balance on 89 of 90 held-out motions across nine stratified pose classes and transfers to a real Unitree G1, demonstrating that the capture point can be made observable from encoders and IMU alone, driving a learned policy directly on hardware.
Closing the Deployability Gap with Dynamic CoM Observation
The paper addresses the deployability gap where crucial balance signals, like the capture point (xCoM), require unmeasurable base linear velocity. The authors solve this by introducing a change of frame
observation: expressing the capture point relative to the support foot causes the unmeasurable base velocity to cancel out exactly. This results in an observation, denoted as support-relative dynamic-CoM, which is reconstructible on-board from encoders and IMU alone. Specifically, it provides the horizontal components of this relative state: obal = (rB, r˙B) ∈ R 4,
which is built only from joint encoders and the gyroscope. This observation is described as the deployable balance state it needs.
A Human-Science Reward Library for Prevention
The policy is paired with a reward library translated term by term from human postural control science, focusing on prevention over repair.
The core of this library includes:
-
Two soft penalties on the capture point: a
stability-margin term
penalizing the xCoM when it nears the support boundary (spatial buffer), and atime-to-boundary (TTB) term
penalizing when the projected time to cross that boundary falls below a reaction threshold (temporal buffer). -
An
ankle→knee→hip response hierarchy,
which encodes graded action-rate penalties on the stance leg, heaviest on the ankle, to steer it toward smooth, sustained torque. -
A
smoothness
penalty based on action jerk (the second difference of the action) to suppress high-frequency motor chatter.
Method-Agnostic Benchmark and Evaluation
To provide a standardized measure of competence, the authors release a method-agnostic, reproducible sim2sim benchmark.
This testbed scores every released policy in a simulator distinct from its training environment. The evaluation uses three outcome tiers:
-
A Perfect hold: A clean single-leg stance across the single-support window without hopping or falling.
-
A Marginal success: The robot does not fall but stays upright only by breaking the clean constraint, such as hopping the support foot or touching the swing foot down (a recovery from a capture-point/support-polygon mismatch).
-
A Failure: The robot loses balance and falls.
Performance and Ablation Results
The DDC policy achieved 98.9% Perfect success on the held-out test set, outperforming eight state-of-the-art general policies, which all scored 0/90 Perfect holds. Ablations confirmed that the deployable dynamic-CoM observation is the single largest driver,
costing 43 points of clean single-leg balance when removed. The continuous metrics also showed DDC had a near-zero fore–aft margin of stability and a capture point out-of-support duration of only 0.09 seconds, indicating biomechanically genuine balance. The policy transfers directly to the real Unitree G1 without distillation or teacher–student methods.
Generalist Policy Limitations
The benchmark demonstrated that even strong generalist policies fail single-leg balance, achieving a Perfect rate of exactly 0/90. While some policies exhibit recovery-dominant behavior (rarely falling but surviving by shuffling), the strongest generalists lack the root-level competence to hold the pose cleanly,
confirming that scale confers robustness, not root-cause single-leg balance. The results show that DDC’s success is rooted in a root-cause balance design
rather than simply absorbing motion breadth.
Deployment Details
The deployed policy runs directly on the Unitree G1 at 50 Hz using ONNX. Every observation consumed by the actor is reconstructed on-board from joint encoders and IMU; the base linear velocity is canceled out, leaving only the support-relative dynamic-CoM state. The support foot selection rule used during training (lower foot supports; feet within 3 cm count as double support) ensures a train-to-deploy continuity,
meaning the policy functions identically on hardware as it did in simulation. The actor observes a total of 463 dimensions, including the key term, "support-relative dynamic-CoM obal.
Improvements for AI systems
Here are the specific improvements for AI systems based on this research, focusing on moving from general whole-body tracking to robust single-leg balance capability:
The primary improvement is shifting policy training from purely motion imitation (tracking) to incorporating a physically grounded, root-cause balance mechanism. This transforms humanoid policies from stumbling
to preventing.
Here are the specific improvements and capabilities of the resulting AI system:
-
A policy that can maintain a clean single-leg stance across a wide range of challenging poses (squat depth and swing-foot height variations) on real hardware, achieving near 99% success rate (89/90).
-
The ability to proactively prevent falls by observing the dynamic Center of Mass (CoM) relative to the support foot, even when absolute base linear velocity is unmeasurable.
-
Robustness against deployment noise: The system maintains high-level balance performance even when subjected to realistic on-board sensor noise (IMU drift and velocity noise), demonstrating a significant resilience compared to generalist policies.
-
A method for objectively measuring and comparing the single-leg balance competence of diverse, pre-existing general whole-body policies using a standardized, reproducible benchmark (the sim2sim testbed).
-
An ability to translate fundamental human postural control principles—such as spatial capture-point margins and time-to-boundary thresholds—directly into the reward function of the learned policy, ensuring the resulting behavior is biomechanically sound and prevents
recovery by repair
instead ofprevention.
In summary, the improved AI system can perform high-stakes, single-leg support tasks reliably in real environments without requiring specialized distillation from a teacher model.
Sources
- LocoMuJoCo: A Comprehensive Imitation Learning Benchmark for Locomotion
- HoloMotion-1 Technical Report
- GMT: General Motion Tracking for Humanoid Whole-Body Control
- Expressive Whole-Body Control for Humanoid Robots
- HumanPlus: Humanoid Shadowing and Imitation from Humans
- OmniH2O: Universal and Dexterous Human-to-Humanoid Whole-Body Teleoperation and Learning
- Switch-JustDance: Benchmarking Whole Body Motion Tracking Controllers Using a Commercial Console Game
- A Kung Fu Athlete Bot That Can Do It All Day: Highly Dynamic, Balance-Challenging Motion Dataset and Autonomous Fall-Resilient Tracking
- BeyondMimic: From Motion Tracking to Versatile Humanoid Control via Guided Diffusion
- SMPLOlympics: Sports Environments for Physically Simulated Humanoids
- SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control
- Agility Meets Stability: Versatile Humanoid Control with Heterogeneous Data
- Embedding Classical Balance Control Principles in Reinforcement Learning for Humanoid Recovery
- Humanoid-GPT: Scaling Data and Structure for Zero-Shot Motion Tracking
- Learning Sim-to-Real Humanoid Locomotion in 15 Minutes
- HumanoidBench: Simulated Humanoid Benchmark for Whole-Body Locomotion and Manipulation
- MOSAIC: Bridging the Sim-to-Real Gap in Generalist Humanoid Motion Tracking and Teleoperation with Rapid Residual Adaptation
- HumanoidArena: Benchmarking Egocentric Hierarchical Whole-body Learning
- OmniXtreme: Breaking the Generality Barrier in High-Dynamic Humanoid Control
- General Humanoid Whole-Body Control via Pretraining and Fast Adaptation
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving