PAC-MAN: Perception-Aware CBF-RL for Whole-Body Safety in Humanoid Dodgeball
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "PAC-MAN: Perception-Aware CBF-RL for Whole-Body Safety in Humanoid Dodgeball".
Dev: We present PAC-MAN, a perception-aware CBF-RL framework that couples control-barrier safety with deployment-realistic onboard sensing for whole-body humanoid dodgeball.
Rosa: First, who's behind it and why it matters.
Title and authors: Rosa: So we’re talking about PAC-MAN today, which is titled "Perception-Aware CBF-RL for Whole-Body Safety in Humanoid Dodgeball." It’s clear from the title that they are focusing on how to make a humanoid robot safe during dodgeball using perception and control barrier functions.
Dev: That title tells you immediately that the core contribution is bridging the gap between safety guarantees provided by control barrier functions and the actual sensing limitations of a deployed system. It’s about making sure it stays safe even when things aren't perfect.
Taro: I think focusing on whole-body safety is important because most robotics papers focus on just keeping the base stable, but PAC-MAN seems to push for safety across every single link of the robot during an evasive maneuver.
Rosa: Precisely, and what they do in simple terms is they train the policy to understand how to move the whole body based on partial, realistic camera input while using training guidance that covers all joints. They want this policy to be robust enough for real-world use where you don't have perfect data.
Dev: The authors are Lizhi Yang, Junheng Li, and Aaron D. Ames, and they’ve clearly done some deep work on the mechanics of how these perception constraints interact with learning safety policies. I think their background in both robotics and control engineering is what makes this paper feel so grounded in reality.
Taro: I'm interested in how they framed the problem as a partially observed Markov decision process; it sounds like they are treating this as a real-time decision-making problem under uncertainty rather than just a static planning task.
Rosa: That’s right, and that formulation is what lets them talk about joint-position targets being emitted at control step t, which is essential for any system running on physical hardware with time constraints.
Dev: And the observation constraint they put on the policy observation o t is very specific: it must only contain signals computable on hardware from onboard sensing, which sets a hard limit on what kind of information the AI can rely on during operation.
Taro: That constraint really highlights the deployment reality; if you can't compute something on-board, the policy simply can't use it to make decisions at runtime.
Rosa: Exactly, and that’s why they have to carefully design what information is fed into the system versus what is only used during training guidance.
Dev: This leads us into the core of their methodology where they introduce Link-CBF and Joint-CBF as different levels of safety enforcement during training versus runtime.
The paper's summary: Rosa: To summarize PAC-MAN, this perception-aware CBF-RL framework couples control barrier safety with deployment-realistic sensing for whole-body humanoid dodgeball. It basically says the robot learns to avoid being hit by a ball by using depth images from a head camera as its main input.
Dev: The key summary point I see is that the training guidance uses clearance information for every body link, but in deployment, the policy only sees segmentation-masked depth, and they use an adversarial motion prior to shape those evasive reflexes into more natural movements like leaning or sidestepping.
Taro: So it’s not just about learning a dodgeball avoidance strategy; it’s about making sure that whatever the policy learns is physically executable and safe across the entire body structure, which is a significant step up from simpler approaches.
Rosa: That’s right, and they evaluate this on two distinct regimes: single throws for quick reactions and a deployment loop where the robot has to walk back to its station and recover between throws.
Dev: The summary also highlights that they found that usable barrier structure depends on perceptual observability; specifically, Joint-CBF performs best when accurate ball states are available, whereas Link-CBF is the most deployable option under limited onboard perception.
Taro: That means the system's performance is directly tied to how well you can track or infer the threat geometry from just that depth data during operation.
Rosa: It seems like they’ve really established a clear hierarchy of safety mechanisms: training uses the whole-body view, but deployment relies on a lighter per-link barrier structure for practical success.
Dev: And they explicitly state that they validate this design choice by deploying the fixed-camera Link-CBF policy zero-shot on a Unitree G1 using only onboard depth and proprioception to prove its deployability.
The paper's improvements: Rosa: One major improvement they highlight is replacing the reliance on just the pelvis-centered safety with a whole-body link safety across all body segments, which means maintaining safety throughout the torso and limbs.
Dev: That’s a big step because limiting it to just the base can lead to instability when you have dynamic interactions or external forces affecting other parts of the body. It shows they are thinking about holistic stability.
Taro: And I think incorporating perception-aware CBF is important because it allows the policy to infer and enforce collision avoidance based on partial or noisy visual input, meaning it can be proactive even if a ball is partially blocked.
Rosa: Right, and then they have this perception-conditioned safety structure that dynamically adjusts its strength based on the quality of perception you are getting during operation. If the vision is bad, the system adjusts accordingly.
Dev: That dynamic adjustment sounds like a very smart way to handle uncertainty; it means you don't rely on a fixed barrier when the input quality changes and perhaps that's a failure mode they've mitigated.
Taro: I also think the adversarial motion prior is valuable because it ensures that the evasive reflexes aren't just arbitrary joint movements but are instead dynamic behaviors like crouching or leaning, which makes the movement more realistic.
Rosa: That regularization helps ensure that when the policy does try to evade, it does so in a way that is actually feasible for a humanoid robot to execute efficiently.
Dev: So, the paper suggests that for deployment, Link-CBF is the best deployable option under limited onboard perception because it balances safety with hardware limitations effectively.
Conclusion: Rosa: To wrap up on PAC-MAN today, the main implication is that perception-aware control barrier methods allow robots to operate safely in dynamic, unpredictable environments like dodgeball without needing perfect information.
Dev: I think the impact is showing that as long as you can design a safety structure that adapts to observation quality, you can achieve high success rates even when operating under real-world perception constraints.
Taro: The takeaway for autonomy research is that the bottleneck in achieving robust evasion isn't necessarily the control theory itself, but rather how much information a policy can internalize when it only receives deployment-realistic observations.
Rosa: Right, and they’ve shown that this approach works well on benchmarks where the policy comes within a few points of an oracle providing perfect state knowledge, suggesting perception is indeed the main hurdle to solving these kinds of problems.
Dev: And for me, it’s about the loop rate; if the deployment requires constant re-evaluation based on imperfect depth data, we need to ensure that latency doesn't cause a failure mode during those critical evasion steps.
Taro: I just think as long as you can quantify that performance against an oracle, like achieving ninety-five percent success rates in real-world scenarios, it validates the method for practical application.
Rosa: Exactly; the PAC-MAN paper gives us a solid foundation on how to design these systems to be robust against perception limitations and ready for deployment on hardware like the Unitree G1.
cs.RO, cs.AI
Submitted: 2026-07-30
Updated: 2026-09-29
Comments: Website at https://lzyang2000.github.io/perceptive_cbf_rl/
Code: https://github.com/ccrpRepo/AMP_mjlab
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 87/100
The gist: We present PAC-MAN, a perception-aware CBF-RL framework that couples control-barrier safety with deployment-realistic onboard sensing for whole-body humanoid dodgeball.
Key concepts
- PAC-MAN
- A perception-aware Control Barrier Function Reinforcement Learning framework designed for whole-body safety in humanoid dodgeball. It couples control barrier safety with deployment-realistic sensing to ensure the robot stays safe during evasive maneuvers.
- Control Barrier Functions (CBF)
- A method used to provide safety guarantees in control systems. In this paper, they use different levels of CBF enforcement: Link-CBF for deployment and Joint-CBF for training guidance, balancing safety with hardware constraints.
- Perception-Awareness
- The ability of the policy to adjust its safety structure based on the quality of onboard perception. The system dynamically adjusts the strength of the safety barrier depending on how accurate or noisy the visual input is during operation.
Terminology
Summary
We present PAC-MAN, a perception-aware CBF-RL framework that couples control-barrier safety with deployment-realistic onboard sensing for whole-body humanoid dodgeball. The deployed policy sees the ball only as segmentation-masked depth from a head-mounted camera, while training-time CBF guidance represents clearance to every body link, and an adversarial motion prior regularizes the resulting evasive reflexes. We evaluate on a controlled anylink contact benchmark with seeded throws in two regimes: single throws and a deployment loop in which the robot walks back to its station and recovers between throws. On this benchmark, the policy comes within a few points of a privileged state oracle: a fixed onboard camera alone is adequate for evasion.
We find that "usable barrier structure depends on perceptual observability: Joint-CBF gives the best performance with accurate ball states, degrades under fixed-camera observations when used only as training guidance, and recovers with a ball-tracking gimbal or privileged runtime filter. We therefore deploy a lightweight Link-CBF policy zero-shot on the Unitree G1 in the real world,
where it tolerates imperfect perception, succeeds on 95% of throws, and uses semantic segmentation to dodge different balls."
PAC-MAN trains a perception-aware policy to map partial onboard observations to whole-body actions shaped by barrier information available only during training. The central question is how much barrier structure a policy can internalize when it receives only deployment-realistic observations.
PAC-MAN studies two levels: Link-CBF, a lightweight per-link reward used during training, and Joint-CBF, a stronger whole-body joint-space projection that can either guide training or stay active at test time as a privileged runtime filter.
The deployed system segments the ball in RGB and masks the depth image down to a compact ball-only observation.
An adversarial motion prior (AMP) regularizes how the humanoid moves, so the evasions come out as crouches, leans, sidesteps, and leaps; the safety structure itself comes from the CBF terms.
The task is formulated as a partially observed Markov decision process where at control step t, the policy πθ(at ot) emits joint-position targets at ∈ R 29 (offsets from a nominal pose, tracked by PD actuators at 50 Hz).
The defining constraint is deployability: the policy observation ot contains only signals computable on hardware from onboard sensing, ot = Dt, qt, q˙t, ωt, gt, at−1,
where Dt is the depth observation.
The reward function is structured as: rt = r core + r cbf + λs r style,
where the task reward pairs the distance-to-core evasion term r core of prior humanoid dodgeball work [15] with th[e control barrier term r cbf].
The CBF terms are implemented at two levels:
-
Link-CBF is a lightweight per-link reward that extends the pelvis-centered task objective to every body link.
It penalizes violations of the barrier conditionh˙i + αhi ≥ 0 at the most-binding link,
whereh i =∥p b − pi∥ − ρ b + ρi
defines the clearance for every robot link i. -
Joint-CBF adds a stronger whole-body joint-space CBF module during training.
This module selects the most-threatened point i⋆ and enforces a constraint on the joint velocity:η⊤v ≤ b,
where η is derived from the positional Jacobian of the most-threatened point. The corresponding barrier-projected velocity isv⋆ = v des − η⊤v des − b +∥η∥2 η.
The policy training utilizes an asymmetric actor-critic paradigm where the value function Vϕ(st) observes a privileged state st that augments the policy observation ot of Eq. (1) with the ground-truth ball relative position/velocity, radius, a visibility gate, and full-body link kinematics.
These privileged signals are used for training and barrier computation but are discarded at deployment.
The perception regimes studied include:
State oracle. A symmetric actor-critic that observes the ground-truth ball states directly (no camera). This is the safety-information upper bound.
Fixed depth camera. The focusing stage alone: the ball is segmented from the RGB-D stream of a single rigidly mounted head camera...
**"Active gimbal (gaze). The second part of human-like perception added on top of the focusing stage: we mount the same camera on a 1-DoF pitch joint... so that it can track the ball through its flight, in analogy to human gaze.
Improvements for AI systems
Here are the specific improvements and capabilities derived from the PAC-MAN framework, tailored for enhancing real-world, perception-constrained humanoid robotics:
-
Replacement of Pelvis-Centric Safety with Whole-Body Link Safety: The system can now maintain safety across all body segments (limbs, torso) rather than just protecting the base.
-
Perception-Aware Control Barrier Functions (CBF): The policy can infer and enforce collision avoidance based on partial or noisy onboard visual input (segmentation-masked depth), allowing for proactive evasion even when a ball is partially obscured or briefly leaves the camera's field of view.
-
Perception-Conditioned Safety Structure: The system's safety mechanism dynamically adjusts its strength based on the quality of perception:
-
Deployment Robustness (Link-CBF): The robot can successfully evade 95% of physical throws in real-world scenarios using only onboard depth and proprioception, demonstrating that a lightweight, per-link barrier structure is sufficient for deployment when accurate ball states are unavailable.
-
Scalable Threat Representation: The system can generalize its evasion to different object types (e.g., dodging a soccer ball vs. a foam ball) by utilizing semantic segmentation of the incoming threat, allowing the same policy to adapt its evasion strategy based on the visual appearance of the object.
-
Regularization for Natural Motion: By incorporating an Adversarial Motion Prior (AMP), the system learns evasive reflexes that are naturally dynamic—such as crouches, leans, and sidesteps—rather than arbitrary joint movements, resulting in more human-like and energy-efficient motion.
-
Adaptive Safety Enforcement (Joint-CBF): During training, the policy can be guided by a Joint-CBF structure that enforces constraints across the entire joint space. This allows the learned policy to internalize a much stronger safety structure than it could achieve with limited observation alone, bridging the gap between training guidance and online enforcement.
-
State-of-the-Art Performance: The improved system can achieve near-optimal safety performance (98% success rate) when equipped with an oracle providing ground-truth ball states, proving that the primary bottleneck is perception, not the control theory itself.
The improved AI system can perform these functions:
The resulting system is a humanoid robot capable of performing high-speed, whole-body evasion maneuvers in unpredictable environments (like human dodgeball) using only a standard head-mounted camera and its internal sensors. It will reliably avoid contact with incoming projectiles by inferring the threat geometry from depth data, coordinating the entire body to move threatened links out of the path, and maintaining dynamic balance—all without requiring perfect, privileged knowledge of where the ball is at all times.
Sources
- GMT: General Motion Tracking for Humanoid Whole-Body Control
- SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control
- BeyondMimic: From Motion Tracking to Versatile Humanoid Control via Guided Diffusion
- VisualMimic: Visual Humanoid Loco-Manipulation via Motion Tracking and Generation
- CReF: Cross-modal and Recurrent Fusion for Depth-conditioned Humanoid Locomotion
- Egocentric Tactile and Proximity Sensors as Observation Priors for Humanoid Collision Avoidance
- Walk the PLANC: Physics-Guided RL for Agile Humanoid Locomotion on Constrained Footholds
- Retargeting Matters: General Motion Retargeting for Humanoid Motion Tracking
- Perceptive Humanoid Parkour: Chaining Dynamic Human Skills via Motion Matching
- Deep Whole-body Parkour
- BAT: Balancing Agility and Stability via Online Policy Switching for Long-Horizon Whole-Body Humanoid Control
- TAGA: Terrain-aware Active Gaze Learning for Generalizable Agile Humanoid Locomotion
- RSL-RL: A Learning Library for Robotics Research
- mjlab: A Lightweight Framework for GPU-Accelerated Robot Learning
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving