PAC-MAN: Perception-Aware CBF-RL for Whole-Body Safety in Humanoid Dodgeball
summary
The gist
We present PAC-MAN, a perception-aware CBF-RL framework that couples control-barrier safety with deployment-realistic onboard sensing for whole-body humanoid dodgeball.
In short
The episode discusses PAC-MAN, a paper presenting a perception-aware Control Barrier Function Reinforcement Learning framework for whole-body safety in humanoid dodgeball. The hosts explore how this method bridges safety guarantees with real-world sensing limitations, focusing on using depth images to enable robust evasion strategies even with imperfect onboard perception.
Key concepts
- PAC-MAN
- A perception-aware Control Barrier Function Reinforcement Learning framework designed for whole-body safety in humanoid dodgeball. It couples control barrier safety with deployment-realistic sensing to ensure the robot stays safe during evasive maneuvers.
- Control Barrier Functions (CBF)
- A method used to provide safety guarantees in control systems. In this paper, they use different levels of CBF enforcement: Link-CBF for deployment and Joint-CBF for training guidance, balancing safety with hardware constraints.
- Perception-Awareness
- The ability of the policy to adjust its safety structure based on the quality of onboard perception. The system dynamically adjusts the strength of the safety barrier depending on how accurate or noisy the visual input is during operation.
Terminology used across episodes
This episode discusses
- PAC-MAN: Perception-Aware CBF-RL for Whole-Body Safety in Humanoid Dodgeball · Paper Radio
- GMT: General Motion Tracking for Humanoid Whole-Body Control
- SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control
- BeyondMimic: From Motion Tracking to Versatile Humanoid Control via Guided Diffusion
- VisualMimic: Visual Humanoid Loco-Manipulation via Motion Tracking and Generation
- CReF: Cross-modal and Recurrent Fusion for Depth-conditioned Humanoid Locomotion
- Egocentric Tactile and Proximity Sensors as Observation Priors for Humanoid Collision Avoidance
- Walk the PLANC: Physics-Guided RL for Agile Humanoid Locomotion on Constrained Footholds
- Retargeting Matters: General Motion Retargeting for Humanoid Motion Tracking
- Perceptive Humanoid Parkour: Chaining Dynamic Human Skills via Motion Matching
- Deep Whole-body Parkour
- BAT: Balancing Agility and Stability via Online Policy Switching for Long-Horizon Whole-Body Humanoid Control · Paper Radio
- TAGA: Terrain-aware Active Gaze Learning for Generalizable Agile Humanoid Locomotion · Paper Radio
- RSL-RL: A Learning Library for Robotics Research
- mjlab: A Lightweight Framework for GPU-Accelerated Robot Learning
The paper
PAC-MAN: Perception-Aware CBF-RL for Whole-Body Safety in Humanoid Dodgeball · Read on arXiv
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "PAC-MAN: Perception-Aware CBF-RL for Whole-Body Safety in Humanoid Dodgeball".
Dev: We present PAC-MAN, a perception-aware CBF-RL framework that couples control-barrier safety with deployment-realistic onboard sensing for whole-body humanoid dodgeball.
Rosa: First, who's behind it and why it matters.
Title and authors: Rosa: So we’re talking about PAC-MAN today, which is titled "Perception-Aware CBF-RL for Whole-Body Safety in Humanoid Dodgeball." It’s clear from the title that they are focusing on how to make a humanoid robot safe during dodgeball using perception and control barrier functions.
Dev: That title tells you immediately that the core contribution is bridging the gap between safety guarantees provided by control barrier functions and the actual sensing limitations of a deployed system. It’s about making sure it stays safe even when things aren't perfect.
Taro: I think focusing on whole-body safety is important because most robotics papers focus on just keeping the base stable, but PAC-MAN seems to push for safety across every single link of the robot during an evasive maneuver.
Rosa: Precisely, and what they do in simple terms is they train the policy to understand how to move the whole body based on partial, realistic camera input while using training guidance that covers all joints. They want this policy to be robust enough for real-world use where you don't have perfect data.
Dev: The authors are Lizhi Yang, Junheng Li, and Aaron D. Ames, and they’ve clearly done some deep work on the mechanics of how these perception constraints interact with learning safety policies. I think their background in both robotics and control engineering is what makes this paper feel so grounded in reality.
Taro: I'm interested in how they framed the problem as a partially observed Markov decision process; it sounds like they are treating this as a real-time decision-making problem under uncertainty rather than just a static planning task.
Rosa: That’s right, and that formulation is what lets them talk about joint-position targets being emitted at control step t, which is essential for any system running on physical hardware with time constraints.
Dev: And the observation constraint they put on the policy observation o t is very specific: it must only contain signals computable on hardware from onboard sensing, which sets a hard limit on what kind of information the AI can rely on during operation.
Taro: That constraint really highlights the deployment reality; if you can't compute something on-board, the policy simply can't use it to make decisions at runtime.
Rosa: Exactly, and that’s why they have to carefully design what information is fed into the system versus what is only used during training guidance.
Dev: This leads us into the core of their methodology where they introduce Link-CBF and Joint-CBF as different levels of safety enforcement during training versus runtime.
The paper's summary: Rosa: To summarize PAC-MAN, this perception-aware CBF-RL framework couples control barrier safety with deployment-realistic sensing for whole-body humanoid dodgeball. It basically says the robot learns to avoid being hit by a ball by using depth images from a head camera as its main input.
Dev: The key summary point I see is that the training guidance uses clearance information for every body link, but in deployment, the policy only sees segmentation-masked depth, and they use an adversarial motion prior to shape those evasive reflexes into more natural movements like leaning or sidestepping.
Taro: So it’s not just about learning a dodgeball avoidance strategy; it’s about making sure that whatever the policy learns is physically executable and safe across the entire body structure, which is a significant step up from simpler approaches.
Rosa: That’s right, and they evaluate this on two distinct regimes: single throws for quick reactions and a deployment loop where the robot has to walk back to its station and recover between throws.
Dev: The summary also highlights that they found that usable barrier structure depends on perceptual observability; specifically, Joint-CBF performs best when accurate ball states are available, whereas Link-CBF is the most deployable option under limited onboard perception.
Taro: That means the system's performance is directly tied to how well you can track or infer the threat geometry from just that depth data during operation.
Rosa: It seems like they’ve really established a clear hierarchy of safety mechanisms: training uses the whole-body view, but deployment relies on a lighter per-link barrier structure for practical success.
Dev: And they explicitly state that they validate this design choice by deploying the fixed-camera Link-CBF policy zero-shot on a Unitree G1 using only onboard depth and proprioception to prove its deployability.
The paper's improvements: Rosa: One major improvement they highlight is replacing the reliance on just the pelvis-centered safety with a whole-body link safety across all body segments, which means maintaining safety throughout the torso and limbs.
Dev: That’s a big step because limiting it to just the base can lead to instability when you have dynamic interactions or external forces affecting other parts of the body. It shows they are thinking about holistic stability.
Taro: And I think incorporating perception-aware CBF is important because it allows the policy to infer and enforce collision avoidance based on partial or noisy visual input, meaning it can be proactive even if a ball is partially blocked.
Rosa: Right, and then they have this perception-conditioned safety structure that dynamically adjusts its strength based on the quality of perception you are getting during operation. If the vision is bad, the system adjusts accordingly.
Dev: That dynamic adjustment sounds like a very smart way to handle uncertainty; it means you don't rely on a fixed barrier when the input quality changes and perhaps that's a failure mode they've mitigated.
Taro: I also think the adversarial motion prior is valuable because it ensures that the evasive reflexes aren't just arbitrary joint movements but are instead dynamic behaviors like crouching or leaning, which makes the movement more realistic.
Rosa: That regularization helps ensure that when the policy does try to evade, it does so in a way that is actually feasible for a humanoid robot to execute efficiently.
Dev: So, the paper suggests that for deployment, Link-CBF is the best deployable option under limited onboard perception because it balances safety with hardware limitations effectively.
Conclusion: Rosa: To wrap up on PAC-MAN today, the main implication is that perception-aware control barrier methods allow robots to operate safely in dynamic, unpredictable environments like dodgeball without needing perfect information.
Dev: I think the impact is showing that as long as you can design a safety structure that adapts to observation quality, you can achieve high success rates even when operating under real-world perception constraints.
Taro: The takeaway for autonomy research is that the bottleneck in achieving robust evasion isn't necessarily the control theory itself, but rather how much information a policy can internalize when it only receives deployment-realistic observations.
Rosa: Right, and they’ve shown that this approach works well on benchmarks where the policy comes within a few points of an oracle providing perfect state knowledge, suggesting perception is indeed the main hurdle to solving these kinds of problems.
Dev: And for me, it’s about the loop rate; if the deployment requires constant re-evaluation based on imperfect depth data, we need to ensure that latency doesn't cause a failure mode during those critical evasion steps.
Taro: I just think as long as you can quantify that performance against an oracle, like achieving ninety-five percent success rates in real-world scenarios, it validates the method for practical application.
Rosa: Exactly; the PAC-MAN paper gives us a solid foundation on how to design these systems to be robust against perception limitations and ready for deployment on hardware like the Unitree G1.
More episodes
- 2610.11768-Narrow and Deep: An Ontology Tower as the Knowledge of an LLM Agent for an Industrial Equipment System
- 2610.11904-Large-Scale Partition-Based RIS Beamforming For Uplink RIS-Equipped Multi-User Systems: Asymptotic Analysis
- 2610.11885-Redefining fuel poverty: Introducing the temporal equity framework (TEF)
- 2610.11900-Reach-Stabilize Control of Control-Affine Systems with Unknown Affine Parameters
- 2610.11964-From Asymptotic to Designer-Assigned-Time Control: A Review of Stability Notions, Design Mechanisms, and Controller Architectures
- 2610.12226-Stabilization of Unidirectional First-Order PDE-ODE Coupled Systems with Boundary and Distributed Input Delays
- 2610.12028-Policy Synthesis for Finite Populations of MDP Agents under Aggregate Reach-Avoid Chance Constraints
- 2610.12103-Predefined-Time Integral Reinforcement Learning for Unknown Nonlinear Systems via Inverse-Optimal Design
- 2610.12110-Adaptive dynamic programming using Lyapunov function constraints
- 2610.12324-Convex Safety Filtering via Spectral Selection for Nonconvex Safe Sets