CSF: Contextual Safety Filtering for Motion Generators

arXiv:2610.12467 · cs.RO, cs.LG · Submitted 2026-10-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "CSF: Contextual Safety Filtering for Motion Generators".

Dev: The gist The Contextual Safety Filtering (CSF) introduces a training-free filter that grounds natural-language safety rules in safe and unsafe reference trajectories produced by motion generators,

Rosa: First, who's behind it and why it matters.

Paper summary: Rosa: So we're looking at this paper called "CSF: Contextual Safety Filtering for Motion Generators," and it sounds like they're tackling a problem where motion generators create whole-body movements, but these movements don't understand the context of the scene.

Dev: That’s right. The core issue is that the same action can mean something totally different depending on whether it’s aimed at an object or a person, and existing safety checks just don't handle how the scene itself changes what a motion means for safety.

Taro: So, before we get into the technical stuff, what is this CSF system actually claiming? What’s the big idea here?

Rosa: The main claim is that they introduce a training-free filter called CSF that grounds natural-language safety rules directly into safe and unsafe reference motions produced by the generator. This means for every rule you set up, they generate one unsafe motion and one safe motion based on your command.

Dev: And those two motions define an affine safety value that a specific type of optimization problem called a CBF-QP enforces to keep the generated motion safe. They claim this works across four different pretrained generators with different architectures, showing it reduces danger events by up to ninety percent while keeping eighty-eight to one hundred percent of benign motions <ref:2610.12467#pg1,pretrained generators with different architectures>.

Taro: Ninety percent reduction is quite a big number for safety in this area. But what’s the mechanism behind how it manages that context dependency?

Rosa: They achieve this by defining a semantic margin, which is essentially the difference between that unsafe reference and the safe reference motion based on your current command. This difference defines a semantic direction, which then tells the QP how far away from safety you are.

Dev: The mathematical formulation involves a CBF-QP with this safe-reference tracking objective over that affine semantic margin, and they prove per-step feasibility, invariance, and optimality for any tracking strength in their work.

Taro: So when the world misbehaves—say, someone suddenly enters the scene—how does this filter react in real time? Does it just stop everything?

Rosa: They show a real-world deployment on a Unitree G1 robot where they have a runtime shield that redirects an executing action as soon as a person enters view. This demonstrates that the system successfully prevents unsafe motions in scenarios involving interactions with both humans and objects.

Dev: And you mentioned it keeps the skill preserved? How does it handle explicit and scene-triggered unsafe cases on that hardware?

Taro: They show that on this robot, these rules block explicit strikes directed at a person while still preserving the same skill in a benign scene. It sounds like it’s selective about when it intervenes.

Paper summary: Rosa: Exactly. And they showed that for benign motion, eighty-eight to one hundred percent of clips either pass through or reach the selected safe counterpart, which is a pretty strong result if you ask me.

Dev: So, looking at the title "CSF: Contextual Safety Filtering for Motion Generators," what does that really imply for how we think about motion generation systems?

Taro: It implies that safety can be built into the filtering layer itself by using natural language rules and scene context to constrain the motion without needing to train a separate auxiliary model.

Rosa: It suggests a way to get safety information directly from the generator's output space, rather than trying to patch it on after it’s already made something. It’s about grounding those safety rules in what the generator is actually producing.

Dev: And from an engineering standpoint, they spend time on inference cost too. They found that for rule-space dual, it takes between zero point zero one and zero point one eight milliseconds for one to six active rules, but a full filter call can take between zero point three and one point two milliseconds on an RTX four thousand ninety GPU.

Taro: That latency is important for real-time control, especially when you're dealing with fast movements like walking or driving. The cost seems manageable for deployment on hardware like the Unitree G1, which they tested on a real robot.

Rosa: The paper shows that they cache references after one-time setup, which adds no measurable per-sample overhead in the fastest cases and at most nineteen percent in the slowest ones. That caching part seems to be key for keeping it practical.

Dev: So, what about the limitations of this CSF approach? Where does it stop working or what’s not covered by this method?

Taro: The authors state that they don't extend CSF to the root trajectory yet, meaning if a rule becomes active during execution, the robot doesn't get to redirect its locomotion or pose while that rule is in effect.

Rosa: That’s a limitation because it means you can’t have that kind of dynamic redirection during an active safety event right now. It's more of a shield that redirects the current action rather than controlling the overall movement strategy.

Dev: So, to wrap up on "CSF: Contextual Safety Filtering for Motion Generators," what is the final practical meaning for someone just listening to this show who isn't deep in the math?

Taro: It means you can take a motion generator that doesn't inherently know about scene safety and layer a filtering system on top that uses simple natural language rules to stop dangerous actions, and it actually works without needing any extra training data or models.

Rosa: It’s about taking the unsafe and safe motions the generator produces and turning them into scene-selected margins enforced by this specific optimization problem, which is a very direct way to control what the motion does.

Conclusion: Rosa: So we're wrapping up on CSF, which stands for Contextual Safety Filtering for Motion Generators.

Dev: Yeah, let's just talk about what this paper actually says about the whole thing and who put it together.

Taro: I saw the authors are Kimodo and their team, focusing on how to make these motion generators safer without needing to retrain anything.

Rosa: Right, so they're taking these text-conditioned motion generators that just spit out movement, like a robot trying to walk or move an arm.

Dev: And what they did was build this system that uses natural language rules—not training data—to filter those movements in real time based on the scene.

Taro: The big idea is using the difference between a safe motion and an unsafe one to define a safety margin that guides the generation process.

Rosa: It basically lets you feed in a rule, and it generates two versions of that action, one safe and one unsafe, which then tells the system how far off track you are.

Dev: And they use this difference to make sure the motion stays within bounds using this specific type of optimization problem called a CBF-QP.

Taro: It’s not just about stopping every bad move; it’s about making sure that if the scene changes, the filter automatically adjusts which rules matter.

Rosa: That's what they showed on a real robot, where you have this shield that stops an action instantly when someone walks into view.

Dev: The numbers are pretty solid too; they saw a danger reduction of up to ninety percent across four different motion generators they tested.

Taro: It means for people who just listen, it suggests we can build safety directly into the filtering layer using just language and scene awareness.

Rosa: It's about grounding those safety rules in what the generator is actually producing, rather than trying to patch it on later.

Dev: But there are still things they didn't tackle yet, like making sure the robot can redirect its whole movement strategy when a rule becomes active during execution.

Taro: So while this is a strong safety filter for immediate actions, we still need to figure out how to make those safety rules change the overall path of the motion itself.

Lizhi Yang, Yiling Hou, Yao Tang, Junheng Li, Daniel Weng, Blake Werner, Aaron D. Ames

California Institute of Technology · New York University

cs.RO, cs.LG

Submitted: 2026-10-08

Updated: 2026-10-08

The gist: The gist The Contextual Safety Filtering (CSF) introduces a training-free filter that grounds natural-language safety rules in safe and unsafe reference trajectories produced by motion generators,

Key concepts

Semantic Margin
This is the mathematical difference between an unsafe reference trajectory and a safe reference trajectory derived from a benign command. It defines a specific direction in motion space that represents the boundary between acceptable and unacceptable actions. This margin is used by the tracking controller to guide the motion toward safety.
Safe-Reference Tracking CBF-QP
This is an optimization framework that uses a quadratic program (QP) to enforce safety. It tracks a safe reference trajectory while minimizing the deviation from it, constrained by a barrier function derived from the semantic margin. This ensures that the generated motion stays close to what is considered safe.
Semantic Context Gate
This mechanism intelligently selects which safety rules apply based on both the command prompt and visual scene entities. It activates specific protected-entity rules only when the command is near an unsafe description or a detected entity matches a protected label, making the filtering process context-aware.

Terminology

Summary

The gist The Contextual Safety Filtering (CSF) introduces a training-free filter that grounds natural-language safety rules in safe and unsafe reference trajectories produced by motion generators, reducing danger-event rates by up to 90% while preserving 88–100% of benign motions

Problem Setting

Pretrained text-conditioned motion generators produce trackable whole-body motions, but they have no notion of scenedependent safety because the same action may target an object or a person Existing safeguards either inspect the prompt, require labeled motion data, or enforce geometric constraints; therefore, they do not directly account for how scene context changes a motion’s meaning. A whole-body controller will track either reference, so the safety decision must account for the scene as well as the motion. Existing safety mechanisms address only parts of this distinction, such as language guardrails being bypassed by adversarial prompts or prompt-only screens not verifying whether a generated motion is directed toward a benign object or a person. CSF specifies unsafe behavior directly through natural-language rules.

Semantic References and Margin

CSF constructs the semantic margin by having the frozen generator produce an unsafe reference x(i)unsafe from the rule description and a safe reference xsafe from a benign version of the current command that preserves the rest of the requested motion. The difference x(i)unsafe − xsafe defines a semantic direction, and this difference is used to define an affine safety value that a safe reference tracking CBF-QP enforces. The semantic margin is defined as the negative projection of the current estimate’s displacement from the safe reference onto the unsafe direction. This margin has three direct properties, including mi(xsafe) = 0.

Safe-Reference Tracking CBF-QP

The system formulates a discrete-time CBF-QP with a safe-reference tracking objective over an affine semantic margin. The formulation is given by xˆfilt = arg min x − (1 − γ)ˆx + γxsafe s.t. hi(x) ≥ (1 − ρt) hi(ˆx), i ∈ A. Here, ρt ∈ [0, 1] is the barrier contraction rate and γ ∈ [0, 1] weights the safe-reference tracking objective. At γ = 0, this is the minimum-norm correction of xˆ and leaves feasible proposals unchanged. Theorem 1 proves per-step feasibility, invariance, and optimality for any tracking strength.

Semantic Context Gate

The semantic context gate uses the command and scene entities to select which protected-entity rules the QP should evaluate. The prompt signal activates the context-sensitive rule set when the command is closest to an unsafe description rather than a benign contrast. The perception signal is rule specific and activates rule i when the label of a visually detected entity is closest to the entity protected by that rule. Concept i is active when either signal fires, defined as gi = g perception i ∨ g prompt i, A = (i: gi = 1).

Application Across Generators

CSF interfaces with each generator through the sampler’s current estimate of its final output. For direct motion-space predictions, the full filtered estimate is FA(ˆx) = MQ M A (ˆx) + (I − M)ˆx. For decoder-aware latent predictions, a worst active margin me k(z) is found and the estimate z is updated via z ← z − η me k(z)∥q∥2 + ε q (8) until the barrier is met or an iteration budget N is exhausted. Finally, a runtime safety shield applies the same rules to the remaining reference when context changes during execution.

Real-World Deployment

The complete perception, generation, and tracking pipeline is deployed on a real-world Unitree G1 where a runtime shield redirects an executing action when a person enters view. The system successfully prevents unsafe motions in scenarios including interactions with humans and objects. CSF turns generator-produced unsafe and safe references into scene-selected affine margins enforced by a safereference tracking CBF-QP. Across four pretrained generators, residual danger falls by up to 90% with better preservation than a standing-pose blend. On a real-world Unitree G1, the same rules block explicit and scene-triggered person-directed punches while preserving the same skill in the benign scene. The generator remains frozen during both offline generation and closed-loop execution. Future work includes extending CSF to the root trajectory, letting the robot redirect its locomotion as well as its pose when a rule becomes active.

Inference-Time Cost

On one RTX 4090, the rule-space dual takes 0.01 to 0.18 ms for one to six active rules, and a complete filter call takes 0.3 to 1.2 ms. ECHO, MotionHiFlow, and ARDY cache their references after a one-time setup of 0.6 to 0.9 s, adding no measurable per-sample overhead in the fastest case and at most 19%. Kimodo constructs step-matched references during sampling, increasing the time needed to generate a four-second motion from 1.61 to 7.60 s, although the QP itself accounts for only 0.14 s.

Benchmark Results

The reported danger ratio divides the filtered event rate by its unfiltered counterpart, so 0.1 represents a 90% relative reduction. Residual danger falls to 0.1 of the unfiltered rate on three backbones and 0.2 on ECHO. Kimodo’s distinct motion-space margin supports minimum intervention, while MotionHiFlow’s nearly flat latent margin needs a stronger safe-reference pull. The gate selects the intended rule for every explicit and scene-triggered unsafe case across all four backbones and leaves it inactive in the matched allowed scene. For benign motion, 88 to 100% of clips either pass through or reach the selected safe counterpart. On Kimodo, the correction is smaller than the natural variation between two unfiltered samples generated from the same command.

Conclusion

CSF turns generator-produced unsafe and safe references into scene-selected affine margins enforced by a safereference tracking CBF-QP. Across four pretrained generators, residual danger falls by up to 90% with better preservation than a standing-pose blend. On a real-world Unitree G1, the same rules block explicit and deceptive unsafe commands, preserve a benign object-directed skill, and interrupt an executing strike when a person enters view. The generator remains frozen during both offline generation and closed-loop execution. The system provides training-free safety filtering that uses naturallanguage rules, generator-produced reference motions, and scene context to constrain generated motion without training an auxiliary model.

References

[1] D. Rempe, M. Petrovich, Y. Yuan, H. Zhang, X. B. Peng, Y. Jiang, T. Wang, U. Iqbal, D. Minor, M. de Ruyter, et al., “Kimodo: Scaling controllable human motion generation,” arXiv preprint arXiv:2603.15546

[2] H. Jia, J. Song, Y. Zhang, H. Jin, Y. Fan, W. Chen, W. Zhang, and Y. Yue, “Echo: Edge-cloud humanoid orchestration for language-tomotion control,” arXiv preprint arXiv:2603.16188

[3] H. Li, X. Lin, L.-A. Zeng, Y. Kang, S. Li, and J.-F. Hu, “Motionhiflow: Text-to-motion via hierarchical flow matching,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 9352–9363

[4] K. Zhao, M. Petrovich, H. Zhang, T. Wang, S. Tang, and D. Rempe, “Autoregressive diffusion with hybrid representation for interactive human motion generation,” ACM Transactions on Graphics (TOG), vol. 45, no. 4, pp. 1–14

[5] A. Robey, Z. Ravichandran, V. Kumar, H. Hassani, and G. J. Pappas, “Jailbreaking LLM-controlled robots,” in IEEE Int. Conf Robot Autom (ICRA), 2025

[6] K.-C Hsu, H.-H Hu, and J.-F Fisac, “The safety filter: A unified view of safety-critical control in autonomous systems,” Annual Review of Control, Robotics, and Autonomous Systems, vol. 7

[7] Z. Luo, Y. Yuan, T.

Improvements for AI systems

  1. The system can produce trackable whole-body motion while incorporating contextual safety filtering (CSF), a training-free filter that grounds natural-language safety rules in safe and unsafe reference trajectories produced by the generator. This allows generated actions to be constrained by scene context without requiring auxiliary models or retraining.

  2. The AI system can enforce semantic constraints using a CBF-QP with a safe-reference tracking objective over an affine semantic margin, which defines the barrier function as hi(x):= mi(x) and uses the constraint set S = “the intersection of the active concepts’ superlevel sets, S = [x: hi(x) ≥ 0, i ∈ A]” to ensure safety.

  3. The system can dynamically adapt its constraints based on scene entities via a Semantic Context Gate defined by gi = g perception i ∨ g prompt i, meaning Concept i is active when either signal fires, A = [i: gi = 1], which supplies the constraints to Eq. (3). This enables the system to select which safety rules are enforced based on both the command and perceived entities.

  4. The system can perform real-time safety intervention during execution through a Runtime Safety Shield that "selects the first violating unexecuted window and updates the reference by j⋆ = min j: sj ≥ l, min i∈Al hi(R[Wj]) < τ," preserving the executed portion while redirecting the remaining reference when context changes.

  5. The system can be deployed across diverse architectures, as demonstrated by evaluating CSF on four pretrained generators with different architectures, showing that it can interface effectively with Kimodo-G1 [1], ECHO [2], MotionHiFlow [3], and ARDY [4].

Sources

Related papers