BabelSafe: A Policy-Grounded Multilingual Safety Benchmark for LLMs

arXiv:2605.00689 · cs.CL, cs.CR · Submitted 2026-05-01 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "BabelSafe: A Policy-Grounded Multilingual Safety Benchmark for LLMs".

Tom: ML-BENCH and ML-GUARD introduce a policy-grounded multilingual safety benchmark and guardrail model, respectively, to address the limitations of existing multilingual safety evaluations that rely on general risk taxonomies.

Jane: First, who's behind it and why it matters.

Paper summary: Tom: So, wrapping up our discussion on BabelSafe: A Policy-Grounded Multilingual Safety Benchmark for LLMs, it seems the main message is that existing multilingual safety evaluations are too broad because they rely too much on general risk taxonomies. The authors addressed this by creating ML-BENCH and ML-GUARD, which are designed to be policy-grounded and language-specific.

Jane: Precisely; the core claim is that this method moves beyond just language transfer by grounding safety assessments directly in region-specific regulations and native-language contexts, allowing for evaluation that is actually aligned with local legal realities.

Lu: The significance lies in how they manage the data pipeline, ensuring that risk categories and rules are expressed directly in their native languages, which eliminates the need for machine translation at any stage when generating the safety queries and responses.

Meng: From a practical standpoint, this means that companies building cross-linguistic AI tools have a much more accurate way to test compliance with actual laws in different jurisdictions instead of just guessing based on general guidelines.

Lalam: I think the real impact is that it sets a new standard for how we approach safety evaluation—it suggests that for truly global deployment, the guardrails need to be rooted in local policy, not just generic AI principles.

Tom: It’s definitely about making safety checks more relevant and actionable in diverse global settings; the paper shows ML-GUARD achieves high performance on both their new benchmark and existing ones, which validates the approach's effectiveness.

Jane: The implication is that we can start developing regulation-aware and culturally aligned multilingual guardrail systems, which is a massive step forward for responsibly deploying these powerful models everywhere.

Conclusion: Tom: So, we've been deep into the details of BabelSafe, and now it’s time to talk about what this whole project actually means for us as listeners and as AI developers out there.

Jane: I think we should start by just saying what "BabelSafe: A Policy-Grounded Multilingual Safety Benchmark for LLMs" actually tells us in plain English.

Lu: From a conceptual standpoint, it’s about moving safety away from just general guidelines and anchoring it directly into the actual legal frameworks of different countries.

Meng: I'm curious how this practical application will look when we start building AI systems that need to operate across multiple legal zones simultaneously.

Lalam: For me, the biggest vision here is that this work could fundamentally improve how we design and build AI to be culturally sensitive, not just technically compliant.

Tom: Exactly, Lalam. It moves the conversation from abstract principles to concrete regulations, which is a huge step for anyone dealing with real-world deployment issues.

Jane: And speaking of those regulations, who are the folks behind this work? Getting to know the authors can give us a sense of their background in this area.

Lu: The team behind it brings together expertise from various regulatory and linguistic fields, which is what makes the construction of ML-BENCH so thorough.

Meng: I’m looking at their methodology, specifically how they built that data pipeline; does it feel like something we could actually replicate in a production setting?

Lalam: The pipeline is clever because it forces the system to deal with native language phrasing and local risk categories, which is where the real cultural alignment happens.

Tom: That's a great point about replication, Meng. It’s not just about building a benchmark; it’s about creating a blueprint for more context-aware safety systems.

Jane: So, to wrap up this segment, BabelSafe isn't just another test; it's providing the necessary structure to align AI with local legal realities across many languages.

Lu: And that alignment potential opens up so many creative avenues for how AI can interact with diverse human cultures legally and ethically.

Meng: It certainly suggests that future guardrails won't work if they aren't built on this kind of localized, rule-based foundation.

Lalam: It really shows that the most impactful advances in AI safety are those that respect the specific cultural and legal nuances of every region they operate in.

Yunhan Zhao, Zhaorun Chen, Xingjun Ma, Yu-Gang Jiang, Bo Li

University of Illinois Urbana-Champaign · Fudan University · University of Chicago

cs.CL, cs.CR

Submitted: 2026-05-01

Updated: 2026-09-27

Code: https://github.com/meta-llama/PurpleLlama

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 82/100

The gist: ML-BENCH and ML-GUARD introduce a policy-grounded multilingual safety benchmark and guardrail model, respectively, to address the limitations of existing multilingual safety evaluations that rely on

Key concepts

ML-BENCH
This is a multilingual safety benchmark created by systematically extracting rules from 17 regional AI regulations spanning 14 countries and 14 languages. It structures these extracted constraints into a hierarchy of high-level risks and fine-grained safety rules, ensuring the resulting data reflects native language legal expressions.
ML-GUARD
This is a diffusion Large Language Model (dLLM) guardrail designed for policy-grounded safety assessment. It comes in two versions: a lightweight 1.5B model for fast checks and a more capable 7B model that can perform standard safety checks and check compliance against specific regulatory rules, providing rationales for decisions.
Policy-Grounded Safety
This approach means safety assessments are not based on general risk ideas but are directly tied to specific, region-specific regulations. By grounding evaluations in native language contexts and legal texts, the system ensures that safety checks align with local laws and cultural expectations.
Attack-Enhanced Queries
These are specially crafted queries used in the benchmark. They involve appending adversarial text like 'reasoning distraction' to unsafe prompts to test how well guardrail models can be bypassed. This helps determine if a model can still evade safety checks even when the prompt is intentionally obscured.

Terminology

Summary

ML-BENCH and ML-GUARD introduce a policy-grounded multilingual safety benchmark and guardrail model, respectively, to address the limitations of existing multilingual safety evaluations that rely on general risk taxonomies. This work matters because it moves beyond language transfer by grounding safety assessments directly in region-specific regulations and native-language contexts, enabling culturally and legally aligned evaluation across diverse linguistic environments.

ML-BENCH Construction

ML-BENCH is a policy-grounded multilingual safety benchmark constructed by systematically extracting, organizing, and refining risk categories and safety rules from 17 regional AI regulations spanning 14 countries and 14 languages. The construction process involves two main stages: first, Article-Level Rule Extraction, where an LLM is prompted to identify concrete constraints from individual articles in a regulation; second, Language-Specific Risk Category Formation, where extracted rules are aggregated and structured into a two-layer hierarchy consisting of high-level risk categories and fine-grained safety rules. This ensures that the resulting ML-BENCH Risk Categories and Safety Rules are expressed directly in their native languages, preserving jurisdiction-specific legal expressions.

Multilingual Data Generation Pipeline

Building on the ML-BENCH structure, a staged data generation pipeline is used to create the benchmark dataset. This pipeline includes:

  1. Query Construction: Queries are constructed in three progressively challenging levels: seed queries, refined queries, and attack-enhanced queries. Seed queries are based directly on the risk categories and rules, while refined unsafe queries preserve policy-violating intent by rephrasing them into legitimate, professional, or bureaucratic inquiries using native-language phrasing.

  2. Response Construction: Paired safe and unsafe responses are constructed to be near regulatory decision boundaries. Unsafe responses are generated by directly prompting a model with refined unsafe queries, while safe responses are generated from those same queries but remain policy-compliant without explicit refusals.

  3. Attack Enhancement: Adversarial suffixes, such as reasoning distraction and risk category shifting, are appended to refined unsafe queries to create attack-enhanced settings designed to bypass guardrail models.

ML-GUARD Model Development

ML-GUARD is a Diffusion Large Language Model (dLLM)-based guardrail model designed for policy-grounded safety assessment. It comprises two variants:

  1. ML-GUARD-1.5B: A lightweight model optimized for fast 'safe/unsafe' checking suitable for latency-sensitive scenarios.

  2. ML-GUARD-7B: A more capable model that supports both standard safety assessment and policy-conditioned compliance checking, allowing users to specify regulatory rules, evaluate violations, and produce rationales to justify decisions. The dLLM architecture is chosen for its parallel inference capability, making it well-suited for structured outputs like policy violation judgments accompanied by rationales.

Evaluation and Performance

Extensive experiments were conducted against 11 strong guardrail baselines across six existing multilingual safety benchmarks and ML-BENCH. The results consistently show that ML-GUARD achieves state-of-the-art performance on both ML-BENCH and existing benchmarks. Specifically, for the Seed Query setting on ML-BENCH, ML-GUARD achieves an F1 score of 0.97, while the attack-enhanced setting shows that ML-GUARD reaches an accuracy of 0.92 for the 7B model. Furthermore, Table 3 demonstrates that ML-GUARD-7B accurately predicts violated rules across both ML-BENCH and existing benchmarks, achieving high F1 scores of 0.94 and 0.87 on Seed and Refined Prompts, maintaining a low false positive rate (FPR) in the Response setting.

Inference Efficiency

The paper also evaluates the inference efficiency of ML-GUARD-7B by comparing its performance against fine-tuned autoregressive models like Qwen2.5-7B. The diffusion-based architecture enables approximately 1.5× lower per-token latency and 1.5× higher throughput on average during inference, highlighting the efficiency benefits of this approach for guardrail tasks requiring structured outputs, such as those involving violated rule prediction with rationales. This efficiency is notably superior to comparable policy-aware baselines, achieving approximately 9× lower average latency than gpt-oss-safeguard-20B.

Conclusion

The work concludes that ML-BENCH captures complementary safety challenges beyond existing evaluations, and ML-GUARD provides an effective solution for policy-grounded multilingual guardrail. The findings suggest that this approach can advance the development of regulation-aware and culturally aligned multilingual guardrail systems. The paper also includes extensive human validation, confirming high agreement (94.3%) between human annotations and ground-truth labels when explanations are provided to the annotators.

Improvements for AI systems

Here are the specific improvements that can be made to AI systems based on this research:

  1. Improved Safety Alignment through Policy Grounding:

  2. Development of a Policy-Grounded Multilingual Safety Benchmark (ML-BENCH): This benchmark is constructed directly from regional regulatory texts (e.g., EU AI Act, US NIST framework) in their native languages, eliminating the reliance on error-prone machine translation for safety data.

  3. Creation of Policy-Conditioned Guardrail Models (ML-GUARD):

  4. Deployment of two specialized guardrail variants:

  5. ML-GUARD-1.5B: For fast, low-latency 'safe/unsafe' binary classification checks suitable for real-time applications like API gateways or quick input filtering.

  6. ML-GUARD-7B: For sophisticated, policy-conditioned compliance assessment that supports detailed explanations and rule violation identification (e.g., specifying which exact regulatory rules were broken).

  7. Enhanced Robustness to Adversarial Attacks: The system is trained on attack-enhanced queries designed to bypass existing guardrails, significantly increasing its resilience against jailbreaking and evasive prompting techniques.

  8. Contextually Aligned Safety Judgments: The models are fine-tuned on native language data derived from regional regulations (e.g., 14 languages), ensuring safety assessments respect specific cultural nuances and legal terminology that general-purpose models often miss.

  9. Dynamic Policy Adaptability: The ML-GUARD-7B architecture supports staged training updates, allowing the system to incorporate newly enacted or updated regional policies without suffering from catastrophic forgetting of previously learned rules.

  10. Enhanced Interpretability and Auditing: The ML-GUARD-7B model provides explicit, rule-based rationales for its decisions (e.g., Violated Rule: National Security…: Do not…use CBRN weapons.), which is crucial for regulatory compliance auditing and debugging safety failures in complex systems.

  11. Efficient Inference Performance: The use of a Diffusion Large Language Model (dLLM)-based architecture allows ML-GUARD-7B to achieve approximately 9× lower average latency compared to large, policy-aware baselines like gpt-oss-safeguard-20B while maintaining superior accuracy and reasoning capabilities.

  12. Multi-Rule Compliance Checking: The system can handle complex regulatory scenarios where a single input might violate multiple rules simultaneously, accurately predicting all violated rules and providing a rationale for each violation.

Sources

Related papers