CHILLGuard: Towards Fine-Grained Chinese LLM Safety Guardrail with Scalable Data Construction and Model-aware Preference Alignment

arXiv:2606.15396 · cs.CL, cs.AI · Submitted 2026-06-13 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "CHILLGuard: Towards Fine-Grained Chinese LLM Safety Guardrail with Scalable Data Construction and Model-aware Preference Alignment".

Jane: Malicious content generated by large language models (LLMs) poses severe safety risks, and existing guardrails often fail to adapt to Chinese cultural context and regulatory nuances.

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So, we're diving into the paper "CHILLGuard: Towards Fine-Grained Chinese LLM Safety Guardrail with Scalable Data Construction and Model-aware Preference Alignment." The main thesis here is that current safety guardrails aren't cutting it for the specific needs of Chinese content because they don't handle the cultural context and regulatory details correctly, which this paper claims to solve.

Jane: Exactly, Tom. Basically, the paper introduces a system called CHILLGuard that builds a dedicated safety guardrail tailored specifically for Chinese LLMs by creating a detailed risk taxonomy and using a multi-stage data construction pipeline to train it effectively.

Lu: What's really compelling about this approach is the fine-grained harm taxonomy they developed, which has five major macro categories and thirty-one micro categories that are explicitly aligned with Chinese regulations and linguistic traits, moving beyond generic English taxonomies.

Meng: From an engineering standpoint, I'm interested in how they managed to build such a detailed taxonomy when the foundational data might be sparse or biased towards certain types of harm.

Lalam: I think this fine-grained structure is crucial because it allows the AI to understand the specific risks within Chinese cultural contexts, which means we can develop safety mechanisms that actually resonate with local policies and user behaviors.

Tom: Right, so they are tackling a major gap in existing safety tools by creating this hyper-specific classification system for Chinese content. What's next is understanding how they actually built the data to feed this system.

Jane: They proposed a scalable multistage data construction pipeline that pulls from multiple sources to overcome the shortage of high-quality, native Chinese safety data, which is a big hurdle in this area.

Lu: That pipeline involves three main stages: using Retrieval-Augmented Generation for prompt construction, gathering real-world user prompts from production environments, and then augmenting the dataset through prompt engineering strategies that account for specific Chinese linguistic quirks.

Meng: I'm curious about the practicalities of that RAG part; how do you ensure that retrieving candidate texts is actually capturing the most relevant harmful queries without introducing noise into the training set?

Lalam: The retrieval-augmented generation stage uses embeddings from bgem3 to construct queries based on those macro and micro categories, which helps them pull out a diverse set of initial candidates before the prompt engineering stage refines them.

Tom: And then they move into prompt engineering to increase the diversity and implicitness of harmful samples by applying rewriting strategies that incorporate things like homophonic substitution or cultural allusion. That sounds like a clever way to stress-test the model's guardrail.

Jane: It’s interesting because that process generates a large dataset, specifically CHILLGuardTrain with four hundred five thousand seven samples and CHILLGuardTest with fifty-one thousand seven hundred forty-five samples, to ensure the system is robustly trained.

Paper summary: Lu: And to make sure the guardrail itself is strong enough to handle these diverse inputs, they utilize a generator-classifier collaborative framework employing Model-aware Direct Preference Optimization or MDPO for training.

Meng: MDPO sounds complex; how does adjusting the KL penalty based on real-time responsiveness actually translate into better detection robustness during the training process?

Lalam: The authors found that standard DPO has limitations, so they dynamically adjust the KL penalty by computing an implicit reward gap Ri, which is then normalized to estimate a responsiveness factor αM and set the coefficient βM = β·αM.

Tom: So they are essentially making the optimization process smarter by tying it directly to how responsive the model is to different training instances. That’s a sophisticated tuning mechanism for safety alignment.

Jane: And their iterative training process involves three iterations, where they first get an initial classifier C(zero), then generate augmented data D(one)gen, and finally fine-tune the generator G(zero) using preference pairs P(two) derived from that generated data to get the optimized G(one).

Lu: This iterative cycle is important because it allows the generator to learn from its own output quality scores assigned by a classifier, which refines the adversarial samples over time.

Meng: It’s impressive that they managed this multi-stage refinement process, but what about the limitations they mentioned regarding their dataset construction? They did note that they are still working on improving coverage for certain types of harmful queries.

Lalam: The authors acknowledge that while the system is state-of-the-art across all scales, the underlying data construction still faces challenges in completely capturing every possible nuance of online harm.

Tom: So, we've seen they’ve built a very strong framework with this CHILLGuard: Towards Fine-Grained Chinese LLM Safety Guardrail with Scalable Data Construction and Model-aware Preference Alignment paper, but the authors themselves admit that it’s not perfect yet regarding data coverage.

Jane: That means while the performance metrics are excellent, the system isn't claiming to cover every single possible edge case in Chinese online environments right now.

Lu: Thinking about the broader impact, if this framework successfully adapts safety tools to these specific regulatory and cultural nuances, it opens up a new level of responsible AI deployment within China.

Meng: If we can build guardrails that actually understand the cultural context rather than just translating rules, that means our models will be safer when deployed in complex local settings.

Lalam: For me, this advance has implications because it allows us to deploy AI systems that respect social norms and avoid generating content that causes specific cultural offense or regulatory violations within China.

Paper summary: Tom: It really does shift the focus from broad multilingual safety to deep, context-aware safety for specific regions. What are the big takeaways we should be thinking about as a team?

Jane: The main implication is that high-quality, culturally attuned data pipelines combined with model-aware training can lead to guardrails that perform significantly better than existing open-source options on benchmarks like CHILLGuardTest.

Lu: I see this as a step toward building truly localized and trustworthy AI systems where the safety mechanisms aren't just imported but are intrinsically shaped by the local environment.

Meng: From a practical standpoint, it means our engineering teams can focus less on manual rule-writing for every region and more on refining these sophisticated data generation pipelines to feed the models correctly.

Lalam: I believe this work supports a future where AI development in China is not just about technical capability but also about building systems that are inherently safe and compliant with local values through better safety alignment.

Tom: So, the paper "CHILLGuard: Towards Fine-Grained Chinese LLM Safety Guardrail with Scalable Data Construction and Model-aware Preference Alignment" shows a very robust method for fine-grained risk classification specific to Chinese contexts, built on scalable data and model-aware training techniques.

Jane: It’s about taking the complexity of Chinese content moderation and creating a structured, high-quality dataset to train guardrails that can handle those fine distinctions effectively.

Lu: The authors are showing how integrating RAG for scale, prompt engineering for diversity, and MDPO for robustness creates a powerful training loop.

Meng: I’m interested in how the 8B variant of CHILLGuard reaches an overall F1 score of eighty-nine point seven seven on the CHILLGuardTest, which they noted outperforms LlamaGuard and Qwen3Guard-8B-Strict by about fifteen point nine two percent.

Lalam: That performance metric shows that this tailored approach delivers tangible improvements over existing mainstream open-source guardrails when tested against a comprehensive set of Chinese safety benchmarks.

Tom: Exactly, so the paper’s core contribution is demonstrating that this specialized methodology yields state-of-the-art results across different model sizes for this specific safety problem.

Jane: The title itself really tells you it's about achieving fine-grained control over safety specifically for Chinese LLMs through a combination of scalable data and preference alignment.

Lu: The paper’s architecture suggests that future work could explore how this framework could be adapted to other complex cultural or regulatory domains outside of the immediate focus.

Meng: For practical application, it suggests we need to prioritize building robust, multi-source data pipelines early on if we want our deployed AI solutions to have high fidelity in nuanced settings.

Lalam: This paper offers a clear path forward for developing AI systems that are not just technically capable but are also deeply aligned with the specific cultural and regulatory landscape of their deployment area.

Conclusion: Tom: So, we've looked at all those intricate details on CHILLGuard, and now we need to wrap up by talking about what this paper actually means for us as a whole team.

Jane: We’ve covered the technical deep dives into the data pipeline and the MDPO training, so let’s bring it back to the title and who wrote this work.

Lu: The authors really did a great job framing their contribution by focusing on fine-grained control for Chinese LLMs specifically, which is a really sharp observation.

Meng: I see how they built this system from the ground up to handle those specific cultural and regulatory nuances, but how does that actually translate into something useful in a real-world deployment scenario?

Lalam: For me, the most impactful vision is seeing how this approach can improve culture because it allows us to build AI systems that respect local values instead of just following generic rules.

Tom: That’s a huge point, Lalam; moving beyond surface-level safety to something deeply contextual is exactly what we need.

Jane: And the authors' focus on the specific Chinese context in their taxonomy shows they really understood the unique challenges there, which is why this work feels so relevant right now.

Meng: It’s exciting to see how they tackled the data scarcity problem with that multistage construction pipeline; that sounds like a practical solution for any engineering team facing similar data hurdles.

Lu: I think their iterative training framework, where the generator learns from its own quality scores, is particularly fascinating because it shows a really sophisticated way to refine adversarial samples over time.

Lalam: That refinement process suggests that future AI can be shaped not just by what we tell it to do, but by how we let it learn from challenging interactions within its environment.

Tom: So, the core idea is this specialized methodology creates a high-fidelity safety net for Chinese LLMs through structured data and smarter training methods.

Jane: And the implication is that this level of detail can lead to AI systems that are much safer and more culturally appropriate in complex local settings.

Meng: If we can replicate their success with the data pipeline, it means our teams could focus less on manual rule writing and more on designing these sophisticated generation processes for any region.

Lu: That would open up incredible possibilities for localized AI applications, allowing us to tailor safety mechanisms precisely where they are needed most.

Lalam: I see this as a step toward building truly localized and trustworthy AI systems where the safety mechanisms aren't just imported but are intrinsically shaped by the local environment.

Tom: It really shows that this paper is about taking complex content moderation and giving it a structured, high-quality dataset to make it work effectively.

Jane: And we have to keep an eye on their limitations regarding complete coverage; even with this progress, there's still room for improvement in capturing every possible harmful query.

Lu: That’s the natural next step for any research: acknowledging the current boundaries so we know exactly where the future work needs to go.

Tsinghua University

cs.CL, cs.AI

Submitted: 2026-06-13

Updated: 2026-10-01

Comments: accepted by EMNLP 2026 findings

Code: https://github.com/cswbyu/CHILLGuard

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 92/100

The gist: Malicious content generated by large language models (LLMs) poses severe safety risks, and existing guardrails often fail to adapt to Chinese cultural context and regulatory nuances.

Key concepts

Fine-Grained Risk Taxonomy
This is a detailed classification system with 5 major categories and 31 subcategories specifically designed for Chinese content risks. It moves beyond generic safety rules by covering issues ranging from national security violations to infringements of individual rights, ensuring tailored moderation.
Scalable Multistage Data Construction
To create high-quality training data, the system uses three steps: using RAG for broad text retrieval, collecting real user prompts from production environments, and applying prompt engineering techniques like cultural allusion to generate diverse, hard-to-classify examples.
Model-aware Direct Preference Optimization (MDPO)
MDPO is a training method that dynamically adjusts the penalty for preference optimization based on how quickly the model responds to different training instances. This makes the guardrail classifier more robust by focusing optimization efforts where the model shows real responsiveness to safety signals.

Terminology

Summary

Malicious content generated by large language models (LLMs) poses severe safety risks, and existing guardrails often fail to adapt to Chinese cultural context and regulatory nuances. This paper introduces CHILLGuard, a dedicated Chinese LLM content safety guardrail system that achieves state-of-the-art performance by integrating a scalable data construction pipeline with Model-aware Direct Preference Optimization (MDPO).

Fine-Grained Risk Taxonomy

The foundation of CHILLGuard is a novel 5-macro, 31-micro category fine-grained harm taxonomy specifically aligned with Chinese regulations and linguistic characteristics. This taxonomy addresses the limitations of existing English or generic multilingual taxonomies by covering risks from national security to individual rights. The five major macro-categories are: (A) Violations of Core Socialist Values, (B) Discriminatory Content, (C) Commercial Violations and Non-compliance, (D) Infringement of Legitimate Rights and Interests, and (E) Failure to Meet Safety Demands of Specific Services. This detailed structure ensures comprehensive coverage tailored for Chinese content moderation needs.

Scalable Multistage Data Construction

To overcome the scarcity of high-quality annotated Chinese safety data, the authors propose a scalable multistage data construction pipeline designed to expand corpus diversity and refine quality. This pipeline integrates three complementary sources:

  1. Retrieval-Augmented Generation (RAG)-based Prompt Construction: This stage expands scale by using targeted social media crawling and encoding multi-language corpora via bgem3 embeddings. It constructs retrieval queries using macro/micro-category labels and keywords to retrieve Top-100 candidate texts, which are then fed into a prompt template instructing the model to minimally modify the original content while preserving its semantic intent.

  2. Real-World Data Acquisition: This involves collecting 46,742 real user prompts from authoritative institutions' production environments to capture authentic harmful queries and edge cases.

  3. Prompt Engineering (PE)-based Data Augmentation: This stage increases implicitness and diversity by applying category-specific rewriting strategies tailored to Chinese linguistic characteristics, including homophonic substitution, cultural allusion, rhetorical irony, and semantic nesting. This process generates 109,312 rewritten samples that are incorporated into the final dataset.

Model-aware Direct Preference Optimization (MDPO)

CHILLGuard is trained under a generator-classifier collaborative framework inspired by DuoGuard (Deng et al., 2025), utilizing Model-aware Direct Preference Optimization (MDPO) to enhance detection robustness. The framework consists of two interdependent components: a rewritten generator designed to expand training data diversity and produce challenging, hard-to-classify adversarial samples, and a guardrail classifier optimized to maximize the separability between safe and unsafe prompts. MDPO addresses the limitation of standard Direct Preference Optimization (DPO) by dynamically adjusting the KL penalty based on the model’s real-time responsiveness to specific training instances. This is quantified by computing the implicit reward gap Ri, which is then normalized to estimate a responsiveness factor αM, leading to a dynamic KL penalty coefficient βM = β·αM.

Iterative Generator-Classifier Collaborative Training

The training process follows an iterative three-iteration framework:

  1. Iteration 0: The classifier C(0) is obtained by performing Supervised Fine-Tuning (SFT) directly on the seed training dataset D(0).

  2. Iteration 1: The original generator G(0) performs Prompt Engineering (PE) rewriting on seed samples to generate an augmented dataset D(1)gen, which is merged with D(0)train to form D(1)train, and C(1) is trained via SFT.

  3. Iteration 2: The classifier C(1) labels the generated prompts from D(1)gen, and G(0) assigns a quality score si to each prompt. Samples are mapped into four difficulty levels (L1-L4), and preference pairs P(2) are constructed following priority rules (e.g., ⟨L1, L4⟩ ≻ ⟨L1, L3⟩). The generator G(0) is then finetuned using P(2) via MDPO to obtain the optimized generator G(1). Finally, the classifier C(2) is obtained by SFT on Qwen3 backbone using D(2)train.

Experimental Results and Performance

Extensive experiments demonstrate that CHILLGuard achieves state-of-the-art (SOTA) performance across all three parameter scales. For instance, the 8B variant of CHILLGuard reaches an overall F1 score of 89.77 on the CHILLGuardTest, surpassing Qwen3Guard-8B-Strict by a significant margin of 15.92%.

Improvements for AI systems

Based on the provided scientific paper, here are specific improvements that can be made to existing LLM safety systems, along with what these improved AI systems could achieve:


Figure 1: The overall improvement of CHILLGuard over baseline models.

  1. Improved F1 Score Performance in Chinese Scenarios:

CHILLGuard demonstrates a significant state-of-the-art (SOTA) performance, achieving an overall F1 score of 89.77 on CHILLGuardTest, surpassing the second-best model (Qwen3Guard-8B-Strict) by 15.92%. This suggests that integrating a fine-grained Chinese harm taxonomy and a scalable data construction pipeline leads to much higher precision and recall in Chinese content moderation compared to general multilingual or English-centric guardrails.

  1. Fine-Grained Risk Classification:

The system utilizes a comprehensive 5-macro, 31-micro category taxonomy (covering topics from national security to individual rights). This allows for nuanced risk classification rather than simple binary safe/unsafe labels, enabling deployment in high-stakes environments where regulatory compliance requires specific risk identification (e.g., distinguishing between A2: Endangering national security and A7: Disseminating false information).

  1. Robustness Against Implicit and Adversarial Harmful Content:

The core innovation lies in the iterative generator-classifier collaborative training framework powered by Model-aware Direct Preference Optimization (MDPO).

CHILLGuardTrain is generated through three stages: Retrieval-Augmented Generation (RAG) for scale, Prompt Engineering (PE) rewriting to generate implicit/obfuscated samples, and human expert calibration via multi-model voting.

The MDPO mechanism dynamically adjusts the KL penalty based on the model’s mastery of specific difficulty levels (L1–L4), specifically prioritizing challenging preference pairs. This directly addresses the critical gap in existing models that fail to detect harmful content expressed through homophonic substitutions, cultural allusions, or rhetorical irony—the very linguistic nuances prevalent in Chinese online environments.

  1. Scalable and Data-Efficient Training:

The proposed pipeline (RAG + PE + Multi-model Voting) creates a large training set of 405,007 samples with a rigorous test set of 51,745 samples. This scalable construction method addresses the scarcity of high-quality annotated Chinese safety data by synthesizing diverse real-world and adversarial examples.

What the Improved AI System Can Do:

The resulting CHILLGuard system can perform highly reliable, context-aware content safety moderation for Chinese LLM deployments, specifically excelling in complex regulatory and cultural contexts. It can achieve the following specific capabilities:

  1. Precise Regulatory Compliance Monitoring: Unlike general guardrails that might flag benign discussions as harmful (high false positives), CHILLGuard can accurately classify content according to China's specific legal frameworks across 31 detailed risk categories, ensuring high precision for legal and governmental applications.

  2. Detection of Implicit Evasion Tactics: The system is specifically trained to recognize sophisticated evasion techniques prevalent in Chinese internet culture, such as homophonic substitution (using similar-sounding characters), cultural allusions, and rhetorical irony, making it superior at detecting borderline harmful content that simpler models miss.

  3. Adversarial Robustness: By incorporating an iterative generator-classifier loop under MDPO, the system gains robustness against adaptive adversarial attacks. It can be trained to identify novel, unseen harmful prompts generated by malicious actors, significantly reducing the vulnerability of LLM applications in real-world deployment scenarios.

  4. High-Quality User Experience Filtering: The fine-grained classification (including Macro E: Failure to Meet Safety Demands of Specific Services) allows developers to tune safety guardrails not just for illegal content, but also for quality issues like inaccurate scientific claims or unhelpful responses, leading to a more reliable and user-friendly LLM service.

  5. Efficient Deployment: The performance results across various model sizes (1.7B to 8B+) indicate that CHILLGuard can be deployed efficiently on resource-constrained hardware while maintaining superior safety performance compared to larger, less specialized models like Qwen3Guard.

Abstract

Malicious content generated from large language models (LLMs) could pose severe safety risks and ethical concerns. While existing LLM safety guardrails excel in English or multilingual settings, they lack adaptation to Chinese-specific regulatory policies, cultural context, and linguistic nuances, failing to support fine-grained risk classification for diverse deployment needs. In this paper, we introduce a 5-macro, 31-micro category fine-grained risk taxonomy for Chinese scenarios, and build CHILLGuard: a dedicated Chinese LLM content safety guardrail. To address the critical scarcity of high-quality annotated Chinese safety data, we propose a scalable multi-stage data construction pipeline: we expand multi-source corpus via retrieval-augmented generation, generate implicit harmful samples through prompt engineering rewriting, and refine high-quality data via multi-model voting-based label calibration. Based on this, we build CHILLGuardTrain, a large-scale training set with 405,007 samples, and CHILLGuardTest, a rigorously curated annotated test set with 51,745 samples. We then train CHILLGuard on CHILLGuardTrain under a generator-classifier collaborative framework via Model-aware Direct Preference Optimization. Extensive experiments under multiple settings demonstrate the state-of-the-art performance of CHILLGuard, e.g., a 15.92% relative improvement of F1 score over Qwen3Guard-8B-Strict on our benchmark. We release our resources at https://github.com/cswbyu/CHILLGuard.

Sources

Related papers