Refusing Intent, Not Form: Wrapper-Based Intent-Group Supervision for LLM Safety
Ping Wu, Haibo Tong, Feifei Zhao, Han Shen, Yu Shi, Yilin Zhao, Sicheng Shen, Guobin Shen, Yun Luo, Yi Zeng
Institute of Automation, Chinese Academy of Sciences · University of Chinese Academy of Sciences · Beijing Key Laboratory of Safe AI and Superalignment · Ant Group
cs.CL
Submitted: 2026-08-13
Updated: 2026-08-14
Comments: 23 pages, 11 figures, 24 tables
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: This paper proposes Wrapper-Based Intent-Form Augmentation (WIFA), an automatic intent-group augmentation method that pairs wrapped harmful examples with structurally matched wrapped benign
Terminology
Summary
This paper proposes Wrapper-Based Intent-Form Augmentation (WIFA), an automatic intent-group augmentation method that pairs wrapped harmful examples with structurally matched wrapped benign counterexamples, requiring no external teacher or manual per-wrapper intent labels. The paper uses WIFA as a common data layer for two complementary fine-tuning routes: WIFA-Boost, a two-stage high-safety recipe, and Anchored Group-Consistent Refusal Training (A-GCRT), which regularizes refusal/compliance decision scores across same-intent wrappers and anchors harmful and benign groups on opposite sides of a margin.
The paper's main contributions are:
-
Automatic matched intent-group augmentation: WIFA groups the same underlying intent across different surface forms and pairs harmful intent groups with structurally matched benign intent groups, making wrapper form a nuisance variable rather than a decision label.
-
WIFA-based training for different safety goals: WIFA-Boost provides a safety-first route for harmful requests written in altered forms, while A-GCRT uses intent-group consistency and directional anchors to reduce unnecessary refusal on benign requests while preserving harmful-request refusal.
-
Empirical gains beyond base models and reproduced defenses: In the main Qwen setting, WIFA-Boost raises SORRY-Bench mutation-average refusal from 22.1 to 63.7, exceeding the strongest reproduced defense on this metric. A low-over-refusal A-GCRT operating point lowers Qwen OR-Bench over-refusal from 25.7 to 17.4, below both the base model and all reproduced defenses, while still improving harmful-request refusal.
In the Qwen setting, WIFA-Boost reaches the strongest transformed-harmful refusal, while A-GCRT reduces OR-Bench over-refusal from 25.7% for the base model to 17.4%; reproduced baselines do not match these operating points. Llama results and ablations over data structure, two-stage order, and A-GCRT components support this intent-group interpretation without claiming universal below-base over-refusal.
The paper evaluates the methods on two model settings (Qwen2.5-7B-Instruct and Llama-3.1-8B-Instruct), seven benchmarks, and a 15-family unseen-attack suite, compares against six reproduced defenses, and ablates data structure, data ratio, stage order, A-GCRT components, margin/anchor settings, and decision-score diagnostics.
The paper concludes that safety supervision should target intent groups rather than isolated prompt forms. WIFA builds the matched harmful/benign data layer, WIFA-Boost turns it into a high-safety training route, and A-GCRT uses anchored group consistency to shape the safety–over-refusal trade-off. The central lesson is refusing harmful intent, not surface form.
Improvements for AI systems
Improvements to AI systems:
-
Intent-group-based safety training: Replace single-prompt safety fine-tuning with wrapper-based intent-group augmentation, where harmful and benign examples are paired across surface forms (e.g., paraphrases, encodings, role-play contexts) to train the model to recognize and refuse the underlying intent rather than specific phrasings. This reduces evasion via novel wrappers.
-
Two-stage safety boosting (WIFA-Boost): Implement a two-stage fine-tuning recipe: first, expose the model to matched harmful/benign intent groups to learn form-invariant refusal; second, fine-tune on standard safety data to consolidate. This raises refusal rates on transformed harmful requests (e.g., from 22.1% to 63.7% on SORRY-Bench) without needing external teachers or manual wrapper labels.
-
Anchored group-consistent refusal training (A-GCRT): Add a regularization term that penalizes inconsistent refusal/compliance scores across same-intent wrappers and anchors harmful groups above a high refusal threshold and benign groups below a low threshold. This reduces over-refusal on benign inputs (e.g., OR-Bench from 25.7% to 17.4%) while preserving harmful-request refusal, improving the safety–utility trade-off.
-
Automatic matched data generation: Use the WIFA data layer to automatically generate paired harmful/benign intent groups with structurally matched wrappers (e.g., same syntactic template, different intent). This eliminates the need for manual per-wrapper intent labels, enabling scalable safety training across diverse attack surfaces.
-
Diagnostic score alignment: Train the model to produce decision scores that are consistent within an intent group and separated by a margin between harmful and benign groups. This makes the model’s refusal behavior more predictable and interpretable, allowing operators to tune the margin for desired safety vs. over-refusal operating points.
What the improved AI system can do:
-
Refuse harmful requests even when rephrased, encoded, or embedded in novel contexts (e.g., jailbreak wrappers, fictional scenarios, indirect prompts), while complying with benign requests that share similar surface forms.
-
Automatically generate its own matched harmful/benign training pairs for new domains or languages, without human labeling of wrapper types.
-
Provide a tunable safety–over-refusal knob (via margin/anchor settings) so deployers can choose between strict safety (high refusal on transformed attacks) or low over-refusal (fewer false positives on benign queries) based on application risk.
-
Generalize to unseen attack families (e.g., 15-family suite) better than models trained on fixed prompt lists, because it learns intent-level boundaries rather than form-level patterns.
-
Maintain high performance on standard safety benchmarks while significantly reducing unnecessary refusals on legitimate user requests, improving user experience and trust.
Sources
- Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models
- Invariant Risk Minimization
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Constitutional AI: Harmlessness from AI Feedback
- OR-Bench: An Over-Refusal Benchmark for Large Language Models
- Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
- The Llama 3 Herd of Models
- Measuring Massive Multitask Language Understanding
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Distributionally Robust Neural Networks for Group Shifts: On the Importance of Regularization for Worst-Case Generalization
- Safety Instincts: LLMs Learn to Trust Their Internal Compass for Self-Defense
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- FalseReject: A Resource for Improving Contextual Safety and Mitigating Over-Refusals in LLMs via Structured Reasoning
- The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
- Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- Qwen2.5 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering