Refusing Intent, Not Form: Wrapper-Based Intent-Group Supervision for LLM Safety

arXiv:2608.13304 · cs.CL · Submitted 2026-08-13 · Read on arXiv

Ping Wu, Haibo Tong, Feifei Zhao, Han Shen, Yu Shi, Yilin Zhao, Sicheng Shen, Guobin Shen, Yun Luo, Yi Zeng

Institute of Automation, Chinese Academy of Sciences · University of Chinese Academy of Sciences · Beijing Key Laboratory of Safe AI and Superalignment · Ant Group

cs.CL

Submitted: 2026-08-13

Updated: 2026-08-14

Comments: 23 pages, 11 figures, 24 tables

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: This paper proposes Wrapper-Based Intent-Form Augmentation (WIFA), an automatic intent-group augmentation method that pairs wrapped harmful examples with structurally matched wrapped benign

Terminology

Summary

This paper proposes Wrapper-Based Intent-Form Augmentation (WIFA), an automatic intent-group augmentation method that pairs wrapped harmful examples with structurally matched wrapped benign counterexamples, requiring no external teacher or manual per-wrapper intent labels. The paper uses WIFA as a common data layer for two complementary fine-tuning routes: WIFA-Boost, a two-stage high-safety recipe, and Anchored Group-Consistent Refusal Training (A-GCRT), which regularizes refusal/compliance decision scores across same-intent wrappers and anchors harmful and benign groups on opposite sides of a margin.

The paper's main contributions are:

  • Automatic matched intent-group augmentation: WIFA groups the same underlying intent across different surface forms and pairs harmful intent groups with structurally matched benign intent groups, making wrapper form a nuisance variable rather than a decision label.

  • WIFA-based training for different safety goals: WIFA-Boost provides a safety-first route for harmful requests written in altered forms, while A-GCRT uses intent-group consistency and directional anchors to reduce unnecessary refusal on benign requests while preserving harmful-request refusal.

  • Empirical gains beyond base models and reproduced defenses: In the main Qwen setting, WIFA-Boost raises SORRY-Bench mutation-average refusal from 22.1 to 63.7, exceeding the strongest reproduced defense on this metric. A low-over-refusal A-GCRT operating point lowers Qwen OR-Bench over-refusal from 25.7 to 17.4, below both the base model and all reproduced defenses, while still improving harmful-request refusal.

In the Qwen setting, WIFA-Boost reaches the strongest transformed-harmful refusal, while A-GCRT reduces OR-Bench over-refusal from 25.7% for the base model to 17.4%; reproduced baselines do not match these operating points. Llama results and ablations over data structure, two-stage order, and A-GCRT components support this intent-group interpretation without claiming universal below-base over-refusal.

The paper evaluates the methods on two model settings (Qwen2.5-7B-Instruct and Llama-3.1-8B-Instruct), seven benchmarks, and a 15-family unseen-attack suite, compares against six reproduced defenses, and ablates data structure, data ratio, stage order, A-GCRT components, margin/anchor settings, and decision-score diagnostics.

The paper concludes that safety supervision should target intent groups rather than isolated prompt forms. WIFA builds the matched harmful/benign data layer, WIFA-Boost turns it into a high-safety training route, and A-GCRT uses anchored group consistency to shape the safety–over-refusal trade-off. The central lesson is refusing harmful intent, not surface form.

Improvements for AI systems

Improvements to AI systems:

  1. Intent-group-based safety training: Replace single-prompt safety fine-tuning with wrapper-based intent-group augmentation, where harmful and benign examples are paired across surface forms (e.g., paraphrases, encodings, role-play contexts) to train the model to recognize and refuse the underlying intent rather than specific phrasings. This reduces evasion via novel wrappers.

  2. Two-stage safety boosting (WIFA-Boost): Implement a two-stage fine-tuning recipe: first, expose the model to matched harmful/benign intent groups to learn form-invariant refusal; second, fine-tune on standard safety data to consolidate. This raises refusal rates on transformed harmful requests (e.g., from 22.1% to 63.7% on SORRY-Bench) without needing external teachers or manual wrapper labels.

  3. Anchored group-consistent refusal training (A-GCRT): Add a regularization term that penalizes inconsistent refusal/compliance scores across same-intent wrappers and anchors harmful groups above a high refusal threshold and benign groups below a low threshold. This reduces over-refusal on benign inputs (e.g., OR-Bench from 25.7% to 17.4%) while preserving harmful-request refusal, improving the safety–utility trade-off.

  4. Automatic matched data generation: Use the WIFA data layer to automatically generate paired harmful/benign intent groups with structurally matched wrappers (e.g., same syntactic template, different intent). This eliminates the need for manual per-wrapper intent labels, enabling scalable safety training across diverse attack surfaces.

  5. Diagnostic score alignment: Train the model to produce decision scores that are consistent within an intent group and separated by a margin between harmful and benign groups. This makes the model’s refusal behavior more predictable and interpretable, allowing operators to tune the margin for desired safety vs. over-refusal operating points.

What the improved AI system can do:

  • Refuse harmful requests even when rephrased, encoded, or embedded in novel contexts (e.g., jailbreak wrappers, fictional scenarios, indirect prompts), while complying with benign requests that share similar surface forms.

  • Automatically generate its own matched harmful/benign training pairs for new domains or languages, without human labeling of wrapper types.

  • Provide a tunable safety–over-refusal knob (via margin/anchor settings) so deployers can choose between strict safety (high refusal on transformed attacks) or low over-refusal (fewer false positives on benign queries) based on application risk.

  • Generalize to unseen attack families (e.g., 15-family suite) better than models trained on fixed prompt lists, because it learns intent-level boundaries rather than form-level patterns.

  • Maintain high performance on standard safety benchmarks while significantly reducing unnecessary refusals on legitimate user requests, improving user experience and trust.

Sources

Related papers