SHARD: Safe and Helpful Alignment via Self-Reframing Distillation
cs.CL
Submitted: 2026-06-14
Updated: 2026-09-02
Comments: EMNLP 2026
Code: https://github.com/Viswonathan06/shard-self-reframing
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models often struggle with sensitive prompts.
Terminology
Abstract
Large language models often struggle with sensitive prompts. They may refuse outright, provide generic safety boilerplate, or fail to address the user's legitimate informational needs that can be answered safely. We introduce SHARD, a self-reframing distillation method to improve safe-helpfulness. It first rewrites sensitive prompts to surface benign intent using philosophical guidelines, then reframes its original responses into safe, more helpful ones, and finally fine-tunes the model on its self-reframed responses. Across DNA and the English subset of LINGUASAFE, SHARD improves helpfulness for most model families while preserving safety. It also remains competitive with distillation from a larger teacher model, suggesting that models can internalize safe and helpful behavior elicited from their own. Warning: This paper contains content that may be offensive or harmful.
Sources
- Phi-4 Technical Report
- A General Language Assistant as a Laboratory for Alignment
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Constitutional AI: Harmlessness from AI Feedback
- The Llama 3 Herd of Models
- Deliberative Alignment: Reasoning Enables Safer Language Models
- THINKSAFE: Self-Generated Safety Alignment for Reasoning Models
- IntentionReasoner: Facilitating Adaptive LLM Safeguards through Intent Reasoning and Selective Query Refinement
- Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior
- LLaMA: Open and Efficient Foundation Language Models
- HuggingFace's Transformers: State-of-the-art Natural Language Processing
- FalseReject: A Resource for Improving Contextual Safety and Mitigating Over-Refusals in LLMs via Structured Reasoning
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering