SafeSteer: Localized On-Policy Distillation for Efficient Safety Alignment
Hao Li, Jingkun An, Zijun Song, Pengyu Zhu, Rui Li, Hao Wang, Wendi Feng, Yesheng Liu, Lijun Li, Jin-Ge Yao, Lei Sha
cs.AI, cs.CL
Submitted: 2026-08-21
Updated: 2026-08-24
Code: https://github.com/tatsu-lab/alpaca_eval
Project page: https://anjingkun.github.io/SafeSteer
Terminology
Sources
- Evaluating Large Language Models Trained on Code
- Training Verifiers to Solve Math Word Problems
- Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment
- Measuring Massive Multitask Language Understanding
- Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable
- Mitigating the Safety Alignment Tax with Null-Space Constrained Policy Optimization
- Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
- Let's Verify Step by Step
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- The Llama 3 Herd of Models
- X-Boundary: Establishing Exact Safety Boundary to Shield LLMs from Multi-Turn Jailbreaks without Compromising Usability
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- Qwen2.5 Technical Report
- LLMs know their vulnerabilities: Uncover Safety Gaps through Natural Distribution Shifts
- Self-Distillation Enables Continual Learning
- A Survey of On-Policy Distillation for Large Language Models
- ALERT: A Comprehensive Benchmark for Assessing Large Language Models' Safety through Red Teaming
- SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
- Qwen3 Technical Report
- GLM-5: from Vibe Coding to Agentic Engineering
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection