AUDITPLAN: Commit, Then Answer for Auditable Safety Alignment
cs.CR, cs.AI, cs.CL
Submitted: 2026-09-16
Updated: 2026-09-16
License: http://creativecommons.org/licenses/by/4.0/
The gist: Safety tuning pipelines judge only the final answer, which makes it difficult to distinguish robust refusal from two undesirable shortcuts: blanket refusal on benign requests and polished but
Terminology
Abstract
Safety tuning pipelines judge only the final answer, which makes it difficult to distinguish robust refusal from two undesirable shortcuts: blanket refusal on benign requests and polished but unfaithful safety rationales that do not actually constrain the answer. We propose AUDITPLAN, a single-model plan-then-answer approach where the model first emits a compact structured safety plan and then answers conditioned on it. The plan records a threat label, intended action, and explicit constraints, enabling machine-checkable auditing while remaining hidden from users at deployment. We train this behavior with supervised fine-tuning followed by reinforcement learning with FAITHGATE, a reward-gating objective that grants answer reward only when the safety plan is correct. This discourages safe-looking but unfaithful behavior and promotes tighter plan-answer coupling. Across Qwen backbones, AUDITPLAN improves both robustness and auditability: on Qwen2.5-3B-Instruct, FAITHGATE reduces ASR from 24.0% to 11.6%, LSR from 1.0% to 0.36%, and over-refusal from 11.0% to 2.0%, outperforming answer-only RL, free-form explanation, and weighted-sum structured rewards. Similar trends hold for Qwen2.5-1.5B-Instruct. Larger-model confirmation runs on Qwen-3-4B-Instruct and Qwen2.5-7B-Instruct preserve the same trend suggesting that explicit internal commitments can make safety alignment more faithful, robust, and auditable.
Sources
- Constitutional AI: Harmlessness from AI Feedback
- JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
- QLoRA: Efficient Finetuning of Quantized LLMs
- Enhancing Chat Language Models by Scaling High-quality Instructional Conversations
- WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
- WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
- Measuring Faithfulness in Chain-of-Thought Reasoning
- Let's Verify Step by Step
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- Training language models to follow instructions with human feedback
- Qwen2.5 Technical Report
- Qwen3 Technical Report
- Deliberative Alignment: Reasoning Enables Safer Language Models
- LoRA: Low-Rank Adaptation of Large Language Models
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- PLeak: Prompt Leaking Attacks against Large Language Model Applications
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs
- Benchmarking and Defending Against Indirect Prompt Injection Attacks on Large Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs