Behind Harmful Compliance: Behavioral and Mechanistic Divergence Across LLM Jailbreaks
cs.CR, cs.AI, cs.CL
Submitted: 2026-04-20
Updated: 2026-09-05
Code: https://github.com/tatsu-lab/stanford_alpaca
License: http://creativecommons.org/licenses/by/4.0/
The gist: Open-weight language models can be rendered unsafe through several parameter-level interventions, yet models with matched harmful compliance can exhibit fundamentally different failure modes.
Terminology
Abstract
Open-weight language models can be rendered unsafe through several parameter-level interventions, yet models with matched harmful compliance can exhibit fundamentally different failure modes. We compare harmful supervised fine-tuning (SFT), harmful reinforcement learning with verifiable rewards (RLVR), and refusal-feature abliteration in Qwen2.5-7B and Llama-3.1-8B using harmfulness, capability, self-audit, safety reflection, representation similarity, and refusal-direction repair. All three routes reach near-ceiling harmfulness, but SFT causes the broadest capability loss and representational drift; abliteration yields localized, family-dependent refusal-feature suppression; and RLVR largely preserves base-model capability, explicit safety judgments, and representation geometry while retargeting behavior toward compliance. RLVR models consequently remain unusually responsive to safety-reflection prompts despite complying under direct prompting. We further find that harmful RLVR induces capability-blind compliance: models claim to complete unavailable actions and fabricate information about nonexistent entities. Targeted RLVR calibration substantially reduces this false acceptance without restoring safety or degrading general capability. These results show that harmful compliance, harm recognition, and capability awareness are separable behavioral axes, and that self-audit and hallucination patterns are not necessarily robust safety signals under adaptive post-training.
Sources
- Constitutional AI: Harmlessness from AI Feedback
- Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation
- There Is More to Refusal in Large Language Models than a Single Direction
- No Two Devils Alike: Unveiling Distinct Mechanisms of Fine-tuning Attacks
- Targeted Vaccine: Safety Alignment for Large Language Models against Harmful Fine-Tuning via Layer-wise Perturbation
- HarmRLVR: Weaponizing Verifiable Rewards for Harmful LLM Alignment
- Mitigating the Safety Alignment Tax with Null-Space Constrained Policy Optimization
- Training language models to follow instructions with human feedback
- GRP-Obliteration: Unaligning LLMs With a Single Unlabeled Prompt
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior
- A StrongREJECT for Empty Jailbreaks
- Safety Alignment as Continual Learning: Mitigating the Alignment Tax via Orthogonal Gradient Projection
- Persistent Instability in LLM's Personality Measurements: Effects of Scale, Reasoning, and Conversation History
- Jailbroken: How Does LLM Safety Training Fail?
- Knowing without Acting: The Disentangled Geometry of Safety Mechanisms in Large Language Models
- AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies
- LLMs Encode Harmfulness and Refusal Separately
- Universal and Transferable Adversarial Attacks on Aligned Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs