Beyond Shallow Alignment: How Post-Training Methods Determine Refusal Circuits And Steering Robustness
cs.CL
Submitted: 2026-09-03
Updated: 2026-09-03
Comments: 27 pages, accepted at EMNLP 2026 Main Conference
Code: https://github.com/hoangcuongnguyen2001/Beyond-Shallow-Alignment
License: http://creativecommons.org/licenses/by/4.0/
The gist: How do the methods used to train language models to refuse harmful requests shape how that refusal actually works inside the model? We compare three post-training methods - supervised fine-tuning,
Terminology
Abstract
How do the methods used to train language models to refuse harmful requests shape how that refusal actually works inside the model? We compare three post-training methods - supervised fine-tuning, reasoning-augmented fine-tuning (training on reasoning chains that justify a safety decision), and preference optimization (ORPO) - across three architecturally distinct models (Llama-3.1-8B, Gemma-2-9B, Qwen3-8B). We find that training method, not just data, reshapes how refusal is computed internally: reasoning-augmented training consistently produces a distinct kind of refusal computation, visible across all three models, while architecture independently shapes internal structure and how reliably refusal can be steered. Most importantly, no method we study achieves all three properties we would want from safe alignment at once: refusal that isn't concentrated in a few fragile components, safety gains that don't cost general capability, and safety behavior correctable through small, targeted edits. We caution against treating current post-training methods as a solved, reliable defense, especially for security-critical use. Code and models are available in https://github.com/hoangcuongnguyen2001/Beyond-Shallow-Alignment.
Sources
- Where to Steer: Input-Dependent Layer Selection for Steering Improves LLM Alignment
- Gemma 2: Improving Open Language Models at a Practical Size
- The Llama 3 Herd of Models
- A Single Neuron Is Sufficient to Bypass Safety Alignment in Large Language Models
- The Hydra Effect: Emergent Self-repair in Language Model Computations
- Mechanistic Interpretability for Large Language Model Alignment: Progress, Challenges, and Future Directions
- Thinking Sparks!: Emergent Attention Heads in Reasoning Models During Post Training
- Qwen3 Technical Report
- OpenAI GPT-5 System Card
- Steering Language Models With Activation Engineering
- Objective Matters: Fine-Tuning Objectives Shape Safety, Robustness, and Persona Drift
- Knowing without Acting: The Disentangled Geometry of Safety Mechanisms in Large Language Models
- OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models
- Representation Engineering: A Top-Down Approach to AI Transparency
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering