There Is More to Refusal in Large Language Models than a Single Direction
cs.CL
Submitted: 2026-02-02
Updated: 2026-09-15
Comments: 37 pages. Accepted for publication in the main track of EMNLP 2026. Updated manuscript
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- A General Language Assistant as a Laboratory for Alignment
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Constitutional AI: Harmlessness from AI Feedback
- Safe RLHF: Safe Reinforcement Learning from Human Feedback
- Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models
- Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
- Aligning AI With Shared Human Values
- Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal
- Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
- Steering Language Model Refusal with Sparse Autoencoders
- Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders
- The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence
- From Instructions to Intrinsic Human Values -- A Survey of Alignment Goals for Big Models
- Understanding Refusal in Language Models with Sparse Autoencoders
- Differentiated Directional Intervention A Framework for Evading LLM Safety Alignment
- Representation Engineering: A Top-Down Approach to AI Transparency
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering