Controlling Refusal Behavior of LLMs via Stiefel-Constrained Rotation Steering
cs.LG, cs.CL
Submitted: 2026-08-31
Updated: 2026-08-31
Comments: 19 pages, 5 figures, 6 tables
License: http://creativecommons.org/licenses/by/4.0/
The gist: Activation steering has emerged as a lightweight approach for controlling model refusal at inference time.
Terminology
Abstract
Activation steering has emerged as a lightweight approach for controlling model refusal at inference time. A growing line of research explores trainable rotations of activations to develop geometrically principled intervention mechanisms. However, existing techniques rely on auxiliary constructs, such as refusal vectors, to define these rotations. In our work, we develop a self-contained methodology for learning parameter-efficient rotational transformations based on Riemannian optimization. We empirically validate the proposed scheme, demonstrating its superiority in intervention efficiency. An extensive ablation study highlights the importance of key design choices in our method. Our results identify the proposed rotation-based steering scheme as a promising direction for more reliable control over the behavior of LLMs.
Sources
- DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Training Verifiers to Solve Math Word Problems
- Selective Steering: Norm-Preserving Control Through Discriminative Layer Selection
- Measuring Massive Multitask Language Understanding
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Refusal in LLMs is an Affine Function
- Pointer Sentinel Mixture Models
- Minimizing Collateral Damage in Activation Steering
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
- AlphaSteer: Learning Refusal Steering with Principled Null-Space Constraint
- Angular Steering: Behavior Control via Rotation in Activation Space
- The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence
- Qwen2.5 Technical Report
- Spherical Steering: Geometry-Aware Activation Rotation for Language Models
- Qwen3Guard Technical Report
- LLMs Encode Harmfulness and Refusal Separately
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks