Steering Interference Reflects the Model's Defaults, Not the Behavior Directions
cs.LG, cs.AI
Submitted: 2026-09-07
Updated: 2026-09-25
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Activation steering promises modular control of language model behavior: a behavior such as politeness corresponds to a direction in a model's activations, and adding that direction while it
Terminology
Abstract
Activation steering promises modular control of language model behavior: a behavior such as politeness corresponds to a direction in a model's activations, and adding that direction while it generates should switch the behavior on and leave everything else alone. It does not. We ask what decides which other behaviors move, and by how much, and find that it is the model rather than the behavior being steered. A steer relaxes the model toward a small set of behaviors it already favors, chiefly refusal, sycophancy, and poeticism, and that set is much the same whatever is steered. Three results across 24 behaviors and ten instruction-tuned models support this, every effect read off the generated text by a language-model judge rather than off a probe. That readout matters: all 24 behaviors are linearly decodable, but only 20 change what the model writes. First, a direction carrying no behavioral content, matched to a real steer only in the size of the vector it adds, moves the same behaviors in the same order as real steers do, while producing none of the behaviors that need a specific direction. Second, most interference runs one way, so it cannot be an overlap between two directions: steering profanity makes the model toxic, while steering toxicity leaves profanity untouched. Third, with a behavior held out entirely, geometry measured on the others explains almost none of the interference it takes part in. The account holds on all ten models, the pull toward defaults strongest below 10B parameters and weakening in each family's largest. Reading a steer as a perturbation whose endpoint the model fixes implies that disentangling behavior directions cannot by itself make steering modular.
Sources
- Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
- Do Personality Traits Interfere? Geometric Limitations of Steering in Large Language Models
- Understanding (Un)Reliability of Steering Vectors in Language Models
- Persona Vectors: Monitoring and Controlling Character Traits in Language Models
- Perfect Detection, Failed Control: The Geometry of Knowing vs. Steering in Language Models
- There Is More to Refusal in Large Language Models than a Single Direction
- The Rogue Scalpel: Activation Steering Compromises LLM Safety
- Analysing the Safety Pitfalls of Steering Vectors
- The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models
- Multi-Attribute Steering of Language Models via Targeted Intervention
- Beyond Linear Steering: Unified Multi-Attribute Control for Language Models
- A Low-Rank Subspace Analysis of LLM Interventions
- SteeringSafety: Benchmarking Representation Steering in LLMs Across Safety Perspectives
- Steering Language Models With Activation Engineering
- In-Distribution Steering: Balancing Control and Coherence in Language Model Generation
- The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence
- Representation Engineering: A Top-Down Approach to AI Transparency
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks