When Does Activation Steering Change What a Model Computes From?
cs.LG, cs.CL
Submitted: 2026-06-28
Updated: 2026-09-06
License: http://creativecommons.org/licenses/by/4.0/
The gist: Activation steering can reliably change an agent's output by modifying its internal activations.
Terminology
Abstract
Activation steering can reliably change an agent's output by modifying its internal activations. Yet arriving at the same answer need not involve the same computation: behavioral equivalence does not imply mechanistic equivalence. We test whether an activation edit changes the state used in subsequent computation or biases that computation toward the desired output. In a controlled state-tracking task, trace supervision creates an editable register: changing its internal value makes the model apply the next operation to the edited state. We next ask whether widely used mean activation steering admits a similar interpretation: does steering recreate the internal configuration the model naturally uses when performing a task? If so, transplanting that activation should be effective, and steering should remain effective at the scale of the natural source-to-target activation change. Neither prediction holds in the two selected model-task settings. Replacing the activation at one layer produces less than 2% of the target-answer margin gain from patching through all remaining layers, while natural-scale steering is similarly ineffective. The steering vectors have norms 77 and 26 times the median natural change in Qwen and Llama, producing 54% and 82% of the reference effect. Thus, a successful steering intervention need not reproduce the natural target activation at the intervention layer. More generally, an intervention should be interpreted as changing the computational state only when a later computation uses the edited value according to the semantics of that state.
Sources
- Not All Language Model Features Are One-Dimensionally Linear
- Arithmetic in the Wild: Llama uses Base-10 Addition to Reason About Cyclic Concepts
- Language Models Use Trigonometry to Do Addition
- Measuring Faithfulness in Chain-of-Thought Reasoning
- Show Your Work: Scratchpads for Intermediate Computation with Language Models
- In-context Learning and Induction Heads
- Attribution Patching Outperforms Automated Circuit Discovery
- Steering Language Models With Activation Engineering
- Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior
- Representation Engineering: A Top-Down Approach to AI Transparency
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks