Steering Vector Fields for Context-Aware Inference-Time Control in Large Language Models
cs.CL
Submitted: 2026-02-02
Updated: 2026-09-19
Comments: EMNLP 2026 Main
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Steering vectors (SVs) offer a lightweight way to control large language models (LLMs) at inference time by shifting hidden activations, providing a practical middle ground between prompting and
Terminology
Abstract
Steering vectors (SVs) offer a lightweight way to control large language models (LLMs) at inference time by shifting hidden activations, providing a practical middle ground between prompting and fine-tuning. Yet SVs can be unreliable in practice. Some concepts are unsteerable, and even when steering helps on average it can backfire for a non-trivial fraction of inputs. Reliability also degrades in long-form generation and multi-attribute steering. We take a geometric view of these failures. A static SV applies the same update vector everywhere in representation space, implicitly assuming that the concept-improving direction is constant across contexts. When the locally effective direction varies with the current activation, a single global vector can become misaligned, which yields weak or reversed effects. Guided by this perspective, we propose Steering Vector Fields (SVF), which learns a differentiable concept scoring function whose local gradient defines the steering direction at each activation, making interventions explicitly context-dependent. This formulation supports coordinated multi-layer interventions in a shared, aligned concept space, and enables efficient long-form and multi-attribute control within a unified framework. Across multiple LLMs and steering tasks, SVF delivers stronger and more reliable control, improving the practicality of inference-time steering.
Sources
- Designing a Dashboard for Transparency and Control of Conversational AI
- CausalGym: Benchmarking causal interpretability methods on linguistic tasks
- LEACE: Perfect linear concept erasure in closed form
- Understanding (Un)Reliability of Steering Vectors in Language Models
- Measuring Massive Multitask Language Understanding
- Self-Control of LLM Behaviors by Compressing Suffix Gradient into Prefix Controller
- Implicit In-context Learning
- TruthfulQA: Measuring How Models Mimic Human Falsehoods
- Discovering Language Model Behaviors with Model-Written Evaluations
- In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space Steering
- Can sparse autoencoders be used to decompose and interpret steering vectors?
- Steering Llama 2 via Contrastive Activation Addition
- Analyzing the Generalization and Reliability of Steering Vectors
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Qwen3 Technical Report
- Steering Language Models With Activation Engineering
- Root Mean Square Layer Normalization
- Representation Engineering: A Top-Down Approach to AI Transparency
- Extending Activation Steering to Broad Skills and Multiple Behaviours
- Taxonomy, Opportunities, and Challenges of Representation Engineering for Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering