IDEEA: training-free Input-Dependent stEEring via Activation cluster matching
cs.CL, cs.LG
Submitted: 2026-09-02
Updated: 2026-09-02
Comments: Accepted to EMNLP 2026 Findings
Code: https://github.com/DSL-Lab/IDEEA
License: http://creativecommons.org/licenses/by/4.0/
The gist: Steering aligns large language models (LLMs) by injecting a bias into selected activations at inference time, offering a far cheaper alternative to weight-update methods such as supervised
Terminology
Abstract
Steering aligns large language models (LLMs) by injecting a bias into selected activations at inference time, offering a far cheaper alternative to weight-update methods such as supervised fine-tuning or reinforcement learning. However, most existing training-free steering methods are input-independent: a single direction is fitted once and shared across all inputs. This is fundamentally limiting as different inputs occupy different regions of the activation space and admit different optimal steering directions toward the same target concept, much as the gradient with respect to a fixed loss varies from input to input. We close this gap with IDEEA (Input-Dependent stEEring via Activation cluster matching), a training-free framework for input-dependent steering. IDEEA clusters the positive and negative activation supports per attention head, and solves an optimal-matching problem to construct a set of cluster-conditional directions, all about the target concept. At inference time, it picks from this pool of directions and uses the one that best matches the input's own activation for steering. IDEEA aligns the model toward the target concept while preserving the input's original representation, evidence that activations encoding a concept occupy several distinct sub-regions of the representation space rather than a single one. IDEEA improves the truth times info rate in TruthfulQA by an average of 9.9% (up to 23.5%) over the best input-independent baseline.
Sources
- Mistral 7B
- Do LLM Agents Exhibit Social Behavior?
- The Llama 3 Herd of Models
- Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
- Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Gemma 2: Improving Open Language Models at a Practical Size
- Scikit-learn: Machine Learning in Python
- Qwen2.5 Technical Report
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- HuggingFace's Transformers: State-of-the-art Natural Language Processing
- Representation Engineering: A Top-Down Approach to AI Transparency
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering