MetaSteer: Context-Conditioned, nonlinear Steering via Attention-Projection Adaptation
cs.CL
Submitted: 2026-09-30
Updated: 2026-09-30
Terminology
Sources
- OPIUM: Mitigating Steering Externalities and Over-Refusal via Dual Objective Latent Optimization
- Mechanistic Interpretability for AI Safety -- A Review
- Improving Steering Vectors by Targeting Sparse Autoencoder Features
- CorrSteer: Generation-Time LLM Steering via Correlated Sparse Autoencoder Features
- When is Your LLM Steerable?
- Reasoning Models Don't Just Think Longer, They Move Differently
- Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders
- SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models
- Measuring Massive Multitask Language Understanding
- Contextual Linear Activation Steering of Language Models
- Scaling Laws for Forgetting When Fine-Tuning Large Language Models
- Emergent World Models and Latent Variable Estimation in Chess-Playing Language Models
- Learning a Generative Meta-Model of LLM Activations
- Don't Lose Focus: Activation Steering via Key-Orthogonal Projections
- The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
- Steered LLM Activations are Non-Surjective
- The Origins of Representation Manifolds in Large Language Models
- Minimizing Collateral Damage in Activation Steering
- Beyond Linear Activation Steering: Invertible Latent Transformations for Controlling LLM Behavior
- Steering Llama 2 via Contrastive Activation Addition
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering