SteerCheck: Attribution Specificity and Alignment Leakage in Activation-Steering Audits
cs.CL
Submitted: 2026-08-25
Updated: 2026-08-25
Terminology
Sources
- Understanding (Un)Reliability of Steering Vectors in Language Models
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- Gemma 3 Technical Report
- Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence
- Towards Reliable Evaluation of Behavior Steering Interventions in LLMs
- Qwen2.5 Technical Report
- Analyzing the Generalization and Reliability of Steering Vectors
- Steering Language Models With Activation Engineering
- Qwen3 Technical Report
- Where Steering Signals Come From: Activation Source Selection in Activation Steering
- Representation Engineering: A Top-Down Approach to AI Transparency
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering