Activation Steering Transfer to Agents: One Gain Ratio Does Not Identify Potency and Efficacy
cs.LG
Submitted: 2026-07-10
Updated: 2026-08-29
Terminology
Sources
- Refusal in Language Models Is Mediated by a Single Direction
- Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering
- Persona Vectors: Monitoring and Controlling Character Traits in Language Models
- Controlling Tool Use with Heading-Specific Activation Steering
- Universal Refusal Circuits Across LLMs: Cross-Model Transfer via Trajectory Replay and Concept-Basis Reconstruction
- GEMS: Geometric Constraints Enable Multi-Semantic Superposition in LLMs
- Internal-State Probes Read the Situation, Not the Action: Three Negative Results for Pre-Action Misalignment Monitoring
- Perfect Detection, Failed Control: The Geometry of Knowing vs. Steering in Language Models
- Gemma 2: Improving Open Language Models at a Practical Size
- Prompt-Activation Duality: Improving Activation Steering via Attention-Level Interventions
- Refused in Chat, Written in Code: Workflow-Level Jailbreak Construction in IDE Coding Agents
- Applying Refusal-Vector Ablation to Llama 3.1 70B Agents
- AgentLens: Interpretable Safety Steering via Mechanistic Subspaces for Multi-Turn Coding Agent
- The Llama 3 Herd of Models
- Tracing Persona Vectors Through LLM Pretraining
- Minimizing Collateral Damage in Activation Steering
- Steering Llama 2 via Contrastive Activation Addition
- Analyzing the Generalization and Reliability of Steering Vectors
- Qwen2.5 Technical Report
- Steering Language Models With Activation Engineering
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks